Context window vs max output tokens: limits for GPT-6, Claude, DeepSeek and Kimi
What context window and max output tokens mean, the limits for GPT-6, Claude, DeepSeek V4.1 Flash and Kimi K3, and how to fix replies that stop early.
Read as MarkdownThe context window is the total number of tokens a model can hold in one request: your prompt, the conversation history, images, and the reply it writes. The max output tokens limit is the most the model may write in that reply, and on reasoning models it includes the tokens spent thinking. The reply has to fit inside the context window, so a long prompt leaves less room for output.
What is the difference between context window and max output tokens?
The context window is shared space. Everything you send and everything the model writes in one request must fit in it.
Max output tokens is a separate, smaller cap on the reply alone. Each model has a ceiling, and each API has a parameter that lets you set a lower limit per request.
Two consequences follow:
- Input and output compete for the same window. DeepSeek's API reference puts it plainly: "The total length of input tokens and generated tokens is limited by the model's context length."
- Reasoning counts as output. On GPT-6, Claude, DeepSeek V4.1 Flash and Kimi K3, the tokens a model spends thinking count toward the output limit and are billed as output. A reply can stop early even though the visible answer is short.
What are the limits for each model?
| Model | Model ID | Context window | Max output | Output parameter | Default if you leave it out |
|---|---|---|---|---|---|
| GPT-6 Astra, Sol, Luna | gpt-6-astra, gpt-6-sol, gpt-6-luna | 1.05M tokens (922K input) | 128K | max_output_tokens (Responses), max_completion_tokens (Chat Completions) | Model maximum |
| Claude Opus 5.5, Fable 5.1, Fable 5 | claude-opus-5-5, claude-fable-5-1, claude-fable-5 | 1M tokens | 128K | max_tokens | None: the field is required |
| DeepSeek V4.1 Flash | deepseek-v4.1-flash | 1M tokens | 393,216 | max_tokens | 8K without thinking, 64K with thinking, 128K at max effort |
| Kimi K3 | kimi-k3 | 1,048,576 tokens | 1,048,576 | max_completion_tokens | 131,072 |
The last column is where most cut-off replies come from. DeepSeek V4.1 Flash can write 393,216 tokens, but if you do not set max_tokens it stops at 8K, or 64K with thinking on. Kimi K3 can write up to 1,048,576 tokens but defaults to 131,072.
Each model's API reference has the full parameter list: GPT-6 Astra, Claude Opus 5.5, DeepSeek V4.1 Flash and Kimi K3.
max_tokens, max_completion_tokens or max_output_tokens?
They all cap the reply. Which one to send depends on the API format and the model:
| API format | Parameter | Notes |
|---|---|---|
OpenAI Responses (/v1/responses) | max_output_tokens | OpenAI: "An upper bound for the number of tokens that can be generated for a response, including visible output tokens and reasoning tokens." |
OpenAI Chat Completions (/v1/chat/completions) | max_completion_tokens | OpenAI marks max_tokens as "deprecated in favor of max_completion_tokens" and "not compatible with o-series models". Use max_completion_tokens for GPT-6. |
Anthropic Messages (/v1/messages) | max_tokens | Required on every request. |
| DeepSeek Chat Completions | max_tokens | 1 to 393216. |
| Kimi Chat Completions | max_completion_tokens | max_tokens is the deprecated name for the same limit. |
If an API answers "max_tokens is not supported with this model", switch to the parameter in this table.
Why did my model stop before finishing?
It hit the output limit. Every API reports this in the response, not as an error:
| API | Field | Value when the output limit was reached |
|---|---|---|
| OpenAI Responses | incomplete_details.reason | "max_output_tokens" |
| OpenAI Chat Completions | choices[].finish_reason | "length" |
| Anthropic Messages | stop_reason | "max_tokens" |
| DeepSeek | choices[].finish_reason | "length" |
Anthropic has a second stop reason, model_context_window_exceeded, for a reply that filled the whole context window. DeepSeek's length covers both cases: the reply exceeded max_tokens or the conversation exceeded the context length.
To fix it:
- Raise the output limit up to the model's maximum in the table above.
- Lower the reasoning effort. Less thinking leaves more of the budget for the answer, and costs less.
- Continue instead of retrying. Send the partial reply back and ask the model to go on. Anthropic's stop-reason guide describes this for
max_tokens.
What does "context window exceeded" mean?
Your request does not fit in the model's window. Where exactly it fails depends on the API:
- Claude. If the input alone is larger than the window, the API returns a 400
invalid_request_error("prompt is too long"). If only input plusmax_tokensis larger, Anthropic's docs say Claude 4.5 and newer models, including the ones in this guide, accept the request and stop withstop_reason: "model_context_window_exceeded"if they run out of room. - DeepSeek. The reply ends with
finish_reason: "length"when the conversation exceeds the context length.
Three fixes, in order:
- Drop or summarise old turns of the conversation. This is the only fix when the input alone is too long.
- Lower the output limit so input plus output fits.
- Move to a model with a larger window. All four families in the table above read about 1M tokens.
How do coding agents handle these limits?
Coding agents set the output limit for you, and their defaults can be lower than the model allows.
Claude Code uses CLAUDE_CODE_MAX_OUTPUT_TOKENS. Its documentation says it "defaults to 32000 for model IDs it doesn't recognize, such as gateway-specific names, and lowers values above a model's cap to the cap". When you run DeepSeek V4.1 Flash or Kimi K3 in Claude Code, set the variable if you need longer replies. The same documentation warns that a higher value "reduces the effective context window available before auto-compaction triggers".
Codex reads model_context_window from config.toml, described as the "context window tokens available to the active model". Set it for models Codex does not know, as in our Kimi K3 setup and DeepSeek V4.1 Flash setup. For a model name it does not recognise, Codex also prints: "Model metadata for <model> not found. Defaulting to fallback metadata; this can degrade performance and cause issues." The session still runs.
Frequently asked questions
Does the context window include output tokens?
Yes. Input, history, images and the reply all share one context window. The max output limit is a separate cap on the reply inside that window.
Do reasoning tokens count toward max output tokens?
Yes, on all four model families here. Thinking tokens count toward the output limit and are billed as output tokens.
Which model writes the longest replies?
Kimi K3 accepts max_completion_tokens up to 1,048,576, and DeepSeek V4.1 Flash accepts max_tokens up to 393,216. GPT-6 and Claude stop at 128K per request.
Does a bigger max output limit cost more?
No. You pay for the tokens the model actually writes, not for the limit you set. A higher limit only matters when the model needs the room.
Where can I see live prices for these models?
On each model page: GPT-6 Astra, Claude Opus 5.5, DeepSeek V4.1 Flash and Kimi K3.



