Claude Opus 5.5 is live on SeedRouter

DeepSeek V4.1 Flash specs: context window, max output, size and defaults

DeepSeek V4.1 Flash specs: 552B-parameter MoE, 1M-token context, 384K max output, image input, thinking modes and the defaults that cut replies short.

Read as Markdown

DeepSeek V4.1 Flash is a 552-billion-parameter mixture-of-experts model with a 1M-token context window and a maximum output of 393,216 tokens (384K) per request. It reads text and images, writes text, and thinks before it answers unless you turn thinking off. DeepSeek released it on September 10, 2026. The spec most people trip over is the default output limit, which is far lower than the maximum.

DeepSeek V4.1 Flash specs at a glance

SpecValue
DeveloperDeepSeek
ReleasedSeptember 10, 2026
ArchitectureMixture of experts, Causal Encoder–Decoder
Total parameters552B
Active parameters8B for input, 16B for output
Context window1M tokens
Max output393,216 tokens (384K), reasoning included
Default output limit8K without thinking, 64K with thinking, 128K at max effort
InputText and images
OutputText
ThinkingOn by default at high effort; none, low, high or max
Model ID on SeedRouterdeepseek-v4.1-flash
Model ID on DeepSeek's APIdeepseek-flash

How big is DeepSeek V4.1 Flash?

552 billion parameters in total. DeepSeek's release note describes a "new Causal Encoder–Decoder architecture: just 8B active parameters for input, 16B for output." Only a small part of the model runs for each token, which is why it is fast and cheap to serve despite its size.

DeepSeek also says its KV cache needs "1/4 the HBM" and "1/8 the SSD storage" of the previous generation. That matters for long agent sessions, where cached input makes up much of the cost.

What is the DeepSeek V4.1 Flash context window?

1M tokens. The prompt, images, conversation history and the reply all share it. DeepSeek's API reference puts it this way: "The total length of input tokens and generated tokens is limited by the model's context length."

For how context and output limits interact across models, see context window vs max output tokens.

What is the DeepSeek V4.1 Flash max output?

393,216 tokens per request, which DeepSeek lists as "MAXIMUM: 384K". The reasoning counts toward it.

The catch is the default. If you do not set max_tokens, DeepSeek's reference says the limit is "8K in non-thinking mode, 64K in thinking mode (128K with reasoning_effort set to max)". A long answer or a long reasoning trace hits that default well before the real ceiling, and the reply ends with finish_reason: "length".

To get longer replies, set max_tokens yourself, anywhere from 1 to 393216:

{
  "model": "deepseek-v4.1-flash",
  "max_tokens": 200000,
  "messages": [{"role": "user", "content": "Write the full migration plan."}]
}

You pay only for the tokens the model writes, so a high limit does not cost more on its own.

Does DeepSeek V4.1 Flash read images?

Yes. DeepSeek calls it a model "with native visual understanding". On SeedRouter, images go in a user message as image_url parts, either a public URL or a base64 data URI. A URL can be up to 8,192 characters long and point to an image of up to 32 MiB. The output is always text.

How do thinking modes work?

Thinking is on by default at high effort. You can:

  • turn it off with "thinking": {"type": "disabled"} or "reasoning_effort": "none", for fast answers with fewer output tokens;
  • set reasoning_effort to low, high or max.

The reasoning comes back in reasoning_content, next to the answer, and is billed as output tokens. With thinking on, temperature has no effect and top_p values below 0.95 run as 0.95.

Which APIs can call it?

On SeedRouter, DeepSeek V4.1 Flash takes three request formats with the same key:

FormatEndpoint
OpenAI Chat CompletionsPOST https://api.seedrouter.ai/v1/chat/completions
OpenAI ResponsesPOST https://api.seedrouter.ai/v1/responses
Anthropic MessagesPOST https://api.seedrouter.ai/v1/messages

Tool calls work in all three. The Anthropic format is what Claude Code uses, and the Responses format is what Codex uses; the Claude Code and Codex setup has tested configs. Every parameter is listed in the DeepSeek V4.1 Flash API reference.

How is it priced?

Per token, with separate rates for input that hits the cache, input that misses it, and output. Rates halve off-peak: peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday, and every other hour is off-peak. The DeepSeek V4.1 Flash page shows the live rates, and the pricing guide works through examples.

Frequently asked questions

Is DeepSeek V4.1 Flash open source?

DeepSeek publishes the weights on Hugging Face; Is DeepSeek V4.1 Flash free? covers what that means for cost.

Is deepseek-v4-flash the same model?

On DeepSeek's own API, the legacy name now routes to V4.1 Flash. DeepSeek's pricing page says "the corresponding models have been retired, their requests are served by the DeepSeek-V4.1-Flash model". V4.1 Flash vs V4 Flash lists what changed.

Why does DeepSeek V4.1 Flash stop at 8K tokens?

That is the default output limit with thinking off. Set max_tokens up to 393216 to allow longer replies.

Does the context window include the output?

Yes. Input and output share the 1M-token window.

How does it compare with DeepSeek V4 Pro?

DeepSeek reports V4.1 Flash ahead of V4 Pro on its benchmarks, and it costs less. See DeepSeek V4.1 Flash vs V4 Pro.

Related guides