# DeepSeek V4.1 Flash specs: context window, max output, size and defaults

By SeedRouter · Published 2026-09-29 · Updated 2026-09-29

DeepSeek V4.1 Flash is a 552-billion-parameter mixture-of-experts model with a 1M-token context window and a maximum output of 393,216 tokens (384K) per request. It reads text and images, writes text, and thinks before it answers unless you turn thinking off. DeepSeek released it on September 10, 2026. The spec most people trip over is the default output limit, which is far lower than the maximum.

## DeepSeek V4.1 Flash specs at a glance

| Spec                       | Value                                                          |
| -------------------------- | -------------------------------------------------------------- |
| Developer                  | DeepSeek                                                       |
| Released                   | September 10, 2026                                             |
| Architecture               | Mixture of experts, Causal Encoder–Decoder                     |
| Total parameters           | 552B                                                           |
| Active parameters          | 8B for input, 16B for output                                   |
| Context window             | 1M tokens                                                      |
| Max output                 | 393,216 tokens (384K), reasoning included                      |
| Default output limit       | 8K without thinking, 64K with thinking, 128K at `max` effort   |
| Input                      | Text and images                                                |
| Output                     | Text                                                           |
| Thinking                   | On by default at `high` effort; `none`, `low`, `high` or `max` |
| Model ID on SeedRouter     | `deepseek-v4.1-flash`                                          |
| Model ID on DeepSeek's API | `deepseek-flash`                                               |

## How big is DeepSeek V4.1 Flash?

552 billion parameters in total. DeepSeek's release note describes a "new Causal Encoder–Decoder architecture: just 8B active parameters for input, 16B for output." Only a small part of the model runs for each token, which is why it is fast and cheap to serve despite its size.

DeepSeek also says its KV cache needs "1/4 the HBM" and "1/8 the SSD storage" of the previous generation. That matters for long agent sessions, where cached input makes up much of the cost.

## What is the DeepSeek V4.1 Flash context window?

1M tokens. The prompt, images, conversation history and the reply all share it. DeepSeek's API reference puts it this way: "The total length of input tokens and generated tokens is limited by the model's context length."

For how context and output limits interact across models, see [context window vs max output tokens](https://seedrouter.ai/blog/context-window-vs-max-output-tokens).

## What is the DeepSeek V4.1 Flash max output?

393,216 tokens per request, which DeepSeek lists as "MAXIMUM: 384K". The reasoning counts toward it.

The catch is the default. If you do not set `max_tokens`, DeepSeek's reference says the limit is "8K in non-thinking mode, 64K in thinking mode (128K with reasoning\_effort set to max)". A long answer or a long reasoning trace hits that default well before the real ceiling, and the reply ends with `finish_reason: "length"`.

To get longer replies, set `max_tokens` yourself, anywhere from 1 to 393216:

```json
{
  "model": "deepseek-v4.1-flash",
  "max_tokens": 200000,
  "messages": [{"role": "user", "content": "Write the full migration plan."}]
}
```

You pay only for the tokens the model writes, so a high limit does not cost more on its own.

## Does DeepSeek V4.1 Flash read images?

Yes. DeepSeek calls it a model "with native visual understanding". On SeedRouter, images go in a user message as `image_url` parts, either a public URL or a base64 data URI. A URL can be up to 8,192 characters long and point to an image of up to 32 MiB. The output is always text.

## How do thinking modes work?

Thinking is on by default at `high` effort. You can:

* turn it off with `"thinking": {"type": "disabled"}` or `"reasoning_effort": "none"`, for fast answers with fewer output tokens;
* set `reasoning_effort` to `low`, `high` or `max`.

The reasoning comes back in `reasoning_content`, next to the answer, and is billed as output tokens. With thinking on, `temperature` has no effect and `top_p` values below 0.95 run as 0.95.

## Which APIs can call it?

On SeedRouter, DeepSeek V4.1 Flash takes three request formats with the same key:

| Format                  | Endpoint                                             |
| ----------------------- | ---------------------------------------------------- |
| OpenAI Chat Completions | `POST https://api.seedrouter.ai/v1/chat/completions` |
| OpenAI Responses        | `POST https://api.seedrouter.ai/v1/responses`        |
| Anthropic Messages      | `POST https://api.seedrouter.ai/v1/messages`         |

Tool calls work in all three. The Anthropic format is what Claude Code uses, and the Responses format is what Codex uses; the [Claude Code and Codex setup](https://seedrouter.ai/blog/deepseek-v4-1-flash-claude-code-codex) has tested configs. Every parameter is listed in the [DeepSeek V4.1 Flash API reference](https://seedrouter.ai/docs/deepseek-v4-1-flash).

## How is it priced?

Per token, with separate rates for input that hits the cache, input that misses it, and output. Rates halve off-peak: peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday, and every other hour is off-peak. The [DeepSeek V4.1 Flash page](https://seedrouter.ai/models/deepseek-v4-1-flash#pricing) shows the live rates, and the [pricing guide](https://seedrouter.ai/blog/deepseek-v4-1-flash-api-pricing) works through examples.

## Frequently asked questions

### Is DeepSeek V4.1 Flash open source?

DeepSeek publishes the weights on Hugging Face; [Is DeepSeek V4.1 Flash free?](https://seedrouter.ai/blog/is-deepseek-v4-1-flash-free) covers what that means for cost.

### Is `deepseek-v4-flash` the same model?

On DeepSeek's own API, the legacy name now routes to V4.1 Flash. DeepSeek's pricing page says "the corresponding models have been retired, their requests are served by the DeepSeek-V4.1-Flash model". [V4.1 Flash vs V4 Flash](https://seedrouter.ai/blog/deepseek-v4-1-flash-vs-v4-flash) lists what changed.

### Why does DeepSeek V4.1 Flash stop at 8K tokens?

That is the default output limit with thinking off. Set `max_tokens` up to 393216 to allow longer replies.

### Does the context window include the output?

Yes. Input and output share the 1M-token window.

### How does it compare with DeepSeek V4 Pro?

DeepSeek reports V4.1 Flash ahead of V4 Pro on its benchmarks, and it costs less. See [DeepSeek V4.1 Flash vs V4 Pro](https://seedrouter.ai/blog/deepseek-v4-1-flash-vs-pro).
