# DeepSeek V4.1 Flash vs V4 Pro: price, vision and benchmarks

By SeedRouter · Published 2026-09-28 · Updated 2026-09-28

DeepSeek V4.1 Flash is the newer, cheaper and, on most of DeepSeek's own agent benchmarks, stronger model. At DeepSeek's list prices, V4 Pro costs 4.4 times as much for input and 3.3 times as much for output. V4.1 Flash also reads images, which V4 Pro does not. V4 Pro keeps one clear lead in DeepSeek's numbers: reasoning without tools on the HLE exam. For most API work, V4.1 Flash is the better buy.

## How do they compare at a glance?

|                                     | DeepSeek V4.1 Flash                           | DeepSeek V4 Pro                               |
| ----------------------------------- | --------------------------------------------- | --------------------------------------------- |
| Version                             | DeepSeek-V4.1-Flash, released Sep 10, 2026    | DeepSeek-V4-Pro-0813, released Aug 13, 2026   |
| DeepSeek model name                 | `deepseek-flash`                              | `deepseek-v4-pro`                             |
| Context / max output                | 1M / 384K tokens                              | 1M / 384K tokens                              |
| Image input                         | Yes                                           | Not supported                                 |
| Thinking                            | On by default, `low` / `high` / `max`, or off | On by default, `low` / `high` / `max`, or off |
| Concurrency limit on DeepSeek's API | 2,500                                         | 500                                           |
| List price, input (cache miss)      | 1×                                            | 4.4×                                          |
| List price, input (cache hit)       | 1×                                            | 7.3×                                          |
| List price, output                  | 1×                                            | 3.3×                                          |

Figures are from DeepSeek's [Models & Pricing page](https://api-docs.deepseek.com/quick_start/pricing) and [change log](https://api-docs.deepseek.com/updates). Price rows show V4 Pro's list price as a multiple of V4.1 Flash's, at the same hour; both models bill off-peak hours at half the peak rate.

## Which one scores higher?

DeepSeek published benchmark results with each release. These are DeepSeek's own numbers, on the benchmarks both releases report (NL2Repo is called NL2Repo-Bench in the V4.1 Flash note):

| Benchmark          | V4.1 Flash | V4 Pro (GA) |
| ------------------ | ---------- | ----------- |
| Terminal Bench 2.1 | 90.6       | 87.9        |
| NL2Repo            | 65.4       | 61.5        |
| CyberGym           | 88.1       | 83.3        |
| Agents' Last Exam  | 31.8       | 25.7        |
| HLE, with tools    | 63.9       | 60.0        |
| HLE, without tools | 36.8       | 42.7        |

V4.1 Flash leads on the agent and coding benchmarks. V4 Pro leads on HLE without tools, a test of hard reasoning and knowledge. DeepSeek's release note for V4.1 Flash says it delivers "benchmark results ahead of flagship models, including DeepSeek-V4-Pro". Vendor benchmarks are a starting point; run your own prompts before you switch.

## Why is Flash so much cheaper?

DeepSeek built V4.1 Flash on a new architecture that reads with 8B active parameters and writes with 16B, out of 552B in total. Its KV cache needs a quarter of the GPU memory and an eighth of the SSD storage of the previous generation, according to the [release note](https://api-docs.deepseek.com/news/news260910). Less compute and less cache per request is what the lower price passes on, and it matters most for agents, where cache hits make up much of the input.

## Is DeepSeek V4 Pro being retired?

Not for now. DeepSeek first announced that V4 Pro requests would move to V4.1 Flash from September 14, 2026. Its change log then says it decided "to continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged". V4 Flash, the previous Flash model, has been retired, and its old names run on V4.1 Flash.

## Which should I use?

* **Use DeepSeek V4.1 Flash** for coding agents, tool use, image input, high-volume jobs and anything where price and throughput matter. That is most API work.
* **Consider DeepSeek V4 Pro** only if your own tests show it answering hard knowledge or reasoning questions without tools noticeably better, and the higher price is worth it.

SeedRouter serves DeepSeek V4.1 Flash. The live rates:

### DeepSeek V4.1 Flash

Current rates are temporarily unavailable. No price estimate is provided; unavailable rates must not be interpreted as free usage.

## Frequently asked questions

### Is DeepSeek V4.1 Flash better than V4 Pro?

On DeepSeek's own agent and coding benchmarks, yes; on HLE without tools, V4 Pro scores higher. It is also 3 to 4 times cheaper per token and reads images.

### Can DeepSeek V4 Pro read images?

No. DeepSeek's pricing page lists vision as "Not supported" for V4 Pro. V4.1 Flash reads images natively.

### Do both models have the same context window?

Yes. Both take up to 1M tokens of context and write up to 384K tokens.

### How do I start with DeepSeek V4.1 Flash?

The [DeepSeek V4.1 Flash API guide](https://seedrouter.ai/blog/deepseek-v4-1-flash-api) shows how to get a key and make the first call.
