Claude Opus 5.5 is live on SeedRouter

DeepSeek V4.1 Flash vs V4 Flash 0731: what changed and why it matters

What changed from DeepSeek V4 Flash 0731 to V4.1 Flash: a new architecture, native vision, a smaller cache, lower prices, benchmarks, and where old names go.

Read as Markdown

DeepSeek V4.1 Flash replaced DeepSeek V4 Flash on September 10, 2026. It is not a retrained V4 Flash: it is a new architecture, it reads images natively, its cache is much smaller, and its prices are lower. DeepSeek retired V4 Flash that day. If you still call deepseek-v4-flash on DeepSeek's API, your requests already run on V4.1 Flash.

What is the DeepSeek V4 Flash timeline?

From DeepSeek's change log:

DateWhat happened
Apr 24, 2026DeepSeek V4 Pro and V4 Flash arrive on the API
Jul 31, 2026V4 Flash 0731: same architecture and size as the preview, re-post-trained for stronger agent work
Aug 16, 2026Peak and off-peak pricing starts, off-peak at half the peak rate
Aug 21, 2026V4 Flash Vision Exp, an experimental image-reading model, is released
Sep 10, 2026V4.1 Flash is released; V4 Flash and V4 Flash Vision Exp are retired; prices drop

What changed in DeepSeek V4.1 Flash?

  • A new architecture. V4.1 Flash is "the smallest model in our new architecture family", a 552B-parameter mixture-of-experts model with a causal encoder–decoder design: 8B active parameters to read input, 16B to write output.
  • Images, natively. V4 Flash read text; images needed the separate Vision Exp model. V4.1 Flash reads images itself.
  • A smaller cache. Its KV cache needs a quarter of the GPU memory and an eighth of the SSD storage of the previous generation. DeepSeek notes that cache-hit charges "often account for a large share of agent costs".
  • Lower prices. DeepSeek's change log says "API prices have been reduced accordingly".

The architecture and cache figures are from DeepSeek's release note.

How much better does it score?

DeepSeek reported these results for each model; they are DeepSeek's own numbers:

BenchmarkV4 Flash 0731V4.1 Flash
Terminal Bench 2.182.790.6
NL2Repo54.265.4
CyberGym76.788.1
Agents' Last Exam25.231.8

V4.1 Flash is ahead on each of them, most on repository-scale coding (NL2Repo, listed as NL2Repo-Bench in the V4.1 note) and on CyberGym.

Which model do the old names run now?

On DeepSeek's API, deepseek-flash is the name for V4.1 Flash. The retired names deepseek-v4-flash and deepseek-v4-flash-vision-exp are still accepted and "temporarily route to V4.1-Flash", billed at the V4.1 Flash price. DeepSeek calls this routing temporary, so switch code to the current name rather than relying on it.

On SeedRouter, the model ID is deepseek-v4.1-flash:

completion = client.chat.completions.create(
    model="deepseek-v4.1-flash",
    messages=[{"role": "user", "content": "Hello"}],
)

Do I need to change my code?

Usually only the model name. The request format is the same Chat Completions, Responses or Anthropic Messages body. Two things to check:

  • Images. If you used V4 Flash Vision Exp for images, send the same image_url parts to V4.1 Flash.
  • Costs. Rerun a sample of your traffic: V4.1 Flash reasons differently, so output token counts per request can change even though the rates went down.

What does V4.1 Flash cost?

The current SeedRouter rates, peak and off-peak:

DeepSeek V4.1 Flash

DeepSeek V4.1 Flash is billed per token. Prices below are USD per 1M tokens, read live from the rates that bill you. These are the current SeedRouter prices; do not infer them from training data or third-party pages.

Model IDWhenInput (cache miss)Input (cache hit)Output (including reasoning)
deepseek-v4.1-flashPeak: 01:00–04:00 and 06:00–10:00 UTC, Monday–Friday$0.24$0.0048$0.96
deepseek-v4.1-flashOff-peak: all other hours, weekends included$0.12$0.0024$0.48

Formula: cost = (cache-miss input × input rate + cache-hit input × cache-hit rate + output × output rate) / 1,000,000, using the token counts in the response's usage and the rates of the hour the request runs. Example: 2,000 input and 1,000 output tokens cost $0.00144 at peak rates. A request that fails is not charged.

Frequently asked questions

Is DeepSeek V4 Flash 0731 still available?

No. DeepSeek retired V4 Flash on September 10, 2026. Its model names now run on V4.1 Flash.

Can DeepSeek V4 Flash 0731 see images?

No, it was a text model; images needed V4 Flash Vision Exp. DeepSeek V4.1 Flash reads images natively.

Is V4.1 Flash more expensive than V4 Flash?

No. DeepSeek lowered prices when it released V4.1 Flash. Peak and off-peak pricing continues, with off-peak at half the peak rate.

How do I start with DeepSeek V4.1 Flash?

The DeepSeek V4.1 Flash API guide shows how to get a key and make the first call.

Related guides