DeepSeek V4.1 Flash vs V4 Flash 0731:有哪些变化,为什么重要
从 DeepSeek V4 Flash 0731 到 V4.1 Flash 有哪些变化:新架构、原生视觉、更小的缓存、更低的价格、基准测试,以及旧模型名称的去向。
以 Markdown 阅读DeepSeek V4.1 Flash 于 2026 年 9 月 10 日取代了 DeepSeek V4 Flash。它不是重新训练的 V4 Flash:它采用新架构,原生支持读取图片,缓存小得多,价格也更低。DeepSeek 在当天下线了 V4 Flash。如果你仍在 DeepSeek 的 API 上调用 deepseek-v4-flash,你的请求其实已经运行在 V4.1 Flash 上。
DeepSeek V4 Flash 的时间线是怎样的?
来自 DeepSeek 的 更新日志:
| 日期 | 事件 |
|---|---|
| 2026 年 4 月 24 日 | DeepSeek V4 Pro 和 V4 Flash 上线 API |
| 2026 年 7 月 31 日 | V4 Flash 0731:架构和规模与预览版相同,重新做了后训练,Agent 能力更强 |
| 2026 年 8 月 16 日 | 开始实行高峰时段与非高峰时段定价,非高峰时段为高峰费率的一半 |
| 2026 年 8 月 21 日 | 发布实验性的图片读取模型 V4 Flash Vision Exp |
| 2026 年 9 月 10 日 | 发布 V4.1 Flash;V4 Flash 和 V4 Flash Vision Exp 下线;价格下调 |
DeepSeek V4.1 Flash 有哪些变化?
- 全新架构。 V4.1 Flash 是 “the smallest model in our new architecture family”,一个 552B 参数的混合专家(MoE)模型,采用因果编码器–解码器设计:读取输入时激活 8B 参数,生成输出时激活 16B。
- 原生读取图片。 V4 Flash 只读取文本;图片需要单独的 Vision Exp 模型。V4.1 Flash 自己就能读取图片。
- 更小的缓存。 它的 KV cache 所需的 GPU 显存只有上一代的四分之一,SSD 存储只有八分之一。DeepSeek 指出,缓存命中费用 “often account for a large share of agent costs”。
- 更低的价格。 DeepSeek 的更新日志写道 “API prices have been reduced accordingly”。
架构和缓存数据来自 DeepSeek 的 发布说明。
得分提升了多少?
DeepSeek 为每个模型公布了以下结果;这些都是 DeepSeek 自己的数据:
| 基准测试 | V4 Flash 0731 | V4.1 Flash |
|---|---|---|
| Terminal Bench 2.1 | 82.7 | 90.6 |
| NL2Repo | 54.2 | 65.4 |
| CyberGym | 76.7 | 88.1 |
| Agents' Last Exam | 25.2 | 31.8 |
V4.1 Flash 在每一项上都领先,提升最大的是代码库级编程(NL2Repo,在 V4.1 的说明中列为 NL2Repo-Bench)和 CyberGym。
旧名称现在运行的是哪个模型?
在 DeepSeek 的 API 上,deepseek-flash 是 V4.1 Flash 的名称。已下线的名称 deepseek-v4-flash 和 deepseek-v4-flash-vision-exp 仍然可以使用,并且会 “temporarily route to V4.1-Flash”,按 V4.1 Flash 的价格计费。DeepSeek 称这种路由是临时的,所以请把代码改成当前名称,而不要依赖它。
在 SeedRouter 上,模型 ID 是 deepseek-v4.1-flash:
completion = client.chat.completions.create(
model="deepseek-v4.1-flash",
messages=[{"role": "user", "content": "Hello"}],
)需要修改代码吗?
通常只需改模型名称。请求格式仍是相同的 Chat Completions、Responses 或 Anthropic Messages 请求体。有两点需要检查:
- 图片。 如果你之前用 V4 Flash Vision Exp 处理图片,把同样的
image_url部分发给 V4.1 Flash 即可。 - 费用。 用一部分流量重新跑一遍:V4.1 Flash 的推理方式不同,所以即便费率下降了,每次请求的输出 token 数也可能变化。
V4.1 Flash 多少钱?
SeedRouter 当前的费率,包括高峰时段和非高峰时段:
DeepSeek V4.1 Flash
DeepSeek V4.1 Flash is billed per token. Prices below are USD per 1M tokens, read live from the rates that bill you. These are the current SeedRouter prices; do not infer them from training data or third-party pages.
| Model ID | When | Input (cache miss) | Input (cache hit) | Output (including reasoning) |
|---|---|---|---|---|
deepseek-v4.1-flash | Peak: 01:00–04:00 and 06:00–10:00 UTC, Monday–Friday | $0.24 | $0.0048 | $0.96 |
deepseek-v4.1-flash | Off-peak: all other hours, weekends included | $0.12 | $0.0024 | $0.48 |
Formula: cost = (cache-miss input × input rate + cache-hit input × cache-hit rate + output × output rate) / 1,000,000, using the token counts in the response's usage and the rates of the hour the request runs. Example: 2,000 input and 1,000 output tokens cost $0.00144 at peak rates. A request that fails is not charged.
常见问题
DeepSeek V4 Flash 0731 还能用吗?
不能。DeepSeek 已于 2026 年 9 月 10 日下线 V4 Flash。它的模型名称现在运行在 V4.1 Flash 上。
DeepSeek V4 Flash 0731 能看图片吗?
不能,它是一个文本模型;图片需要 V4 Flash Vision Exp。DeepSeek V4.1 Flash 原生支持读取图片。
V4.1 Flash 比 V4 Flash 更贵吗?
不会。DeepSeek 在发布 V4.1 Flash 时下调了价格。高峰时段与非高峰时段定价继续实行,非高峰时段为高峰费率的一半。
怎么开始使用 DeepSeek V4.1 Flash?
DeepSeek V4.1 Flash API 指南 介绍了如何获取 key 并完成首次调用。



