DeepSeek V4.1 Flash — 快速、低价、支持图片的推理模型
DeepSeek V4.1 Flash 是 DeepSeek 推出的快速、低价模型,支持 100 万 token 上下文和图片输入。高峰时段每百万 token 价格:$0.24 / 1M 输入、$0.96 / 1M 输出;非高峰时段半价。
- 思考可开可关
- 百万 token 上下文
- 图片输入
- 非高峰时段价格
- 流式输出
试试 DeepSeek V4.1 Flash
和 DeepSeek V4.1 Flash 聊天、附加图片,实时查看每个回答的 token 数和费用。
deepseek-v4.1-flash$0.24 输入 · $0.96 输出,每百万 token (当前为高峰时段)
开启思考时,temperature 不生效,低于 0.95 的 top_p 按 0.95 执行。
参数会被接受,但不会返回对数概率。
DeepSeek V4.1 Flash 的真实输出
一个提示词,在关闭思考的情况下发送给 DeepSeek V4.1 Flash。未经编辑。
Explain the difference between a mutex and a semaphore in two short paragraphs, and say when to use each.
A mutex is a locking mechanism designed for mutual exclusion: only the thread that locked it may unlock it, and it has a notion of ownership. It's binary (locked/unlocked) and is used to protect a shared resource so that only one thread accesses it at a time. A semaphore is a signaling mechanism built around a counter: it tracks how many permits are available, any thread can wait (decrement) or post (increment) it, and it doesn't care which thread does what. A binary semaphore looks similar to a mutex, but lacks ownership, so it can be released by a different thread—useful for signaling between threads rather than guarding a critical section. Use a mutex when you need to protect a shared resource (a data structure, file, or device) from concurrent access, which is the common case for locking critical sections. Use a semaphore when you need to control access to a pool of N identical resources (e.g., limiting concurrent database connections) or to signal between threads, such as a producer notifying a consumer that data is ready or a task-completion event. As a rule of thumb: mutex for mutual exclusion, semaphore for counting or signaling.
生成于 2026-09-28。每次运行的回答都可能不同。
DeepSeek V4.1 Flash 价格
DeepSeek V4.1 Flash 按 token 计费,从你的余额中扣除,按请求所在时段适用高峰或非高峰费率。下方费率为实时价格,与你的账户实际扣费费率相同。
示例用量仅供说明,按高峰时段费率计算。实际账单以每个请求报告的 token 数及其运行时段为准。
| 项目 | 计为 | 价格 | 计费方式 |
|---|---|---|---|
| 输入,高峰时段 | 未命中缓存的提示 token,UTC 周一至周五 01:00–04:00 和 06:00–10:00 | $0.24 / 1M | 每百万 token |
| 输出,高峰时段 | 高峰时段的回答 token 与推理 token | $0.96 / 1M | 每百万 token |
| 缓存输入,高峰时段 | 高峰时段命中缓存的提示 token | $0.0048 / 1M | 每百万 token |
| 输入,非高峰时段 | 未命中缓存的提示 token,其余所有时段,含周末 | $0.12 / 1M | 每百万 token |
| 输出,非高峰时段 | 非高峰时段的回答 token 与推理 token | $0.48 / 1M | 每百万 token |
| 缓存输入,非高峰时段 | 非高峰时段命中缓存的提示 token | $0.0024 / 1M | 每百万 token |
| 失败的请求 | 任何返回错误的请求 | 免费 | 不收费 |
DeepSeek V4.1 Flash 是什么?
DeepSeek V4.1 Flash 是 DeepSeek 推出的快速、低价模型,于 2026 年 9 月 10 日发布。它是一个 5520 亿参数的混合专家(MoE)模型,读取时激活 80 亿参数、生成时激活 160 亿参数,原生支持图片输入,最多可接收 100 万 token 的上下文。
DeepSeek V4.1 Flash 默认先思考再作答,推理强度为 high。简单快速的步骤可以关闭思考,难题则把推理强度调到 max。在 DeepSeek 自己的 API 中,这个模型叫 deepseek-flash。
模型 ID:deepseek-v4.1-flash · 输入:文本、图片 · 输出:文本 · 接口格式:Chat Completions、Responses、Anthropic Messages。
使用 DeepSeek V4.1 Flash 的两种方式
从你的代码调用它,或让你的编码助手来做。
从后端调用
把 OpenAI SDK 指向 SeedRouter,代码无需改动。DeepSeek V4.1 Flash 支持 Chat Completions、Responses 和 Anthropic Messages 三种格式。
- 1创建 API Key
- 2把 base URL 设为 https://api.seedrouter.ai/v1
- 3将 model 设为 deepseek-v4.1-flash
交给你的编码助手
复制一个现成的提示词。你的助手会在执行任何操作前显示请求和费用。
- 1复制提示词
- 2粘贴到你的助手
- 3批准请求
DeepSeek V4.1 Flash 擅长什么
DeepSeek V4.1 Flash 专为看重速度和费用的大批量工作打造。
代理循环
DeepSeek V4.1 Flash 能以低廉的价格运行大量步骤,缓存的上下文让重复读取保持低价。
日常编码
在 Codex、Claude Code 或你自己的工具中,用 DeepSeek V4.1 Flash 编写、审查和修复代码。
提示中带图片
DeepSeek V4.1 Flash 能读懂通过 URL 或 base64 传入的截图、图表和照片。
结构化输出
让 DeepSeek V4.1 Flash 输出 json_object,直接解析回答。
思考可开可关
简单问题跳过 DeepSeek V4.1 Flash 的推理快速作答,难题则调高推理强度。
非高峰时段费率
在高峰时段之外运行的批量任务,只需支付高峰价格的一半。
团队用 DeepSeek V4.1 Flash 构建什么
大规模分类
关闭思考,用 DeepSeek V4.1 Flash 给工单、评论和日志打标签,费用最低。
编码代理
在 Codex 中运行 DeepSeek V4.1 Flash,实现快速、低价的“编辑—运行”循环。
截图识别
把截图和图表转换成文本或 JSON。
四步开始使用 DeepSeek V4.1 Flash
- 01
创建 Key
登录并创建 API Key。随时添加余额;没有订阅。
- 02
更改基础 URL
代码保持不变,只需把 OpenAI SDK 的 base_url 指向 https://api.seedrouter.ai/v1。
- 03
选择 DeepSeek V4.1 Flash
将 model 设为 deepseek-v4.1-flash。可按请求关闭思考或设置推理强度。
- 04
检查账单
每个请求在你的使用记录中显示其 token 数和费用。
SeedRouter 上的 DeepSeek V4.1 Flash 一览
你现在可以使用的功能。
| 功能 | DeepSeek V4.1 Flash |
|---|---|
| Chat Completions | 支持,官方格式 |
| Responses | 支持,可用于 Codex |
| Anthropic Messages | 支持,可用于 Claude Code |
| 图片输入 | 是,支持 URL 或 base64 |
| 流式输出 | 是 |
| 上下文缓存 | 支持,自动生效,命中缓存按更低费率计费 |
| JSON 输出 | 支持,json_object |
| 失败的请求 | 不收费 |
不提供 FIM 与对话前缀续写(beta)、Files API 和对数概率。
DeepSeek V4.1 Flash 的使用限制
百万 token 上下文
在 DeepSeek V4.1 Flash 上,提示、图片和历史记录共用一个 100 万 token 的窗口。
384K token 输出
单次 DeepSeek V4.1 Flash 请求最多输出 393,216 token,包含推理内容。
开启思考时的采样参数
开启思考时,temperature 不生效,top_p 保持在 0.95 或以上。
图片大小
DeepSeek V4.1 Flash 的图片 URL 所指向的文件最大可达 32 MiB。
为什么通过 SeedRouter 调用 DeepSeek V4.1 Flash
一把 key,三种格式
用同一个 Key,即可通过 Chat Completions、Responses 或 Anthropic Messages 调用 DeepSeek V4.1 Flash。
按 token 付费
一次充值,任意模型都能用。无套餐,无月费。
失败请求不收费
返回错误的请求不会被扣费。
实时价格
本页显示的高峰与非高峰时段费率,就是你的账户支付的费率。
上下文缓存
重复的提示前缀会从缓存读取,按更低的 DeepSeek V4.1 Flash 费率计费。
回复中返回推理内容
开启思考时,推理内容会在回答旁边的 reasoning_content 中返回。
流式输出
设置 stream: true,即可边生成边接收 DeepSeek V4.1 Flash 的 token。
相关模型
使用同一把 key 可以调用的其他模型。
用 OpenAI 或 Anthropic SDK 调用 Kimi K3:100 万 token 上下文,推理始终开启,强度可选 low、high 或 max,按 token 实时计价,Playground 在线试用。
查看价格DeepSeek V4.1 Flash 使用指南
该模型的教程和对比。
DeepSeek V4.1 Flash API 调用教程:获取 API key,用 OpenAI SDK 调用,开启或关闭思考模式、流式输出、发送图片,并解决新手常见错误。
DeepSeek V4.1 Flash API 价格(每百万 token):高峰时段与非高峰时段费率、缓存命中、各费率适用的时间、计费公式以及一次请求的实际成本。
如何通过 SeedRouter 把 DeepSeek V4.1 Flash 设为 Claude Code 和 OpenAI Codex CLI 的模型:环境变量、config.toml 配置以及我们的实测结果。
DeepSeek V4.1 Flash 与 DeepSeek V4 Pro 对比:每 token 价格、图片输入、并发上限和 DeepSeek 公布的基准测试成绩,并给出不同任务的明确选择。
从 DeepSeek V4 Flash 0731 到 V4.1 Flash 有哪些变化:新架构、原生视觉、更小的缓存、更低的价格、基准测试,以及旧模型名称的去向。
DeepSeek V4.1 Flash 免费吗?权重可以免费下载,API 按 token 低价计费,这里告诉你今天如何几乎零成本地试用它。
DeepSeek V4.1 Flash 常见问题
DeepSeek V4.1 Flash 是什么?+
DeepSeek V4.1 Flash 是 DeepSeek 推出的快速、低价模型,于 2026 年 9 月 10 日发布。它是一个 5520 亿参数的混合专家(MoE)模型,原生支持图片输入,拥有 100 万 token 的上下文窗口,思考可开可关。
DeepSeek V4.1 Flash 多少钱?+
DeepSeek V4.1 Flash 在高峰时段(UTC 周一至周五 01:00–04:00 和 06:00–10:00)按每百万 token $0.24 / 1M 输入、$0.96 / 1M 输出计费。其余所有时段为非高峰时段,费率减半;命中缓存的价格只是输入的一小部分。
DeepSeek V4.1 Flash 可以免费用吗?+
每个新注册的 SeedRouter 账户都自带 $0.10 免费余额,足够发送许多条简短请求来试用 DeepSeek V4.1 Flash。之后按 token 付费,没有订阅,失败的请求不收费。
DeepSeek V4.1 Flash 支持图片吗?+
支持。DeepSeek V4.1 Flash 原生支持读取图片,可在请求中通过公开 URL 或 base64 传入;它返回文本。
怎么关闭 DeepSeek V4.1 Flash 的思考?+
把 thinking 设为 disabled,或把 reasoning_effort 设为 none。这样回答会立即返回,并且消耗更少的输出 token。
DeepSeek V4.1 Flash 和 DeepSeek V4 Flash 是同一个模型吗?+
不是。DeepSeek V4.1 Flash 于 2026 年 9 月 10 日取代了 V4 Flash,采用新架构并原生支持图片输入。DeepSeek 已下线 V4 Flash,并把它的旧模型名称指向 DeepSeek V4.1 Flash。
DeepSeek V4.1 Flash 能本地部署吗?+
它的权重已发布在 Hugging Face 上。它拥有 5520 亿参数,需要多 GPU 硬件才能运行,所以大多数团队会通过 API 调用 DeepSeek V4.1 Flash。
DeepSeek V4.1 Flash 的模型 ID 是什么?怎么调用?+
在 SeedRouter 上,模型 ID 是 deepseek-v4.1-flash。创建一个 SeedRouter API Key,把 OpenAI SDK 指向 https://api.seedrouter.ai/v1,然后将 model 设为 deepseek-v4.1-flash 发送请求。
DeepSeek V4.1 Flash 请求失败会收费吗?+
不会。返回错误的 DeepSeek V4.1 Flash 请求不收费。你只需为成功完成的请求所报告的 token 付费。
立即试用 DeepSeek V4.1 Flash
几分钟内,就能在 Playground 或你自己的代码里发出第一个 DeepSeek V4.1 Flash 请求。
