Claude Opus 5.5 已在 SeedRouter 上线
DeepSeek文本生成

DeepSeek V4.1 Flash — 快速、低价、支持图片的推理模型

DeepSeek V4.1 Flash 是 DeepSeek 推出的快速、低价模型,支持 100 万 token 上下文和图片输入。高峰时段每百万 token 价格:$0.24 / 1M 输入、$0.96 / 1M 输出;非高峰时段半价。

可用$0.24 / 1M 输入 · $0.96 / 1M 输出$0.30 / 1M 输入 · $1.20 / 1M 输出-20%
  • 思考可开可关
  • 百万 token 上下文
  • 图片输入
  • 非高峰时段价格
  • 流式输出
上下文窗口
100 万 token
最大输出
384K token
推理强度等级
none、low、high、max
发布日期
2026 年 9 月 10 日

试试 DeepSeek V4.1 Flash

和 DeepSeek V4.1 Flash 聊天、附加图片,实时查看每个回答的 token 数和费用。

参数

deepseek-v4.1-flash$0.24 输入 · $0.96 输出,每百万 token (当前为高峰时段)

思考
Reasoning effort

开启思考时,temperature 不生效,低于 0.95 的 top_p 按 0.95 执行。

Log probabilities

参数会被接受,但不会返回对数概率。

对话总计$0.000000
就绪

试试下面这几条提示词

每条消息会从你的余额中扣费,费率如下所示,和调用 API 一样。失败的请求不收费。

DeepSeek V4.1 Flash 的真实输出

一个提示词,在关闭思考的情况下发送给 DeepSeek V4.1 Flash。未经编辑。

Explain the difference between a mutex and a semaphore in two short paragraphs, and say when to use each.

A mutex is a locking mechanism designed for mutual exclusion: only the thread that locked it may unlock it, and it has a notion of ownership. It's binary (locked/unlocked) and is used to protect a shared resource so that only one thread accesses it at a time. A semaphore is a signaling mechanism built around a counter: it tracks how many permits are available, any thread can wait (decrement) or post (increment) it, and it doesn't care which thread does what. A binary semaphore looks similar to a mutex, but lacks ownership, so it can be released by a different thread—useful for signaling between threads rather than guarding a critical section. Use a mutex when you need to protect a shared resource (a data structure, file, or device) from concurrent access, which is the common case for locking critical sections. Use a semaphore when you need to control access to a pool of N identical resources (e.g., limiting concurrent database connections) or to signal between threads, such as a producer notifying a consumer that data is ready or a task-completion event. As a rule of thumb: mutex for mutual exclusion, semaphore for counting or signaling.

生成于 2026-09-28。每次运行的回答都可能不同。

DeepSeek V4.1 Flash 价格

DeepSeek V4.1 Flash 按 token 计费,从你的余额中扣除,按请求所在时段适用高峰或非高峰费率。下方费率为实时价格,与你的账户实际扣费费率相同。

一次典型的请求2000 个输入 token、1000 个输出 token,高峰时段
$0.0014每次请求
在非高峰时段,同样的请求只需一半费用。
$20.00 能买多少按高峰时段费率计算输出 token
20,833,333输出 token 数
非高峰时段可买到两倍的量。缓存输入的价格只是输入的一小部分。

示例用量仅供说明,按高峰时段费率计算。实际账单以每个请求报告的 token 数及其运行时段为准。

项目计为价格计费方式
输入,高峰时段未命中缓存的提示 token,UTC 周一至周五 01:00–04:00 和 06:00–10:00$0.24 / 1M每百万 token
输出,高峰时段高峰时段的回答 token 与推理 token$0.96 / 1M每百万 token
缓存输入,高峰时段高峰时段命中缓存的提示 token$0.0048 / 1M每百万 token
输入,非高峰时段未命中缓存的提示 token,其余所有时段,含周末$0.12 / 1M每百万 token
输出,非高峰时段非高峰时段的回答 token 与推理 token$0.48 / 1M每百万 token
缓存输入,非高峰时段非高峰时段命中缓存的提示 token$0.0024 / 1M每百万 token
失败的请求任何返回错误的请求免费不收费

DeepSeek V4.1 Flash 是什么?

DeepSeek V4.1 Flash 是 DeepSeek 推出的快速、低价模型,于 2026 年 9 月 10 日发布。它是一个 5520 亿参数的混合专家(MoE)模型,读取时激活 80 亿参数、生成时激活 160 亿参数,原生支持图片输入,最多可接收 100 万 token 的上下文。

DeepSeek V4.1 Flash 默认先思考再作答,推理强度为 high。简单快速的步骤可以关闭思考,难题则把推理强度调到 max。在 DeepSeek 自己的 API 中,这个模型叫 deepseek-flash。

模型 ID:deepseek-v4.1-flash · 输入:文本、图片 · 输出:文本 · 接口格式:Chat Completions、Responses、Anthropic Messages。

模型 ID
deepseek-v4.1-flash
上下文窗口
100 万 token
最大输出
384K token
推理强度等级
none、low、high、max
输入
文本、图片

使用 DeepSeek V4.1 Flash 的两种方式

从你的代码调用它,或让你的编码助手来做。

API

从后端调用

适合应用和服务

把 OpenAI SDK 指向 SeedRouter,代码无需改动。DeepSeek V4.1 Flash 支持 Chat Completions、Responses 和 Anthropic Messages 三种格式。

  1. 1创建 API Key
  2. 2把 base URL 设为 https://api.seedrouter.ai/v1
  3. 3将 model 设为 deepseek-v4.1-flash
代理

交给你的编码助手

适合 Codex、Claude Code 和其他代理

复制一个现成的提示词。你的助手会在执行任何操作前显示请求和费用。

  1. 1复制提示词
  2. 2粘贴到你的助手
  3. 3批准请求

DeepSeek V4.1 Flash 擅长什么

DeepSeek V4.1 Flash 专为看重速度和费用的大批量工作打造。

代理循环

DeepSeek V4.1 Flash 能以低廉的价格运行大量步骤,缓存的上下文让重复读取保持低价。

日常编码

在 Codex、Claude Code 或你自己的工具中,用 DeepSeek V4.1 Flash 编写、审查和修复代码。

提示中带图片

DeepSeek V4.1 Flash 能读懂通过 URL 或 base64 传入的截图、图表和照片。

结构化输出

让 DeepSeek V4.1 Flash 输出 json_object,直接解析回答。

思考可开可关

简单问题跳过 DeepSeek V4.1 Flash 的推理快速作答,难题则调高推理强度。

非高峰时段费率

在高峰时段之外运行的批量任务,只需支付高峰价格的一半。

团队用 DeepSeek V4.1 Flash 构建什么

大规模分类

关闭思考,用 DeepSeek V4.1 Flash 给工单、评论和日志打标签,费用最低。

编码代理

在 Codex 中运行 DeepSeek V4.1 Flash,实现快速、低价的“编辑—运行”循环。

截图识别

把截图和图表转换成文本或 JSON。

四步开始使用 DeepSeek V4.1 Flash

  1. 01

    创建 Key

    登录并创建 API Key。随时添加余额;没有订阅。

  2. 02

    更改基础 URL

    代码保持不变,只需把 OpenAI SDK 的 base_url 指向 https://api.seedrouter.ai/v1。

  3. 03

    选择 DeepSeek V4.1 Flash

    将 model 设为 deepseek-v4.1-flash。可按请求关闭思考或设置推理强度。

  4. 04

    检查账单

    每个请求在你的使用记录中显示其 token 数和费用。

获取 API Key

SeedRouter 上的 DeepSeek V4.1 Flash 一览

你现在可以使用的功能。

功能DeepSeek V4.1 Flash
Chat Completions支持,官方格式
Responses支持,可用于 Codex
Anthropic Messages支持,可用于 Claude Code
图片输入是,支持 URL 或 base64
流式输出是
上下文缓存支持,自动生效,命中缓存按更低费率计费
JSON 输出支持,json_object
失败的请求不收费

不提供 FIM 与对话前缀续写(beta)、Files API 和对数概率。

DeepSeek V4.1 Flash 的使用限制

百万 token 上下文

在 DeepSeek V4.1 Flash 上,提示、图片和历史记录共用一个 100 万 token 的窗口。

384K token 输出

单次 DeepSeek V4.1 Flash 请求最多输出 393,216 token,包含推理内容。

开启思考时的采样参数

开启思考时,temperature 不生效,top_p 保持在 0.95 或以上。

图片大小

DeepSeek V4.1 Flash 的图片 URL 所指向的文件最大可达 32 MiB。

为什么通过 SeedRouter 调用 DeepSeek V4.1 Flash

一把 key,三种格式

用同一个 Key,即可通过 Chat Completions、Responses 或 Anthropic Messages 调用 DeepSeek V4.1 Flash。

按 token 付费

一次充值,任意模型都能用。无套餐,无月费。

失败请求不收费

返回错误的请求不会被扣费。

实时价格

本页显示的高峰与非高峰时段费率,就是你的账户支付的费率。

上下文缓存

重复的提示前缀会从缓存读取,按更低的 DeepSeek V4.1 Flash 费率计费。

回复中返回推理内容

开启思考时,推理内容会在回答旁边的 reasoning_content 中返回。

流式输出

设置 stream: true,即可边生成边接收 DeepSeek V4.1 Flash 的 token。

相关模型

使用同一把 key 可以调用的其他模型。

Kimi K3
kimi-k3

用 OpenAI 或 Anthropic SDK 调用 Kimi K3:100 万 token 上下文,推理始终开启,强度可选 low、high 或 max,按 token 实时计价,Playground 在线试用。

查看价格
claude-fable-5
claude-fable-5

Anthropic

查看价格
claude-fable-5-1
claude-fable-5-1

Anthropic

查看价格
claude-opus-5-5
claude-opus-5-5

Anthropic

查看价格
dreamina-seedance-2-0
dreamina-seedance-2-0

ByteDance

查看价格
dreamina-seedance-2-0-fast
dreamina-seedance-2-0-fast

ByteDance

查看价格

DeepSeek V4.1 Flash 使用指南

该模型的教程和对比。

DeepSeek V4.1 Flash 常见问题

DeepSeek V4.1 Flash 是什么?+

DeepSeek V4.1 Flash 是 DeepSeek 推出的快速、低价模型,于 2026 年 9 月 10 日发布。它是一个 5520 亿参数的混合专家(MoE)模型,原生支持图片输入,拥有 100 万 token 的上下文窗口,思考可开可关。

DeepSeek V4.1 Flash 多少钱?+

DeepSeek V4.1 Flash 在高峰时段(UTC 周一至周五 01:00–04:00 和 06:00–10:00)按每百万 token $0.24 / 1M 输入、$0.96 / 1M 输出计费。其余所有时段为非高峰时段,费率减半;命中缓存的价格只是输入的一小部分。

DeepSeek V4.1 Flash 可以免费用吗?+

每个新注册的 SeedRouter 账户都自带 $0.10 免费余额,足够发送许多条简短请求来试用 DeepSeek V4.1 Flash。之后按 token 付费,没有订阅,失败的请求不收费。

DeepSeek V4.1 Flash 支持图片吗?+

支持。DeepSeek V4.1 Flash 原生支持读取图片,可在请求中通过公开 URL 或 base64 传入;它返回文本。

怎么关闭 DeepSeek V4.1 Flash 的思考?+

把 thinking 设为 disabled,或把 reasoning_effort 设为 none。这样回答会立即返回,并且消耗更少的输出 token。

DeepSeek V4.1 Flash 和 DeepSeek V4 Flash 是同一个模型吗?+

不是。DeepSeek V4.1 Flash 于 2026 年 9 月 10 日取代了 V4 Flash,采用新架构并原生支持图片输入。DeepSeek 已下线 V4 Flash,并把它的旧模型名称指向 DeepSeek V4.1 Flash。

DeepSeek V4.1 Flash 能本地部署吗?+

它的权重已发布在 Hugging Face 上。它拥有 5520 亿参数,需要多 GPU 硬件才能运行,所以大多数团队会通过 API 调用 DeepSeek V4.1 Flash。

DeepSeek V4.1 Flash 的模型 ID 是什么?怎么调用?+

在 SeedRouter 上,模型 ID 是 deepseek-v4.1-flash。创建一个 SeedRouter API Key,把 OpenAI SDK 指向 https://api.seedrouter.ai/v1,然后将 model 设为 deepseek-v4.1-flash 发送请求。

DeepSeek V4.1 Flash 请求失败会收费吗?+

不会。返回错误的 DeepSeek V4.1 Flash 请求不收费。你只需为成功完成的请求所报告的 token 付费。

立即试用 DeepSeek V4.1 Flash

几分钟内,就能在 Playground 或你自己的代码里发出第一个 DeepSeek V4.1 Flash 请求。