DeepSeek V4.1 Flash
DeepSeek V4.1 Flash API 串接文件:透過官方 Chat Completions、Responses 或 Anthropic Messages API 呼叫,支援 100 萬 token 上下文,思考可開可關,支援圖片輸入。
DeepSeek V4.1 Flash 是 DeepSeek 快速、低成本的模型(DeepSeek 自家的 API 稱它為 deepseek-flash)。它預設先思考再回答,你可以在每個請求中關閉思考或設定思考強度。傳送官方 DeepSeek 請求到 SeedRouter:更改基礎 URL 和 API 金鑰,保留請求主體。
模型 ID
| 模型 ID | 上下文視窗 | 最大輸出 | 推理強度 | 預設值 |
|---|---|---|---|---|
deepseek-v4.1-flash | 1M tokens | 384K tokens(393,216) | none、low、high、max | 思考開啟,high |
輸入:文字和圖片。輸出:文字。查看模型頁面了解目前價格。
快速範例
curl https://api.seedrouter.ai/v1/chat/completions \
-H "Authorization: Bearer $SEEDROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v4.1-flash",
"messages": [{"role": "user", "content": "Give me three names for a coffee shop."}]
}'端點
| 格式 | 方法和路徑 | 驗證方式 |
|---|---|---|
| Chat Completions | POST https://api.seedrouter.ai/v1/chat/completions | Authorization: Bearer <key> |
| Responses | POST https://api.seedrouter.ai/v1/responses | Authorization: Bearer <key> |
| Anthropic Messages | POST https://api.seedrouter.ai/v1/messages | x-api-key: <key> 或 Authorization: Bearer <key>,再加上 anthropic-version |
三個端點都回傳 DeepSeek 的官方回應格式,串流與非串流皆可。請把 API 金鑰保存在伺服器端程式碼中。
參數
Chat Completions 欄位:
| 名稱 | 型別 | 必填 | 預設值 | 說明 |
|---|---|---|---|---|
model | string | 是 | — | deepseek-v4.1-flash。 |
messages | object[] | 是 | — | 文字訊息;圖片以 image_url 片段傳入(見圖片輸入)。 |
thinking.type | enum | 否 | enabled | enabled 或 disabled。 |
reasoning_effort | enum | 否 | high | none(關閉思考)、low、high 或 max。minimal 以 low 執行,medium 和 xhigh 以 high 執行。 |
max_tokens | integer | 否 | 8K;開啟思考時 64K(max 推理強度時 128K) | 1–393216。包括推理內容。 |
stop | string or string[] | 否 | — | 停止序列。 |
response_format | object | 否 | {"type": "text"} | text 或 json_object。json_schema 會回傳 400。 |
tools | object[] | 否 | — | 函式工具;接受 strict。 |
tool_choice | string or object | 否 | 無工具時為 none,有工具時為 auto | auto 與 none 會生效。required 和指定函式會被接受,但不會強制呼叫。 |
stream | boolean | 否 | false | 以伺服器傳送事件的形式串流。 |
stream_options.include_usage | boolean | 否 | false | 每個資料塊都帶有 usage,除最後一個外皆為 null。 |
temperature | number | 否 | 1 | 0–2。思考模式下沒有效果。 |
top_p | number | 否 | 1 | 0–1。思考模式下低於 0.95 的值以 0.95 執行;不開思考時固定為 1。 |
user_id | string | 否 | — | 你的終端使用者識別碼。 |
logprobs、top_logprobs | — | 否 | — | 接受(top_logprobs 為 0–20),但不會回傳對數機率。 |
frequency_penalty、presence_penalty | — | 否 | — | 已被 DeepSeek 棄用:接受,但沒有效果。 |
思考與思考強度
思考預設開啟,強度為 high。用 "thinking": {"type": "disabled"} 或 "reasoning_effort": "none" 關閉思考;答案會立即回傳,輸出 token 也更少。max 會在困難問題上投入最多推理。推理內容在 reasoning_content 中回傳,與 content 並列,按輸出 token 計費。
當請求帶有 tools 時,請把先前每則 assistant 訊息連同其 reasoning_content 一起傳回,這是 DeepSeek 對工具呼叫對話的要求。
圖片輸入
圖片放在使用者訊息的 content 中,以 image_url 片段傳入,可以是公開的 http(s) URL,也可以是 base64 data URI:
{"role": "user", "content": [
{"type": "image_url", "image_url": {"url": "https://example.com/chart.png"}},
{"type": "text", "text": "What does this chart show?"}
]}URL 最長 8192 個字元,指向的圖片最大 32 MiB。請把範例 URL 換成你自己可公開存取的圖片。
計費維度
查看模型頁面上的目前費率。請求依使用的 token 計費:
- 未命中快取的輸入 token(
prompt_cache_miss_tokens)、 - 命中快取的輸入 token(
prompt_cache_hit_tokens)、 - 輸出 token(包括推理)。
費率取決於請求執行的時間。尖峰時段為 UTC 週一至週五的 01:00–04:00 與 06:00–10:00;其他所有時間(包括週末)為離峰時段,費率是尖峰時段的一半。費用依完成回應所回報的 usage 計算。失敗的請求不計費。你的帳戶使用記錄會顯示每個請求的確切費用。
輸出
非串流 Chat Completions 請求回傳:
{
"id": "bc86988e-...",
"object": "chat.completion",
"created": 1790585983,
"model": "deepseek-v4.1-flash",
"choices": [{
"index": 0,
"finish_reason": "stop",
"logprobs": null,
"message": {"role": "assistant", "reasoning_content": "...", "content": "..."}
}],
"usage": {
"prompt_tokens": 36,
"completion_tokens": 39,
"total_tokens": 75,
"prompt_cache_hit_tokens": 0,
"prompt_cache_miss_tokens": 36,
"prompt_tokens_details": {"cached_tokens": 0},
"completion_tokens_details": {"reasoning_tokens": 0}
}
}使用 "stream": true 時,每個資料塊都帶有一個 delta,內容是 reasoning_content 或 content;data: [DONE] 之前的最後一個資料塊帶有用量資訊。
Responses API 與 Codex
POST /v1/responses 接受 Responses 請求主體:input、instructions、max_output_tokens、reasoning.effort(取值同上文的 reasoning_effort)、text.format(text 或 json_object;json_schema 會被接受但不會強制執行)、tools(function 與 apply_patch 自訂工具)、tool_choice、temperature、top_p、top_logprobs、user 與 stream。推理內容以帶有 reasoning_text 內容的 reasoning 項目回傳,串流則帶有從 response.created 到 response.completed 的編號事件,推理內容在 response.reasoning_text.delta 事件中。這個 API 是無狀態的:previous_response_id、conversation 以及 web_search 等內建工具會被忽略,因此請在 input 中傳送完整對話。
要在 Codex 中使用 DeepSeek V4.1 Flash,請在 ~/.codex/config.toml 中新增一個 provider,並設定 SEEDROUTER_API_KEY:
model = "deepseek-v4.1-flash"
model_provider = "seedrouter"
show_raw_agent_reasoning = true
[model_providers.seedrouter]
name = "SeedRouter"
base_url = "https://api.seedrouter.ai/v1"
env_key = "SEEDROUTER_API_KEY"
wire_api = "responses"Anthropic Messages 格式
為 Anthropic Messages API 撰寫的程式碼也能呼叫 DeepSeek V4.1 Flash:把 Messages 請求主體以 "model": "deepseek-v4.1-flash" 傳送到 /v1/messages。system、max_tokens、tools、tool_choice(auto、none)、thinking(enabled、disabled)與 temperature(0–2)會生效;output_config.effort 與 metadata.user_id 會被接受;top_k、stop_sequences 與 tool_choice any 沒有效果。推理內容以 thinking 區塊回傳。圖片以 base64 或 url 來源傳入。
錯誤
錯誤使用 {"error": {"code": ..., "message": "..."}} 格式(Messages 端點使用 Anthropic 的錯誤格式)。code 是來自共用錯誤目錄的代碼。失敗的請求不計費。
實用建議
- 分類、擷取等簡單快速的步驟可以關閉思考;推理、數學和程式碼任務請保持開啟。
- 將長的、重複使用的上下文放在提示開頭:快取輸入的費率只有輸入費率的一小部分。
- 大量批次工作請在離峰時段執行,屆時所有費率都減半。
