DeepSeek V4.1 Flash
Call DeepSeek V4.1 Flash with the official Chat Completions, Responses or Anthropic Messages API: a 1M-token context, thinking on or off and images.
DeepSeek V4.1 Flash is DeepSeek's fast, low-cost model (DeepSeek's own API calls it deepseek-flash). It thinks before it answers by default, and you can turn thinking off or set its effort per request. Send the official DeepSeek request to SeedRouter: change the base URL and the API key, keep the body.
Model ID
| Model ID | Context window | Max output | Reasoning effort | Default |
|---|---|---|---|---|
deepseek-v4.1-flash | 1M tokens | 384K tokens (393,216) | none, low, high, max | Thinking on, high |
Input: text and images. Output: text. See the model page for current prices.
Quick example
curl https://api.seedrouter.ai/v1/chat/completions \
-H "Authorization: Bearer $SEEDROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v4.1-flash",
"messages": [{"role": "user", "content": "Give me three names for a coffee shop."}]
}'Endpoints
| Format | Method and path | Authentication |
|---|---|---|
| Chat Completions | POST https://api.seedrouter.ai/v1/chat/completions | Authorization: Bearer <key> |
| Responses | POST https://api.seedrouter.ai/v1/responses | Authorization: Bearer <key> |
| Anthropic Messages | POST https://api.seedrouter.ai/v1/messages | x-api-key: <key> or Authorization: Bearer <key>, plus anthropic-version |
All three return DeepSeek's official response format, streaming or not. Keep the API key in server-side code.
Parameters
Chat Completions fields:
| Name | Type | Required | Default | Notes |
|---|---|---|---|---|
model | string | Yes | — | deepseek-v4.1-flash. |
messages | object[] | Yes | — | Text messages; images as image_url parts (see Image input). |
thinking.type | enum | No | enabled | enabled or disabled. |
reasoning_effort | enum | No | high | none (thinking off), low, high or max. minimal runs as low, medium and xhigh as high. |
max_tokens | integer | No | 8K, or 64K with thinking (128K at max effort) | 1–393216. Includes the reasoning. |
stop | string or string[] | No | — | Stop sequences. |
response_format | object | No | {"type": "text"} | text or json_object. json_schema returns 400. |
tools | object[] | No | — | Function tools; strict is accepted. |
tool_choice | string or object | No | none without tools, auto with tools | auto and none are applied. required and a named function are accepted but do not force a call. |
stream | boolean | No | false | Stream server-sent events. |
stream_options.include_usage | boolean | No | false | Every chunk carries usage, null except on the last. |
temperature | number | No | 1 | 0–2. No effect in thinking mode. |
top_p | number | No | 1 | 0–1. In thinking mode values below 0.95 run as 0.95; without thinking it stays 1. |
user_id | string | No | — | Your end-user identifier. |
logprobs, top_logprobs | — | No | — | Accepted (top_logprobs 0–20), but no log probabilities are returned. |
frequency_penalty, presence_penalty | — | No | — | Deprecated by DeepSeek: accepted, no effect. |
Thinking and effort
Thinking is on by default at high effort. Turn it off with "thinking": {"type": "disabled"} or "reasoning_effort": "none"; the answer then comes straight away and costs fewer output tokens. max spends the most reasoning on hard problems. The reasoning comes back in reasoning_content, next to content, and is billed as output tokens.
When a request carries tools, send every earlier assistant message back with its reasoning_content, as DeepSeek requires in tool-call conversations.
Image input
Images go in a user message's content as image_url parts, either a public http(s) URL or a base64 data URI:
{"role": "user", "content": [
{"type": "image_url", "image_url": {"url": "https://example.com/chart.png"}},
{"type": "text", "text": "What does this chart show?"}
]}A URL may be at most 8192 characters and point to an image of at most 32 MiB. Replace the example URL with a publicly reachable image of your own.
Billing dimensions
See the current rates on the model page. A request is billed by the tokens it uses:
- input tokens that miss the cache (
prompt_cache_miss_tokens), - input tokens that hit the cache (
prompt_cache_hit_tokens), - output tokens, including reasoning.
Rates depend on when the request runs. Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday; every other hour, weekends included, is off-peak at half the peak rates. The charge is taken from the usage reported with the finished response. A request that fails is not charged. Your account's usage records show the exact charge for every request.
Output
A non-streaming Chat Completions request returns:
{
"id": "bc86988e-...",
"object": "chat.completion",
"created": 1790585983,
"model": "deepseek-v4.1-flash",
"choices": [{
"index": 0,
"finish_reason": "stop",
"logprobs": null,
"message": {"role": "assistant", "reasoning_content": "...", "content": "..."}
}],
"usage": {
"prompt_tokens": 36,
"completion_tokens": 39,
"total_tokens": 75,
"prompt_cache_hit_tokens": 0,
"prompt_cache_miss_tokens": 36,
"prompt_tokens_details": {"cached_tokens": 0},
"completion_tokens_details": {"reasoning_tokens": 0}
}
}With "stream": true each chunk carries a delta with reasoning_content or content, and the last chunk before data: [DONE] carries the usage.
Responses API and Codex
POST /v1/responses takes the Responses body: input, instructions, max_output_tokens, reasoning.effort (as reasoning_effort above), text.format (text or json_object; json_schema is accepted but not enforced), tools (function and the apply_patch custom tool), tool_choice, temperature, top_p, top_logprobs, user and stream. The reasoning comes back as a reasoning item with reasoning_text content, and a stream carries numbered events from response.created to response.completed, with the reasoning in response.reasoning_text.delta events. The API is stateless: previous_response_id, conversation and built-in tools such as web_search are ignored, so send the whole conversation in input.
To use DeepSeek V4.1 Flash in Codex, add a provider to ~/.codex/config.toml and set SEEDROUTER_API_KEY:
model = "deepseek-v4.1-flash"
model_provider = "seedrouter"
show_raw_agent_reasoning = true
[model_providers.seedrouter]
name = "SeedRouter"
base_url = "https://api.seedrouter.ai/v1"
env_key = "SEEDROUTER_API_KEY"
wire_api = "responses"Anthropic Messages format
Code written for the Anthropic Messages API can call DeepSeek V4.1 Flash too: send the Messages body to /v1/messages with "model": "deepseek-v4.1-flash". system, max_tokens, tools, tool_choice (auto, none), thinking (enabled, disabled) and temperature (0–2) are applied; output_config.effort and metadata.user_id are accepted; top_k, stop_sequences and tool_choice any have no effect. The reasoning comes back as thinking blocks. Images go in as base64 or url sources.
Errors
Errors use {"error": {"code": ..., "message": "..."}} (the Messages endpoint uses Anthropic's error shape). The code is a code from the common error catalog. Failed requests are not charged.
Tips
- Turn thinking off for simple, fast steps such as classification or extraction; keep it on for reasoning, math and code.
- Keep long, reused context at the start of the prompt: cached input is billed at a fraction of the input rate.
- Run large batch jobs off-peak, when every rate is half.
