Claude Opus 5.5 is live on SeedRouter
SeedRouter Docs

DeepSeek V4.1 Flash

Call DeepSeek V4.1 Flash with the official Chat Completions, Responses or Anthropic Messages API: a 1M-token context, thinking on or off and images.

View Markdown

DeepSeek V4.1 Flash is DeepSeek's fast, low-cost model (DeepSeek's own API calls it deepseek-flash). It thinks before it answers by default, and you can turn thinking off or set its effort per request. Send the official DeepSeek request to SeedRouter: change the base URL and the API key, keep the body.

Model ID

Model IDContext windowMax outputReasoning effortDefault
deepseek-v4.1-flash1M tokens384K tokens (393,216)none, low, high, maxThinking on, high

Input: text and images. Output: text. See the model page for current prices.

Quick example

curl https://api.seedrouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $SEEDROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-v4.1-flash",
    "messages": [{"role": "user", "content": "Give me three names for a coffee shop."}]
  }'

Endpoints

FormatMethod and pathAuthentication
Chat CompletionsPOST https://api.seedrouter.ai/v1/chat/completionsAuthorization: Bearer <key>
ResponsesPOST https://api.seedrouter.ai/v1/responsesAuthorization: Bearer <key>
Anthropic MessagesPOST https://api.seedrouter.ai/v1/messagesx-api-key: <key> or Authorization: Bearer <key>, plus anthropic-version

All three return DeepSeek's official response format, streaming or not. Keep the API key in server-side code.

Parameters

Chat Completions fields:

NameTypeRequiredDefaultNotes
modelstringYes—deepseek-v4.1-flash.
messagesobject[]Yes—Text messages; images as image_url parts (see Image input).
thinking.typeenumNoenabledenabled or disabled.
reasoning_effortenumNohighnone (thinking off), low, high or max. minimal runs as low, medium and xhigh as high.
max_tokensintegerNo8K, or 64K with thinking (128K at max effort)1–393216. Includes the reasoning.
stopstring or string[]No—Stop sequences.
response_formatobjectNo{"type": "text"}text or json_object. json_schema returns 400.
toolsobject[]No—Function tools; strict is accepted.
tool_choicestring or objectNonone without tools, auto with toolsauto and none are applied. required and a named function are accepted but do not force a call.
streambooleanNofalseStream server-sent events.
stream_options.include_usagebooleanNofalseEvery chunk carries usage, null except on the last.
temperaturenumberNo10–2. No effect in thinking mode.
top_pnumberNo10–1. In thinking mode values below 0.95 run as 0.95; without thinking it stays 1.
user_idstringNo—Your end-user identifier.
logprobs, top_logprobs—No—Accepted (top_logprobs 0–20), but no log probabilities are returned.
frequency_penalty, presence_penalty—No—Deprecated by DeepSeek: accepted, no effect.

Thinking and effort

Thinking is on by default at high effort. Turn it off with "thinking": {"type": "disabled"} or "reasoning_effort": "none"; the answer then comes straight away and costs fewer output tokens. max spends the most reasoning on hard problems. The reasoning comes back in reasoning_content, next to content, and is billed as output tokens.

When a request carries tools, send every earlier assistant message back with its reasoning_content, as DeepSeek requires in tool-call conversations.

Image input

Images go in a user message's content as image_url parts, either a public http(s) URL or a base64 data URI:

{"role": "user", "content": [
  {"type": "image_url", "image_url": {"url": "https://example.com/chart.png"}},
  {"type": "text", "text": "What does this chart show?"}
]}

A URL may be at most 8192 characters and point to an image of at most 32 MiB. Replace the example URL with a publicly reachable image of your own.

Billing dimensions

See the current rates on the model page. A request is billed by the tokens it uses:

  • input tokens that miss the cache (prompt_cache_miss_tokens),
  • input tokens that hit the cache (prompt_cache_hit_tokens),
  • output tokens, including reasoning.

Rates depend on when the request runs. Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday; every other hour, weekends included, is off-peak at half the peak rates. The charge is taken from the usage reported with the finished response. A request that fails is not charged. Your account's usage records show the exact charge for every request.

Output

A non-streaming Chat Completions request returns:

{
  "id": "bc86988e-...",
  "object": "chat.completion",
  "created": 1790585983,
  "model": "deepseek-v4.1-flash",
  "choices": [{
    "index": 0,
    "finish_reason": "stop",
    "logprobs": null,
    "message": {"role": "assistant", "reasoning_content": "...", "content": "..."}
  }],
  "usage": {
    "prompt_tokens": 36,
    "completion_tokens": 39,
    "total_tokens": 75,
    "prompt_cache_hit_tokens": 0,
    "prompt_cache_miss_tokens": 36,
    "prompt_tokens_details": {"cached_tokens": 0},
    "completion_tokens_details": {"reasoning_tokens": 0}
  }
}

With "stream": true each chunk carries a delta with reasoning_content or content, and the last chunk before data: [DONE] carries the usage.

Responses API and Codex

POST /v1/responses takes the Responses body: input, instructions, max_output_tokens, reasoning.effort (as reasoning_effort above), text.format (text or json_object; json_schema is accepted but not enforced), tools (function and the apply_patch custom tool), tool_choice, temperature, top_p, top_logprobs, user and stream. The reasoning comes back as a reasoning item with reasoning_text content, and a stream carries numbered events from response.created to response.completed, with the reasoning in response.reasoning_text.delta events. The API is stateless: previous_response_id, conversation and built-in tools such as web_search are ignored, so send the whole conversation in input.

To use DeepSeek V4.1 Flash in Codex, add a provider to ~/.codex/config.toml and set SEEDROUTER_API_KEY:

model = "deepseek-v4.1-flash"
model_provider = "seedrouter"
show_raw_agent_reasoning = true

[model_providers.seedrouter]
name = "SeedRouter"
base_url = "https://api.seedrouter.ai/v1"
env_key = "SEEDROUTER_API_KEY"
wire_api = "responses"

Anthropic Messages format

Code written for the Anthropic Messages API can call DeepSeek V4.1 Flash too: send the Messages body to /v1/messages with "model": "deepseek-v4.1-flash". system, max_tokens, tools, tool_choice (auto, none), thinking (enabled, disabled) and temperature (0–2) are applied; output_config.effort and metadata.user_id are accepted; top_k, stop_sequences and tool_choice any have no effect. The reasoning comes back as thinking blocks. Images go in as base64 or url sources.

Errors

Errors use {"error": {"code": ..., "message": "..."}} (the Messages endpoint uses Anthropic's error shape). The code is a code from the common error catalog. Failed requests are not charged.

Tips

  • Turn thinking off for simple, fast steps such as classification or extraction; keep it on for reasoning, math and code.
  • Keep long, reused context at the start of the prompt: cached input is billed at a fraction of the input rate.
  • Run large batch jobs off-peak, when every rate is half.