DeepSeek V4.1 Flash API: get a key and make your first call
How to use the DeepSeek V4.1 Flash API: get a key, call it with the OpenAI SDK, turn thinking on or off, stream, send images and fix the errors new users hit.
Read as MarkdownTo call DeepSeek V4.1 Flash, you need an API key from a platform that serves it and a request with its model ID. On DeepSeek's own API the model is deepseek-flash, and the old names deepseek-v4-flash and deepseek-v4-flash-vision-exp are routed to it. On SeedRouter the model ID is deepseek-v4.1-flash, and one key calls it pay-as-you-go in the official request format: point the OpenAI SDK at https://api.seedrouter.ai/v1 and keep your code.
This guide uses SeedRouter; the request bodies are the same as DeepSeek's own API.
How do I get a DeepSeek V4.1 Flash API key?
- Sign in to SeedRouter and open API keys.
- Create a key and copy it; it is shown once.
- Add credit when you need it. New accounts start with a small free balance, and there is no subscription.
Keep the key in an environment variable such as SEEDROUTER_API_KEY, and only use it from server-side code.
How do I call DeepSeek V4.1 Flash from Python?
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["SEEDROUTER_API_KEY"],
base_url="https://api.seedrouter.ai/v1",
)
completion = client.chat.completions.create(
model="deepseek-v4.1-flash",
messages=[{"role": "user", "content": "Give me three names for a coffee shop."}],
)
print(completion.choices[0].message.content)Thinking is on by default, so the message also carries the model's reasoning in reasoning_content, next to the answer in content.
How do I call it from Node.js or cURL?
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.SEEDROUTER_API_KEY,
baseURL: "https://api.seedrouter.ai/v1",
});
const completion = await client.chat.completions.create({
model: "deepseek-v4.1-flash",
messages: [{ role: "user", content: "Give me three names for a coffee shop." }],
});
console.log(completion.choices[0].message.content);curl https://api.seedrouter.ai/v1/chat/completions \
-H "Authorization: Bearer $SEEDROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "deepseek-v4.1-flash", "messages": [{"role": "user", "content": "Give me three names for a coffee shop."}]}'The same key also answers the Responses API (/v1/responses) and the Anthropic Messages format (/v1/messages) for deepseek-v4.1-flash.
How do I turn thinking off or set its effort?
Thinking is on by default at high effort. Turn it off with thinking, or choose an effort with reasoning_effort. The OpenAI SDK passes thinking through extra_body. The first request turns thinking off for the fastest, cheapest answer; the second uses the hardest effort:
completion = client.chat.completions.create(
model="deepseek-v4.1-flash",
messages=[{"role": "user", "content": "Classify: 'my card was charged twice'"}],
extra_body={"thinking": {"type": "disabled"}},
)
completion = client.chat.completions.create(
model="deepseek-v4.1-flash",
messages=[{"role": "user", "content": "How many primes are there below 150?"}],
reasoning_effort="max",
)reasoning_effort | Effect |
|---|---|
none | Thinking off |
low | Short reasoning |
high (default) | Most tasks |
max | The hardest problems |
DeepSeek also accepts minimal (runs as low), medium and xhigh (run as high). In our test on a prime-counting question, thinking off used 2 output tokens, low 258 and max 319. Reasoning is billed as output.
How do I stream the answer?
Add stream=True. With thinking on, the reasoning arrives first in delta.reasoning_content, then the answer in delta.content, and the last chunk carries the token usage:
stream = client.chat.completions.create(
model="deepseek-v4.1-flash",
messages=[{"role": "user", "content": "Write a haiku about latency."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)Can I send images?
Yes. DeepSeek V4.1 Flash reads images natively. Send a public URL or a base64 data URI in an image_url part:
completion = client.chat.completions.create(
model="deepseek-v4.1-flash",
messages=[{
"role": "user",
"content": [
{"type": "image_url", "image_url": {"url": "https://example.com/chart.png"}},
{"type": "text", "text": "What does this chart show?"},
],
}],
)An image URL may be up to 8,192 characters long and point to a file of up to 32 MiB.
What errors should I expect?
| Error | Cause | Fix |
|---|---|---|
400 on response_format | json_schema is not supported | Use {"type": "json_object"} and describe the shape in the prompt |
400 on temperature or top_p | Above 2 or above 1 | Keep them in range; with thinking on they have little effect anyway |
| 400 on an image URL | The file could not be downloaded as an image | Check the URL is public and points to an image |
| 401 | Missing or wrong key | Check the Authorization header |
Errors return {"error": {"code": ..., "message": "..."}}, and a request that fails is not charged.
Frequently asked questions
Is the DeepSeek V4.1 Flash API OpenAI-compatible?
Yes. It takes the Chat Completions and Responses formats, so the OpenAI SDK works with only the base URL and model changed. It also takes the Anthropic Messages format.
Why is the model ID different from DeepSeek's?
DeepSeek names the model deepseek-flash on its own API. SeedRouter uses deepseek-v4.1-flash so the version is part of the name. The request body is otherwise the same.
How much does a DeepSeek V4.1 Flash request cost?
It is billed per token, at peak or off-peak rates by the hour. The DeepSeek V4.1 Flash pricing guide has the live rates and worked examples.
Where is the full parameter list?
The DeepSeek V4.1 Flash API reference lists every field, with the DeepSeek V4.1 Flash page for a playground and live prices.



