Claude Opus 5.5 is live on SeedRouter

DeepSeek V4.1 Flash API: get a key and make your first call

How to use the DeepSeek V4.1 Flash API: get a key, call it with the OpenAI SDK, turn thinking on or off, stream, send images and fix the errors new users hit.

Read as Markdown

To call DeepSeek V4.1 Flash, you need an API key from a platform that serves it and a request with its model ID. On DeepSeek's own API the model is deepseek-flash, and the old names deepseek-v4-flash and deepseek-v4-flash-vision-exp are routed to it. On SeedRouter the model ID is deepseek-v4.1-flash, and one key calls it pay-as-you-go in the official request format: point the OpenAI SDK at https://api.seedrouter.ai/v1 and keep your code.

This guide uses SeedRouter; the request bodies are the same as DeepSeek's own API.

How do I get a DeepSeek V4.1 Flash API key?

  1. Sign in to SeedRouter and open API keys.
  2. Create a key and copy it; it is shown once.
  3. Add credit when you need it. New accounts start with a small free balance, and there is no subscription.

Keep the key in an environment variable such as SEEDROUTER_API_KEY, and only use it from server-side code.

How do I call DeepSeek V4.1 Flash from Python?

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["SEEDROUTER_API_KEY"],
    base_url="https://api.seedrouter.ai/v1",
)

completion = client.chat.completions.create(
    model="deepseek-v4.1-flash",
    messages=[{"role": "user", "content": "Give me three names for a coffee shop."}],
)
print(completion.choices[0].message.content)

Thinking is on by default, so the message also carries the model's reasoning in reasoning_content, next to the answer in content.

How do I call it from Node.js or cURL?

import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.SEEDROUTER_API_KEY,
  baseURL: "https://api.seedrouter.ai/v1",
});

const completion = await client.chat.completions.create({
  model: "deepseek-v4.1-flash",
  messages: [{ role: "user", content: "Give me three names for a coffee shop." }],
});
console.log(completion.choices[0].message.content);
curl https://api.seedrouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $SEEDROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "deepseek-v4.1-flash", "messages": [{"role": "user", "content": "Give me three names for a coffee shop."}]}'

The same key also answers the Responses API (/v1/responses) and the Anthropic Messages format (/v1/messages) for deepseek-v4.1-flash.

How do I turn thinking off or set its effort?

Thinking is on by default at high effort. Turn it off with thinking, or choose an effort with reasoning_effort. The OpenAI SDK passes thinking through extra_body. The first request turns thinking off for the fastest, cheapest answer; the second uses the hardest effort:

completion = client.chat.completions.create(
    model="deepseek-v4.1-flash",
    messages=[{"role": "user", "content": "Classify: 'my card was charged twice'"}],
    extra_body={"thinking": {"type": "disabled"}},
)

completion = client.chat.completions.create(
    model="deepseek-v4.1-flash",
    messages=[{"role": "user", "content": "How many primes are there below 150?"}],
    reasoning_effort="max",
)
reasoning_effortEffect
noneThinking off
lowShort reasoning
high (default)Most tasks
maxThe hardest problems

DeepSeek also accepts minimal (runs as low), medium and xhigh (run as high). In our test on a prime-counting question, thinking off used 2 output tokens, low 258 and max 319. Reasoning is billed as output.

How do I stream the answer?

Add stream=True. With thinking on, the reasoning arrives first in delta.reasoning_content, then the answer in delta.content, and the last chunk carries the token usage:

stream = client.chat.completions.create(
    model="deepseek-v4.1-flash",
    messages=[{"role": "user", "content": "Write a haiku about latency."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Can I send images?

Yes. DeepSeek V4.1 Flash reads images natively. Send a public URL or a base64 data URI in an image_url part:

completion = client.chat.completions.create(
    model="deepseek-v4.1-flash",
    messages=[{
        "role": "user",
        "content": [
            {"type": "image_url", "image_url": {"url": "https://example.com/chart.png"}},
            {"type": "text", "text": "What does this chart show?"},
        ],
    }],
)

An image URL may be up to 8,192 characters long and point to a file of up to 32 MiB.

What errors should I expect?

ErrorCauseFix
400 on response_formatjson_schema is not supportedUse {"type": "json_object"} and describe the shape in the prompt
400 on temperature or top_pAbove 2 or above 1Keep them in range; with thinking on they have little effect anyway
400 on an image URLThe file could not be downloaded as an imageCheck the URL is public and points to an image
401Missing or wrong keyCheck the Authorization header

Errors return {"error": {"code": ..., "message": "..."}}, and a request that fails is not charged.

Frequently asked questions

Is the DeepSeek V4.1 Flash API OpenAI-compatible?

Yes. It takes the Chat Completions and Responses formats, so the OpenAI SDK works with only the base URL and model changed. It also takes the Anthropic Messages format.

Why is the model ID different from DeepSeek's?

DeepSeek names the model deepseek-flash on its own API. SeedRouter uses deepseek-v4.1-flash so the version is part of the name. The request body is otherwise the same.

How much does a DeepSeek V4.1 Flash request cost?

It is billed per token, at peak or off-peak rates by the hour. The DeepSeek V4.1 Flash pricing guide has the live rates and worked examples.

Where is the full parameter list?

The DeepSeek V4.1 Flash API reference lists every field, with the DeepSeek V4.1 Flash page for a playground and live prices.

Related guides