# DeepSeek V4.1 Flash

DeepSeek V4.1 Flash is DeepSeek's fast, low-cost model (DeepSeek's own API calls it `deepseek-flash`). It thinks before it answers by default, and you can turn thinking off or set its effort per request. Send the official DeepSeek request to SeedRouter: change the base URL and the API key, keep the body.

## Model ID

| Model ID              | Context window | Max output            | Reasoning effort             | Default             |
| --------------------- | -------------- | --------------------- | ---------------------------- | ------------------- |
| `deepseek-v4.1-flash` | 1M tokens      | 384K tokens (393,216) | `none`, `low`, `high`, `max` | Thinking on, `high` |

Input: text and images. Output: text. See the [model page](https://seedrouter.ai/models/deepseek-v4-1-flash#pricing) for current prices.

## Quick example

```bash
curl https://api.seedrouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $SEEDROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-v4.1-flash",
    "messages": [{"role": "user", "content": "Give me three names for a coffee shop."}]
  }'
```

```python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["SEEDROUTER_API_KEY"],
    base_url="https://api.seedrouter.ai/v1",
)

completion = client.chat.completions.create(
    model="deepseek-v4.1-flash",
    messages=[{"role": "user", "content": "Give me three names for a coffee shop."}],
)
print(completion.choices[0].message.content)
```

```javascript
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.SEEDROUTER_API_KEY,
  baseURL: "https://api.seedrouter.ai/v1",
});

const completion = await client.chat.completions.create({
  model: "deepseek-v4.1-flash",
  messages: [{ role: "user", content: "Give me three names for a coffee shop." }],
});
console.log(completion.choices[0].message.content);
```

```go
package main

import (
	"bytes"
	"fmt"
	"io"
	"net/http"
	"os"
)

func main() {
	body := []byte(`{"model": "deepseek-v4.1-flash",
		"messages": [{"role": "user", "content": "Give me three names for a coffee shop."}]}`)
	req, _ := http.NewRequest("POST", "https://api.seedrouter.ai/v1/chat/completions", bytes.NewReader(body))
	req.Header.Set("Authorization", "Bearer "+os.Getenv("SEEDROUTER_API_KEY"))
	req.Header.Set("Content-Type", "application/json")
	resp, err := http.DefaultClient.Do(req)
	if err != nil {
		panic(err)
	}
	defer resp.Body.Close()
	out, _ := io.ReadAll(resp.Body)
	fmt.Println(string(out))
}
```

## Endpoints

| Format             | Method and path                                      | Authentication                                                                |
| ------------------ | ---------------------------------------------------- | ----------------------------------------------------------------------------- |
| Chat Completions   | `POST https://api.seedrouter.ai/v1/chat/completions` | `Authorization: Bearer <key>`                                                 |
| Responses          | `POST https://api.seedrouter.ai/v1/responses`        | `Authorization: Bearer <key>`                                                 |
| Anthropic Messages | `POST https://api.seedrouter.ai/v1/messages`         | `x-api-key: <key>` or `Authorization: Bearer <key>`, plus `anthropic-version` |

All three return DeepSeek's official response format, streaming or not. Keep the API key in server-side code.

## Parameters

Chat Completions fields:

| Name                                    | Type                | Required | Default                                         | Notes                                                                                                   |
| --------------------------------------- | ------------------- | -------- | ----------------------------------------------- | ------------------------------------------------------------------------------------------------------- |
| `model`                                 | string              | Yes      | —                                               | `deepseek-v4.1-flash`.                                                                                  |
| `messages`                              | object\[]           | Yes      | —                                               | Text messages; images as `image_url` parts (see [Image input](#image-input)).                           |
| `thinking.type`                         | enum                | No       | `enabled`                                       | `enabled` or `disabled`.                                                                                |
| `reasoning_effort`                      | enum                | No       | `high`                                          | `none` (thinking off), `low`, `high` or `max`. `minimal` runs as `low`, `medium` and `xhigh` as `high`. |
| `max_tokens`                            | integer             | No       | 8K, or 64K with thinking (128K at `max` effort) | 1–393216. Includes the reasoning.                                                                       |
| `stop`                                  | string or string\[] | No       | —                                               | Stop sequences.                                                                                         |
| `response_format`                       | object              | No       | `{"type": "text"}`                              | `text` or `json_object`. `json_schema` returns 400.                                                     |
| `tools`                                 | object\[]           | No       | —                                               | Function tools; `strict` is accepted.                                                                   |
| `tool_choice`                           | string or object    | No       | `none` without tools, `auto` with tools         | `auto` and `none` are applied. `required` and a named function are accepted but do not force a call.    |
| `stream`                                | boolean             | No       | `false`                                         | Stream server-sent events.                                                                              |
| `stream_options.include_usage`          | boolean             | No       | `false`                                         | Every chunk carries `usage`, `null` except on the last.                                                 |
| `temperature`                           | number              | No       | 1                                               | 0–2. No effect in thinking mode.                                                                        |
| `top_p`                                 | number              | No       | 1                                               | 0–1. In thinking mode values below 0.95 run as 0.95; without thinking it stays 1.                       |
| `user_id`                               | string              | No       | —                                               | Your end-user identifier.                                                                               |
| `logprobs`, `top_logprobs`              | —                   | No       | —                                               | Accepted (`top_logprobs` 0–20), but no log probabilities are returned.                                  |
| `frequency_penalty`, `presence_penalty` | —                   | No       | —                                               | Deprecated by DeepSeek: accepted, no effect.                                                            |

## Thinking and effort

Thinking is on by default at `high` effort. Turn it off with `"thinking": {"type": "disabled"}` or `"reasoning_effort": "none"`; the answer then comes straight away and costs fewer output tokens. `max` spends the most reasoning on hard problems. The reasoning comes back in `reasoning_content`, next to `content`, and is billed as output tokens.

When a request carries `tools`, send every earlier assistant message back with its `reasoning_content`, as DeepSeek requires in tool-call conversations.

## Image input

Images go in a user message's `content` as `image_url` parts, either a public `http(s)` URL or a base64 data URI:

```json
{"role": "user", "content": [
  {"type": "image_url", "image_url": {"url": "https://example.com/chart.png"}},
  {"type": "text", "text": "What does this chart show?"}
]}
```

A URL may be at most 8192 characters and point to an image of at most 32 MiB. Replace the example URL with a publicly reachable image of your own.

## Billing dimensions

See the [current rates on the model page](https://seedrouter.ai/models/deepseek-v4-1-flash#pricing). A request is billed by the tokens it uses:

* input tokens that miss the cache (`prompt_cache_miss_tokens`),
* input tokens that hit the cache (`prompt_cache_hit_tokens`),
* output tokens, including reasoning.

Rates depend on when the request runs. Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday; every other hour, weekends included, is off-peak at half the peak rates. The charge is taken from the `usage` reported with the finished response. A request that fails is not charged. Your account's usage records show the exact charge for every request.

## Output

A non-streaming Chat Completions request returns:

```json
{
  "id": "bc86988e-...",
  "object": "chat.completion",
  "created": 1790585983,
  "model": "deepseek-v4.1-flash",
  "choices": [{
    "index": 0,
    "finish_reason": "stop",
    "logprobs": null,
    "message": {"role": "assistant", "reasoning_content": "...", "content": "..."}
  }],
  "usage": {
    "prompt_tokens": 36,
    "completion_tokens": 39,
    "total_tokens": 75,
    "prompt_cache_hit_tokens": 0,
    "prompt_cache_miss_tokens": 36,
    "prompt_tokens_details": {"cached_tokens": 0},
    "completion_tokens_details": {"reasoning_tokens": 0}
  }
}
```

With `"stream": true` each chunk carries a `delta` with `reasoning_content` or `content`, and the last chunk before `data: [DONE]` carries the usage.

## Responses API and Codex

`POST /v1/responses` takes the Responses body: `input`, `instructions`, `max_output_tokens`, `reasoning.effort` (as `reasoning_effort` above), `text.format` (`text` or `json_object`; `json_schema` is accepted but not enforced), `tools` (`function` and the `apply_patch` custom tool), `tool_choice`, `temperature`, `top_p`, `top_logprobs`, `user` and `stream`. The reasoning comes back as a `reasoning` item with `reasoning_text` content, and a stream carries numbered events from `response.created` to `response.completed`, with the reasoning in `response.reasoning_text.delta` events. The API is stateless: `previous_response_id`, `conversation` and built-in tools such as `web_search` are ignored, so send the whole conversation in `input`.

To use DeepSeek V4.1 Flash in Codex, add a provider to `~/.codex/config.toml` and set `SEEDROUTER_API_KEY`:

```toml
model = "deepseek-v4.1-flash"
model_provider = "seedrouter"
show_raw_agent_reasoning = true

[model_providers.seedrouter]
name = "SeedRouter"
base_url = "https://api.seedrouter.ai/v1"
env_key = "SEEDROUTER_API_KEY"
wire_api = "responses"
```

## Anthropic Messages format

Code written for the Anthropic Messages API can call DeepSeek V4.1 Flash too: send the Messages body to `/v1/messages` with `"model": "deepseek-v4.1-flash"`. `system`, `max_tokens`, `tools`, `tool_choice` (`auto`, `none`), `thinking` (`enabled`, `disabled`) and `temperature` (0–2) are applied; `output_config.effort` and `metadata.user_id` are accepted; `top_k`, `stop_sequences` and `tool_choice` `any` have no effect. The reasoning comes back as `thinking` blocks. Images go in as `base64` or `url` sources.

## Errors

Errors use `{"error": {"code": ..., "message": "..."}}` (the Messages endpoint uses Anthropic's error shape). The `code` is a code from the [common error catalog](https://seedrouter.ai/docs/api/errors). Failed requests are not charged.

## Tips

* Turn thinking off for simple, fast steps such as classification or extraction; keep it on for reasoning, math and code.
* Keep long, reused context at the start of the prompt: cached input is billed at a fraction of the input rate.
* Run large batch jobs off-peak, when every rate is half.

## Related

* [DeepSeek V4.1 Flash pricing](https://seedrouter.ai/models/deepseek-v4-1-flash)
* [Download OpenAPI](https://seedrouter.ai/docs/deepseek-v4-1-flash.openapi.json)
* [Copyable Markdown](https://seedrouter.ai/docs/deepseek-v4-1-flash.md)
