# DeepSeek V4.1 Flash API pricing: peak, off-peak and cache hits

By SeedRouter · Published 2026-09-28 · Updated 2026-09-28

DeepSeek V4.1 Flash is billed per token, and the rate depends on the hour. Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday; every other hour, weekends included, is off-peak at half the peak rates. Input that hits the cache costs a small fraction of fresh input, and there is no charge for writing to the cache. On SeedRouter there is no subscription, and a request that fails is not charged.

The table below is read from the rate table that bills your account, so it shows what a request costs today.

## What does DeepSeek V4.1 Flash cost per million tokens?

### DeepSeek V4.1 Flash

Current rates are temporarily unavailable. No price estimate is provided; unavailable rates must not be interpreted as free usage.

Off-peak covers 133 of the 168 hours in a week: all of Saturday and Sunday, and every weekday hour outside the two peak windows.

## What are DeepSeek's list prices?

For reference, this is the shape of DeepSeek's own prices for the model, from its [Models & Pricing page](https://api-docs.deepseek.com/quick_start/pricing). Each line is a multiple of the peak input rate:

| Line              | Peak  | Off-peak |
| ----------------- | ----- | -------- |
| Input, cache miss | 1×    | 0.5×     |
| Input, cache hit  | 0.02× | 0.01×    |
| Output            | 4×    | 2×       |

The live table above shows SeedRouter's rate for every line, and the [DeepSeek V4.1 Flash page](https://seedrouter.ai/models/deepseek-v4-1-flash#pricing) shows them next to DeepSeek's list price. DeepSeek's own table also marks Chinese public holidays as off-peak; SeedRouter applies the weekday schedule every week.

## How is a DeepSeek V4.1 Flash request billed?

Every response reports its token counts in `usage`, and the charge follows them:

* `prompt_cache_miss_tokens`: prompt tokens read fresh.
* `prompt_cache_hit_tokens`: prompt tokens served from the cache.
* `completion_tokens`: the answer plus the reasoning, when thinking is on.

The charge, at the rates of the hour the request runs:

```text
cost = ( cache-miss input × input rate
       + cache-hit input × cache-hit rate
       + completion × output rate ) / 1,000,000
```

## How much does one DeepSeek V4.1 Flash task cost?

Thinking changes the output side most. In our test on the same prime-counting question, thinking off used 2 output tokens, `low` effort 258 and `max` 319. Three illustrative shapes of work, in tokens:

| Task                                  | Input                            | Output      |
| ------------------------------------- | -------------------------------- | ----------- |
| A classification with thinking off    | about 150                        | about 10    |
| A code review at `high` effort        | about 5,000                      | about 1,500 |
| An agent step over a large repository | about 200,000, mostly cache hits | about 3,000 |

Multiply each by the live rates above, and halve it off-peak. Output costs four times input per token, so reasoning and answers usually decide the bill. For an agent that re-reads the same files, cache hits cut the input side to about 2% of fresh input.

## Did the DeepSeek V4 Flash price change?

Twice in 2026. On August 16, DeepSeek moved its V4 models to peak and off-peak pricing, with off-peak at half the peak rate. On September 10, when V4.1 Flash replaced V4 Flash, its [change log](https://api-docs.deepseek.com/updates) says "API prices have been reduced accordingly". The most recent change was a cut, and the old model names now run on V4.1 Flash at the V4.1 Flash price.

## How do I pay less for DeepSeek V4.1 Flash?

* **Run flexible work off-peak.** Batch jobs, evaluations and backfills cost half outside the peak windows.
* **Turn thinking off for simple steps.** Classification and extraction rarely need reasoning, and reasoning is output.
* **Keep reused context at the start of the prompt.** Caching is automatic and needs no setup; repeated prefixes are billed at the cache-hit rate.
* **Set `max_tokens`** so a long answer cannot run past what you meant to pay for.

## Frequently asked questions

### When are DeepSeek V4.1 Flash peak hours?

01:00–04:00 and 06:00–10:00 UTC, Monday to Friday. Everything else is off-peak.

### Is there a charge for writing to the cache?

No. DeepSeek V4.1 Flash bills cache misses and cache hits; writing to the cache is not billed separately.

### Is there a DeepSeek V4.1 Flash subscription?

Not on SeedRouter. You top up a balance and pay per token, with no plan and no monthly fee.

### How do I get started?

The [DeepSeek V4.1 Flash API guide](https://seedrouter.ai/blog/deepseek-v4-1-flash-api) shows how to get a key and make the first call.
