DeepSeek V4.1 Flash — fast, low-cost reasoning with images
DeepSeek V4.1 Flash is DeepSeek's fast, low-cost model with a 1M-token context and image input. $0.24 / 1M input, $0.96 / 1M output per 1M tokens at peak hours; half that off-peak.
- Thinking on or off
- 1M-token context
- Image input
- Off-peak pricing
- Streaming
Try DeepSeek V4.1 Flash
Chat with DeepSeek V4.1 Flash, attach images, and see the tokens and cost of every answer.
deepseek-v4.1-flash$0.24 in · $0.96 out per 1M tokens (peak now)
With thinking on, temperature has no effect and top_p below 0.95 runs as 0.95.
Accepted, but no log probabilities are returned.
Try one of these prompts
Messages are billed from your balance at the rates below, like an API call. Failed requests are free.
Real DeepSeek V4.1 Flash output
One prompt, sent to DeepSeek V4.1 Flash with thinking off. Unedited.
Explain the difference between a mutex and a semaphore in two short paragraphs, and say when to use each.
A mutex is a locking mechanism designed for mutual exclusion: only the thread that locked it may unlock it, and it has a notion of ownership. It's binary (locked/unlocked) and is used to protect a shared resource so that only one thread accesses it at a time. A semaphore is a signaling mechanism built around a counter: it tracks how many permits are available, any thread can wait (decrement) or post (increment) it, and it doesn't care which thread does what. A binary semaphore looks similar to a mutex, but lacks ownership, so it can be released by a different thread—useful for signaling between threads rather than guarding a critical section. Use a mutex when you need to protect a shared resource (a data structure, file, or device) from concurrent access, which is the common case for locking critical sections. Use a semaphore when you need to control access to a pool of N identical resources (e.g., limiting concurrent database connections) or to signal between threads, such as a producer notifying a consumer that data is ready or a task-completion event. As a rule of thumb: mutex for mutual exclusion, semaphore for counting or signaling.
Generated 2026-09-28. Answers vary from run to run.
DeepSeek V4.1 Flash pricing
DeepSeek V4.1 Flash is billed per token from your balance, at peak or off-peak rates by the hour. Rates below are live, the same ones your account is charged.
Example sizes are for illustration and use peak rates. Your bill follows the tokens each request reports and the hour it runs.
| Item | Counted as | Rate | Billing |
|---|---|---|---|
| Input, peak | Prompt tokens that miss the cache, 01:00–04:00 and 06:00–10:00 UTC, Mon–Fri | $0.24 / 1M | Per 1M tokens |
| Output, peak | Answer and reasoning tokens, peak hours | $0.96 / 1M | Per 1M tokens |
| Cached input, peak | Prompt tokens that hit the cache, peak hours | $0.0048 / 1M | Per 1M tokens |
| Input, off-peak | Prompt tokens that miss the cache, all other hours, weekends included | $0.12 / 1M | Per 1M tokens |
| Output, off-peak | Answer and reasoning tokens, off-peak hours | $0.48 / 1M | Per 1M tokens |
| Cached input, off-peak | Prompt tokens that hit the cache, off-peak hours | $0.0024 / 1M | Per 1M tokens |
| Failed request | Any request that returns an error | Free | Not charged |
What is DeepSeek V4.1 Flash?
DeepSeek V4.1 Flash is DeepSeek's fast, low-cost model, released on September 10, 2026. It is a 552B-parameter mixture-of-experts model that activates 8B parameters to read and 16B to write, reads images natively and takes up to 1M tokens of context.
DeepSeek V4.1 Flash thinks before it answers by default, at high effort. Turn thinking off for fast, simple steps, or raise the effort to max for hard problems. DeepSeek's own API calls the model deepseek-flash.
Model ID: deepseek-v4.1-flash · Input: text, images · Output: text · Formats: Chat Completions, Responses, Anthropic Messages.
Two ways to use DeepSeek V4.1 Flash
Call it from your code, or let your coding agent do it for you.
Call it from your backend
Point the OpenAI SDK at SeedRouter and keep your code. DeepSeek V4.1 Flash answers Chat Completions, Responses and Anthropic Messages.
- 1Create an API key
- 2Set the base URL to https://api.seedrouter.ai/v1
- 3Send deepseek-v4.1-flash as the model
Hand it to your coding agent
Copy a ready prompt. Your agent shows the request and cost before anything is spent.
- 1Copy the prompt
- 2Paste it into your agent
- 3Approve the request
What DeepSeek V4.1 Flash is good at
DeepSeek V4.1 Flash is built for high-volume work where speed and cost matter.
Agent loops
DeepSeek V4.1 Flash runs many cheap steps, and cached context keeps repeat reads low-cost.
Everyday coding
Write, review and fix code with DeepSeek V4.1 Flash in Codex, Claude Code or your own tools.
Images in the prompt
DeepSeek V4.1 Flash reads screenshots, charts and photos by URL or base64.
Structured output
Ask DeepSeek V4.1 Flash for json_object output and parse the answer directly.
Thinking on or off
Skip the reasoning of DeepSeek V4.1 Flash for quick answers, or turn it up for hard ones.
Off-peak rates
Batch jobs outside peak hours pay half the peak price.
What teams build with DeepSeek V4.1 Flash
Classification at scale
Tag tickets, reviews and logs with DeepSeek V4.1 Flash, thinking off, for the lowest cost.
Coding agents
Run DeepSeek V4.1 Flash in Codex for quick, cheap edit-and-run loops.
Screenshot reading
Turn screenshots and charts into text or JSON.
Start using DeepSeek V4.1 Flash in four steps
- 01
Create a key
Sign in and create an API key. Add credit whenever you need it; there is no subscription.
- 02
Change the base URL
Point the OpenAI SDK at https://api.seedrouter.ai/v1. Your code stays the same.
- 03
Pick DeepSeek V4.1 Flash
Send deepseek-v4.1-flash as the model. Turn thinking off or set its effort per request.
- 04
Check the bill
Every request shows its tokens and charge in your usage records.
DeepSeek V4.1 Flash on SeedRouter at a glance
What you can use today.
| Feature | DeepSeek V4.1 Flash |
|---|---|
| Chat Completions | Yes, official format |
| Responses | Yes, works with Codex |
| Anthropic Messages | Yes, works with Claude Code |
| Image input | Yes, by URL or base64 |
| Streaming | Yes |
| Context caching | Yes, automatic, cache hits billed lower |
| JSON output | Yes, json_object |
| Failed requests | Not charged |
FIM and chat prefix completion (beta), the Files API and log probabilities are not offered.
DeepSeek V4.1 Flash limits
1M-token context
On DeepSeek V4.1 Flash, prompt, images and history share one 1M-token window.
384K tokens out
One DeepSeek V4.1 Flash request writes up to 393,216 tokens, reasoning included.
Sampling with thinking on
With thinking on, temperature has no effect and top_p stays at 0.95 or above.
Image size
A DeepSeek V4.1 Flash image URL may point to a file of up to 32 MiB.
Why call DeepSeek V4.1 Flash through SeedRouter
One key, three formats
Call DeepSeek V4.1 Flash with Chat Completions, Responses or Anthropic Messages, using the same key.
Pay per token
Top up once and spend it on any model. No plan, no monthly fee.
Failures are free
A request that errors is not charged.
Live prices
The rates on this page are the rates your account pays, peak and off-peak.
Context caching
Repeated prompt prefixes are read from the cache at a lower DeepSeek V4.1 Flash rate.
Reasoning back in the reply
With thinking on, the reasoning comes back in reasoning_content next to the answer.
Streaming
Stream DeepSeek V4.1 Flash tokens as they arrive with stream: true.
Related models
Other models you can call with the same key.
Call Kimi K3 with the OpenAI or Anthropic SDK. 1M-token context, always-on reasoning with low, high or max effort, live per-token pricing and a playground.
View pricingDeepSeek V4.1 Flash guides
Walkthroughs and comparisons for this model.
How to use the DeepSeek V4.1 Flash API: get a key, call it with the OpenAI SDK, turn thinking on or off, stream, send images and fix the errors new users hit.
DeepSeek V4.1 Flash API pricing per million tokens: peak and off-peak rates, cache hits, the hours each applies, the billing formula and what a request costs.
How to run DeepSeek V4.1 Flash as the model in Claude Code and OpenAI Codex CLI through SeedRouter: environment variables, config.toml and our test results.
DeepSeek V4.1 Flash vs DeepSeek V4 Pro compared on price per token, image input, concurrency and DeepSeek-reported benchmarks, with a clear pick for each job.
What changed from DeepSeek V4 Flash 0731 to V4.1 Flash: a new architecture, native vision, a smaller cache, lower prices, benchmarks, and where old names go.
Is DeepSeek V4.1 Flash free? The weights are free to download, the API is paid per token at low rates, and here is how to try it for next to nothing today.
DeepSeek V4.1 Flash FAQ
What is DeepSeek V4.1 Flash?+
DeepSeek V4.1 Flash is DeepSeek's fast, low-cost model, released on September 10, 2026. It is a 552B-parameter mixture-of-experts model with native image input, a 1M-token context window and thinking you can turn on or off.
How much does DeepSeek V4.1 Flash cost?+
DeepSeek V4.1 Flash is billed at $0.24 / 1M input and $0.96 / 1M output per 1M tokens during peak hours (01:00–04:00 and 06:00–10:00 UTC, Monday to Friday). Every other hour is off-peak at half those rates, and cache hits cost a fraction of input.
Can I use DeepSeek V4.1 Flash for free?+
Every new SeedRouter account starts with $0.10 of free balance, enough to try DeepSeek V4.1 Flash with many short requests. After that you pay per token, with no subscription, and a failed request is not charged.
Does DeepSeek V4.1 Flash support images?+
Yes. DeepSeek V4.1 Flash reads images natively. Send them by public URL or as base64 in the request; it returns text.
How do I turn off thinking in DeepSeek V4.1 Flash?+
Set thinking to disabled, or reasoning_effort to none. The answer then comes straight away and uses fewer output tokens.
Is DeepSeek V4.1 Flash the same as DeepSeek V4 Flash?+
No. DeepSeek V4.1 Flash replaced V4 Flash on September 10, 2026, with a new architecture and native image input. DeepSeek retired V4 Flash and routes its old model names to DeepSeek V4.1 Flash.
Can I run DeepSeek V4.1 Flash locally?+
Its weights are on Hugging Face. At 552B parameters it needs multi-GPU hardware, so most teams call DeepSeek V4.1 Flash through an API.
What is the DeepSeek V4.1 Flash model ID and how do I call it?+
On SeedRouter the model ID is deepseek-v4.1-flash. Create a SeedRouter API key, point the OpenAI SDK at https://api.seedrouter.ai/v1 and send deepseek-v4.1-flash as the model.
Am I charged if a DeepSeek V4.1 Flash request fails?+
No. A DeepSeek V4.1 Flash request that returns an error is not charged. You pay only for the tokens a finished request reports.
Try DeepSeek V4.1 Flash now
Send your first DeepSeek V4.1 Flash request from the playground or your own code in minutes.
