Claude Opus 5.5 is live on SeedRouter
DeepSeekText generation

DeepSeek V4.1 Flash — fast, low-cost reasoning with images

DeepSeek V4.1 Flash is DeepSeek's fast, low-cost model with a 1M-token context and image input. $0.24 / 1M input, $0.96 / 1M output per 1M tokens at peak hours; half that off-peak.

Available$0.24 / 1M input · $0.96 / 1M output$0.30 / 1M input · $1.20 / 1M output-20%
  • Thinking on or off
  • 1M-token context
  • Image input
  • Off-peak pricing
  • Streaming
Context window
1M tokens
Max output
384K tokens
Effort levels
none, low, high, max
Released
Sep 10, 2026

Try DeepSeek V4.1 Flash

Chat with DeepSeek V4.1 Flash, attach images, and see the tokens and cost of every answer.

Parameters

deepseek-v4.1-flash$0.24 in · $0.96 out per 1M tokens (peak now)

Thinking
Reasoning effort

With thinking on, temperature has no effect and top_p below 0.95 runs as 0.95.

Log probabilities

Accepted, but no log probabilities are returned.

Conversation total$0.000000
Ready

Try one of these prompts

Messages are billed from your balance at the rates below, like an API call. Failed requests are free.

Real DeepSeek V4.1 Flash output

One prompt, sent to DeepSeek V4.1 Flash with thinking off. Unedited.

Explain the difference between a mutex and a semaphore in two short paragraphs, and say when to use each.

A mutex is a locking mechanism designed for mutual exclusion: only the thread that locked it may unlock it, and it has a notion of ownership. It's binary (locked/unlocked) and is used to protect a shared resource so that only one thread accesses it at a time. A semaphore is a signaling mechanism built around a counter: it tracks how many permits are available, any thread can wait (decrement) or post (increment) it, and it doesn't care which thread does what. A binary semaphore looks similar to a mutex, but lacks ownership, so it can be released by a different thread—useful for signaling between threads rather than guarding a critical section. Use a mutex when you need to protect a shared resource (a data structure, file, or device) from concurrent access, which is the common case for locking critical sections. Use a semaphore when you need to control access to a pool of N identical resources (e.g., limiting concurrent database connections) or to signal between threads, such as a producer notifying a consumer that data is ready or a task-completion event. As a rule of thumb: mutex for mutual exclusion, semaphore for counting or signaling.

Generated 2026-09-28. Answers vary from run to run.

DeepSeek V4.1 Flash pricing

DeepSeek V4.1 Flash is billed per token from your balance, at peak or off-peak rates by the hour. Rates below are live, the same ones your account is charged.

One typical request2,000 input tokens, 1,000 output tokens, peak
$0.0014per request
Off-peak, the same request costs half.
What $20.00 buysOutput tokens at the peak rate
20,833,333output tokens
Twice as many off-peak. Cached input costs a fraction of input.

Example sizes are for illustration and use peak rates. Your bill follows the tokens each request reports and the hour it runs.

ItemCounted asRateBilling
Input, peakPrompt tokens that miss the cache, 01:00–04:00 and 06:00–10:00 UTC, Mon–Fri$0.24 / 1MPer 1M tokens
Output, peakAnswer and reasoning tokens, peak hours$0.96 / 1MPer 1M tokens
Cached input, peakPrompt tokens that hit the cache, peak hours$0.0048 / 1MPer 1M tokens
Input, off-peakPrompt tokens that miss the cache, all other hours, weekends included$0.12 / 1MPer 1M tokens
Output, off-peakAnswer and reasoning tokens, off-peak hours$0.48 / 1MPer 1M tokens
Cached input, off-peakPrompt tokens that hit the cache, off-peak hours$0.0024 / 1MPer 1M tokens
Failed requestAny request that returns an errorFreeNot charged

What is DeepSeek V4.1 Flash?

DeepSeek V4.1 Flash is DeepSeek's fast, low-cost model, released on September 10, 2026. It is a 552B-parameter mixture-of-experts model that activates 8B parameters to read and 16B to write, reads images natively and takes up to 1M tokens of context.

DeepSeek V4.1 Flash thinks before it answers by default, at high effort. Turn thinking off for fast, simple steps, or raise the effort to max for hard problems. DeepSeek's own API calls the model deepseek-flash.

Model ID: deepseek-v4.1-flash · Input: text, images · Output: text · Formats: Chat Completions, Responses, Anthropic Messages.

Model ID
deepseek-v4.1-flash
Context window
1M tokens
Max output
384K tokens
Effort levels
none, low, high, max
Input
Text, images

Two ways to use DeepSeek V4.1 Flash

Call it from your code, or let your coding agent do it for you.

API

Call it from your backend

Best for apps and services

Point the OpenAI SDK at SeedRouter and keep your code. DeepSeek V4.1 Flash answers Chat Completions, Responses and Anthropic Messages.

  1. 1Create an API key
  2. 2Set the base URL to https://api.seedrouter.ai/v1
  3. 3Send deepseek-v4.1-flash as the model
Agent

Hand it to your coding agent

Best for Codex, Claude Code and other agents

Copy a ready prompt. Your agent shows the request and cost before anything is spent.

  1. 1Copy the prompt
  2. 2Paste it into your agent
  3. 3Approve the request

What DeepSeek V4.1 Flash is good at

DeepSeek V4.1 Flash is built for high-volume work where speed and cost matter.

Agent loops

DeepSeek V4.1 Flash runs many cheap steps, and cached context keeps repeat reads low-cost.

Everyday coding

Write, review and fix code with DeepSeek V4.1 Flash in Codex, Claude Code or your own tools.

Images in the prompt

DeepSeek V4.1 Flash reads screenshots, charts and photos by URL or base64.

Structured output

Ask DeepSeek V4.1 Flash for json_object output and parse the answer directly.

Thinking on or off

Skip the reasoning of DeepSeek V4.1 Flash for quick answers, or turn it up for hard ones.

Off-peak rates

Batch jobs outside peak hours pay half the peak price.

What teams build with DeepSeek V4.1 Flash

Classification at scale

Tag tickets, reviews and logs with DeepSeek V4.1 Flash, thinking off, for the lowest cost.

Coding agents

Run DeepSeek V4.1 Flash in Codex for quick, cheap edit-and-run loops.

Screenshot reading

Turn screenshots and charts into text or JSON.

Start using DeepSeek V4.1 Flash in four steps

  1. 01

    Create a key

    Sign in and create an API key. Add credit whenever you need it; there is no subscription.

  2. 02

    Change the base URL

    Point the OpenAI SDK at https://api.seedrouter.ai/v1. Your code stays the same.

  3. 03

    Pick DeepSeek V4.1 Flash

    Send deepseek-v4.1-flash as the model. Turn thinking off or set its effort per request.

  4. 04

    Check the bill

    Every request shows its tokens and charge in your usage records.

Get an API key

DeepSeek V4.1 Flash on SeedRouter at a glance

What you can use today.

FeatureDeepSeek V4.1 Flash
Chat CompletionsYes, official format
ResponsesYes, works with Codex
Anthropic MessagesYes, works with Claude Code
Image inputYes, by URL or base64
StreamingYes
Context cachingYes, automatic, cache hits billed lower
JSON outputYes, json_object
Failed requestsNot charged

FIM and chat prefix completion (beta), the Files API and log probabilities are not offered.

DeepSeek V4.1 Flash limits

1M-token context

On DeepSeek V4.1 Flash, prompt, images and history share one 1M-token window.

384K tokens out

One DeepSeek V4.1 Flash request writes up to 393,216 tokens, reasoning included.

Sampling with thinking on

With thinking on, temperature has no effect and top_p stays at 0.95 or above.

Image size

A DeepSeek V4.1 Flash image URL may point to a file of up to 32 MiB.

Why call DeepSeek V4.1 Flash through SeedRouter

One key, three formats

Call DeepSeek V4.1 Flash with Chat Completions, Responses or Anthropic Messages, using the same key.

Pay per token

Top up once and spend it on any model. No plan, no monthly fee.

Failures are free

A request that errors is not charged.

Live prices

The rates on this page are the rates your account pays, peak and off-peak.

Context caching

Repeated prompt prefixes are read from the cache at a lower DeepSeek V4.1 Flash rate.

Reasoning back in the reply

With thinking on, the reasoning comes back in reasoning_content next to the answer.

Streaming

Stream DeepSeek V4.1 Flash tokens as they arrive with stream: true.

Related models

Other models you can call with the same key.

Kimi K3
kimi-k3

Call Kimi K3 with the OpenAI or Anthropic SDK. 1M-token context, always-on reasoning with low, high or max effort, live per-token pricing and a playground.

View pricing
claude-fable-5
claude-fable-5

Anthropic

View pricing
claude-fable-5-1
claude-fable-5-1

Anthropic

View pricing
claude-opus-5-5
claude-opus-5-5

Anthropic

View pricing
dreamina-seedance-2-0
dreamina-seedance-2-0

ByteDance

View pricing
dreamina-seedance-2-0-fast
dreamina-seedance-2-0-fast

ByteDance

View pricing

DeepSeek V4.1 Flash FAQ

What is DeepSeek V4.1 Flash?+

DeepSeek V4.1 Flash is DeepSeek's fast, low-cost model, released on September 10, 2026. It is a 552B-parameter mixture-of-experts model with native image input, a 1M-token context window and thinking you can turn on or off.

How much does DeepSeek V4.1 Flash cost?+

DeepSeek V4.1 Flash is billed at $0.24 / 1M input and $0.96 / 1M output per 1M tokens during peak hours (01:00–04:00 and 06:00–10:00 UTC, Monday to Friday). Every other hour is off-peak at half those rates, and cache hits cost a fraction of input.

Can I use DeepSeek V4.1 Flash for free?+

Every new SeedRouter account starts with $0.10 of free balance, enough to try DeepSeek V4.1 Flash with many short requests. After that you pay per token, with no subscription, and a failed request is not charged.

Does DeepSeek V4.1 Flash support images?+

Yes. DeepSeek V4.1 Flash reads images natively. Send them by public URL or as base64 in the request; it returns text.

How do I turn off thinking in DeepSeek V4.1 Flash?+

Set thinking to disabled, or reasoning_effort to none. The answer then comes straight away and uses fewer output tokens.

Is DeepSeek V4.1 Flash the same as DeepSeek V4 Flash?+

No. DeepSeek V4.1 Flash replaced V4 Flash on September 10, 2026, with a new architecture and native image input. DeepSeek retired V4 Flash and routes its old model names to DeepSeek V4.1 Flash.

Can I run DeepSeek V4.1 Flash locally?+

Its weights are on Hugging Face. At 552B parameters it needs multi-GPU hardware, so most teams call DeepSeek V4.1 Flash through an API.

What is the DeepSeek V4.1 Flash model ID and how do I call it?+

On SeedRouter the model ID is deepseek-v4.1-flash. Create a SeedRouter API key, point the OpenAI SDK at https://api.seedrouter.ai/v1 and send deepseek-v4.1-flash as the model.

Am I charged if a DeepSeek V4.1 Flash request fails?+

No. A DeepSeek V4.1 Flash request that returns an error is not charged. You pay only for the tokens a finished request reports.

Try DeepSeek V4.1 Flash now

Send your first DeepSeek V4.1 Flash request from the playground or your own code in minutes.