Claude Opus 5.5 is live on SeedRouter

Kimi K3 API pricing: cost per million tokens, cache and a task

Kimi K3 API pricing per million tokens: input, output, cached input and 5-minute or 1-hour cache writes, with the billing formula and what a task costs.

Read as Markdown

Kimi K3 is billed per token, with separate rates for input, output, cached input and cache writes. The price does not change with prompt length: a 900K-token prompt pays the same per-token rate as a 900-token one. Output costs the most, and Kimi K3 always reasons, so reasoning tokens are part of the output you pay for. On SeedRouter there is no subscription, and a request that fails is not charged.

The table below is read from the rate table that bills your account, so it shows what a request costs today.

What does Kimi K3 cost per million tokens?

Kimi K3

Current rates are temporarily unavailable. No price estimate is provided; unavailable rates must not be interpreted as free usage.

Cached input is what a repeated prompt prefix costs when it is read from the cache. The two cache-write columns are what writing a prefix to the cache costs, by how long it stays there.

What are Moonshot's list prices?

For reference, this is the shape of Moonshot AI's own prices for Kimi K3, from its pricing page. Each line is a multiple of the input rate:

LineMultiple of input
Input (cache miss)1×
Cached input (cache hit)0.1×
Cache write, 5-minute TTL1×
Cache write, 1-hour TTL2×
Output5×

The live table above shows SeedRouter's rate for every line, and the Kimi K3 page shows them next to Moonshot's list price.

How is a Kimi K3 request billed?

Every response reports its token counts in usage, and the charge follows them:

  • prompt_tokens: the whole prompt, including any cached or cache-written part.
  • prompt_tokens_details.cached_tokens: the part read from the cache.
  • prompt_tokens_details.cache_write_tokens: the part written to the cache.
  • completion_tokens: the answer plus the reasoning.

The charge is:

cost = ( (prompt - cached - cache writes) × input rate
       + cached × cached-input rate
       + cache writes × cache-write rate for the TTL
       + completion × output rate ) / 1,000,000

How much does one Kimi K3 task cost?

It depends on how many tokens the task reads and writes, and effort changes the output side a lot. In our test on the same short question, low effort used 25 output tokens and max used 146, almost six times the output for the same prompt. Three illustrative shapes of work, in tokens:

TaskInputOutput
A short question at low effortabout 100about 30
A code review of one file at high effortabout 5,000about 1,500
An agent step over a large repositoryabout 200,000, mostly cachedabout 3,000

Multiply each by the live rates above. Output weighs five times as much as input per token, so the reasoning and the answer usually decide the bill. For an agent that re-reads the same repository, cache hits cut the input side to a tenth.

How much do 1,000 Kimi K3 tokens cost?

A thousand tokens cost the per-million rate divided by 1,000. Moonshot counts roughly 3 to 4 English characters per token, so 1,000 tokens are a few thousand characters of text. At the cached-input rate they cost a tenth of fresh input.

Should I use the 5-minute or the 1-hour cache?

Caching is automatic. By default a written prefix stays for 5 minutes, and each hit refreshes it. Set prompt_cache_options.ttl to 1h when your requests are spread out:

completion = client.chat.completions.create(
    model="kimi-k3",
    messages=messages,
    prompt_cache_options={"mode": "implicit", "ttl": "1h"},
)

The 1-hour write costs twice the input rate, against once for the 5-minute write. If your requests come more than five minutes apart, a prefix read three or more times within the hour is cheaper on the 1-hour TTL: one write at 2× and two hits at 0.1× (2.2×) against three 5-minute writes (3×).

Is there a cheaper way to run Kimi K3?

Lower effort is the biggest lever, since output dominates the bill. Keep long, reused context at the start of the prompt so it is read from the cache. Kimi K3's weights are open, but at 2.8 trillion parameters running it yourself needs data-center hardware, which only pays off at very high volume.

Frequently asked questions

Does Kimi K3 charge more for long prompts?

No. Kimi K3 uses flat per-token pricing: there is no long-context tier.

Is there a Kimi K3 subscription?

Not on SeedRouter. You top up a balance and pay per token. Moonshot's own API also bills per token and unlocks Kimi K3 only after a first top-up.

Are failed requests charged?

No. A Kimi K3 request that returns an error is not charged.

How do I get started?

The Kimi K3 API guide shows how to get a key and make the first call.

Related guides