Kimi K3 — open frontier model for long coding and agent work
Kimi K3 is Moonshot AI's flagship model for long-horizon coding and knowledge work, with a 1M-token context. $2.40 / 1M input, $12.00 / 1M output per 1M tokens.
- Always-on reasoning
- 1M-token context
- Image input
- Prompt caching
- Streaming
Try Kimi K3
Chat with Kimi K3, attach images, and see the tokens and cost of every answer.
kimi-k3$2.40 in · $12.00 out per 1M tokens
Accepted, but no log probabilities are returned.
Try one of these prompts
Messages are billed from your balance at the rates below, like an API call. Failed requests are free.
Real Kimi K3 output
One prompt, sent to Kimi K3 at low reasoning effort. Unedited.
Explain the difference between a mutex and a semaphore in two short paragraphs, and say when to use each.
A mutex (mutual exclusion) is a locking mechanism with ownership: the thread that locks it must be the one to unlock it, and only one thread can hold it at a time. It exists to protect shared resources from concurrent access. A semaphore is a counter with no ownership: one thread can signal (increment) it while another waits (decrements) it. A binary semaphore acts like a mutex, but counting semaphores allow N threads to access a pool of resources simultaneously. When to use each: Use a mutex when you need exclusive access to shared data—like protecting a critical section or data structure from race conditions. Use a semaphore when you're coordinating or signaling between threads—like limiting access to a fixed pool of connections, implementing producer-consumer queues, or notifying one thread that work is available.
Generated 2026-09-28. Answers vary from run to run.
Kimi K3 pricing
Kimi K3 is billed per token from your balance. Rates below are live, the same ones your account is charged.
Example sizes are for illustration. Your bill follows the tokens each request reports.
| Item | Counted as | Rate | Billing |
|---|---|---|---|
| Input | Prompt tokens not read from or written to the cache | $2.40 / 1M | Per 1M tokens |
| Output | Answer and reasoning tokens | $12.00 / 1M | Per 1M tokens |
| Cached input | Prompt tokens served from the cache | $0.24 / 1M | Per 1M tokens |
| Cache write, 5 min | Prompt tokens written to the cache with the default 5-minute TTL | $2.40 / 1M | Per 1M tokens |
| Cache write, 1 hour | Prompt tokens written to the cache with a 1-hour TTL | $4.80 / 1M | Per 1M tokens |
| Failed request | Any request that returns an error | Free | Not charged |
What is Kimi K3?
Kimi K3 is Moonshot AI's flagship model for long-horizon coding, agents and knowledge work. It is an open-weight model with 2.8 trillion parameters, and Kimi K3 reads up to 1M tokens of context in one request.
Kimi K3 always reasons before it answers. You choose how hard with reasoning effort: low for quick steps, high for most work and max, the default, for the hardest problems.
Model ID: kimi-k3 · Input: text, images · Output: text · Formats: Chat Completions, Responses, Anthropic Messages.
Two ways to use Kimi K3
Call it from your code, or let your coding agent do it for you.
Call it from your backend
Point the OpenAI SDK at SeedRouter and keep your code. Kimi K3 answers Chat Completions, Responses and Anthropic Messages.
- 1Create an API key
- 2Set the base URL to https://api.seedrouter.ai/v1
- 3Send kimi-k3 as the model
Hand it to your coding agent
Copy a ready prompt. Your agent shows the request and cost before anything is spent.
- 1Copy the prompt
- 2Paste it into your agent
- 3Approve the request
What Kimi K3 is good at
Kimi K3 is built for long tasks that need a lot of context and many steps.
Long coding sessions
Kimi K3 keeps a large codebase in view and works through multi-step changes.
Agent workflows
Tool calls and a 1M-token context let an agent run long without losing the thread.
Knowledge work
Hand Kimi K3 reports, contracts or logs and get a structured read on them.
Images in the prompt
Kimi K3 reads screenshots and charts alongside your text.
Effort you control
Low for quick answers, max for hard problems. Kimi K3 reasons on every request.
Cheaper repeat context
Cached input on Kimi K3 costs a tenth of the input rate.
What teams build with Kimi K3
Coding agents
Run Kimi K3 in Codex or your own agent loop for multi-file changes.
Code review bots
Check every pull request with the whole repository in context.
Structured extraction
Turn tickets and documents into JSON with a json_schema response format.
Start using Kimi K3 in four steps
- 01
Create a key
Sign in and create an API key. Add credit whenever you need it; there is no subscription.
- 02
Change the base URL
Point the OpenAI SDK at https://api.seedrouter.ai/v1. Your code stays the same.
- 03
Pick Kimi K3
Send kimi-k3 as the model. Set reasoning effort to trade depth for speed and cost.
- 04
Check the bill
Every request shows its tokens and charge in your usage records.
Kimi K3 on SeedRouter at a glance
What you can use today.
| Feature | Kimi K3 |
|---|---|
| Chat Completions | Yes, official format |
| Responses | Yes, works with Codex |
| Anthropic Messages | Yes |
| Image input | Yes, as base64 |
| Streaming | Yes |
| Prompt caching | Yes, 5-minute or 1-hour TTL |
| Structured output | Yes, json_object and json_schema |
| Failed requests | Not charged |
Video input, file uploads, web search, batch jobs and log probabilities are not offered.
Kimi K3 limits
1M-token context
On Kimi K3, prompt, images and history share one 1,048,576-token window.
Reasoning uses the output budget
Kimi K3 writes up to 131,072 tokens by default, reasoning included; raise max_completion_tokens for more.
Fixed sampling
Kimi K3 fixes temperature at 1.0 and top_p at 0.95. Any other value returns an error.
Images as base64
Kimi K3 takes images as data URIs, not public URLs.
Why call Kimi K3 through SeedRouter
One key, three formats
Call Kimi K3 with Chat Completions, Responses or Anthropic Messages, using the same key.
Pay per token
Top up once and spend it on any model. No plan, no monthly fee.
Failures are free
A request that errors is not charged.
Live prices
The rates on this page are the rates your account pays.
Prompt caching
Repeated prompt prefixes are read from the cache at a lower Kimi K3 rate.
Reasoning back in the reply
Kimi K3 returns its reasoning in reasoning_content next to the answer.
Streaming
Stream Kimi K3 tokens as they arrive with stream: true.
Related models
Other models you can call with the same key.
Call DeepSeek V4.1 Flash with the OpenAI or Anthropic SDK. 1M-token context, thinking on or off, image input, live peak and off-peak pricing and a playground.
View pricingKimi K3 guides
Walkthroughs and comparisons for this model.
Is Kimi K3 free to use? The weights are free to download, but running them needs data-center hardware, and the API is paid per token. What you can try free.
How to access the Kimi K3 API: get a key, call kimi-k3 with the OpenAI SDK, set reasoning effort, stream, send images and fix the errors new users hit first.
Kimi K3 API pricing per million tokens: input, output, cached input and 5-minute or 1-hour cache writes, with the billing formula and what a task costs.
How to run Kimi K3 as the model in Claude Code and OpenAI Codex CLI through SeedRouter: the environment variables, the config.toml and test results.
Kimi K3 vs Claude Fable 5 and Claude Opus 5.5 compared on list price, context, output limits, reasoning control, open weights and API formats, with a verdict.
Kimi K3 FAQ
What is Kimi K3?+
Kimi K3 is Moonshot AI's flagship model for long-horizon coding, agents and knowledge work. It is an open-weight model with 2.8 trillion parameters, a 1M-token context window and reasoning on every request.
What is Kimi K3 good for?+
Kimi K3 suits long coding sessions, agents that call tools over many steps, and work on large documents or codebases, where the 1M-token context keeps everything in view.
How much does Kimi K3 cost?+
Kimi K3 is billed at $2.40 / 1M input and $12.00 / 1M output per 1M tokens. Cached input costs a tenth of the input rate, and cache writes are billed by their TTL. The price does not change with prompt length.
Can I use Kimi K3 for free?+
Every new SeedRouter account starts with $0.10 of free balance, enough to try Kimi K3 with a few requests. After that you pay per token from a prepaid balance, with no subscription, and a request that fails is not charged.
Is Kimi K3 open source, and can I run it locally?+
Kimi K3's weights are published by Moonshot AI. With 2.8 trillion parameters it needs data-center hardware to run, so most teams call Kimi K3 through an API instead.
Who makes Kimi K3?+
Kimi K3 is made by Moonshot AI, the Chinese company behind the Kimi assistant.
Can I turn off Kimi K3's reasoning?+
No. Kimi K3 always reasons. You can set reasoning effort to low, high or max; the default is max, and lower effort answers faster with fewer output tokens.
Can Kimi K3 read images?+
Yes. Kimi K3 accepts images alongside text as base64 data URIs. It does not take public image URLs, and it returns text only.
What is the Kimi K3 model ID and how do I call it?+
The model ID is kimi-k3. Create a SeedRouter API key, point the OpenAI SDK at https://api.seedrouter.ai/v1 and send kimi-k3 as the model with Chat Completions or Responses.
Am I charged if a Kimi K3 request fails?+
No. A Kimi K3 request that returns an error is not charged. You pay only for the tokens a finished request reports.
Try Kimi K3 now
Send your first Kimi K3 request from the playground or your own code in minutes.
