Kimi K3
Call Kimi K3 with the official Chat Completions, Responses or Anthropic Messages API: a 1M-token context window, always-on reasoning and effort you choose.
Kimi K3 is Moonshot AI's flagship model for long-horizon coding, agents and knowledge work. It always reasons before it answers, and you choose how hard with reasoning_effort. Send the official Kimi request to SeedRouter: change the base URL and the API key, keep the body.
Model ID
| Model ID | Context window | Max output | Reasoning effort | Default effort |
|---|---|---|---|---|
kimi-k3 | 1,048,576 tokens | 1,048,576 tokens (default 131,072) | low, high, max | max |
Input: text and images. Output: text. See the model page for current prices.
Quick example
curl https://api.seedrouter.ai/v1/chat/completions \
-H "Authorization: Bearer $SEEDROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "kimi-k3",
"messages": [{"role": "user", "content": "Explain context caching in one sentence."}]
}'Endpoints
| Format | Method and path | Authentication |
|---|---|---|
| Chat Completions | POST https://api.seedrouter.ai/v1/chat/completions | Authorization: Bearer <key> |
| Responses | POST https://api.seedrouter.ai/v1/responses | Authorization: Bearer <key> |
| Anthropic Messages | POST https://api.seedrouter.ai/v1/messages | x-api-key: <key> or Authorization: Bearer <key>, plus anthropic-version |
All three return Kimi's official response format, streaming or not. Keep the API key in server-side code.
Parameters
Chat Completions fields:
| Name | Type | Required | Default | Notes |
|---|---|---|---|---|
model | string | Yes | — | kimi-k3. |
messages | object[] | Yes | — | Text messages; images as image_url parts (see Image input). |
max_completion_tokens | integer | No | 131072 | Up to 1048576. Includes reasoning tokens. max_tokens is the deprecated name for the same limit. |
reasoning_effort | enum | No | max | low, high or max. Any other value returns 400. |
stop | string or string[] | No | — | Up to 5 sequences. |
response_format | object | No | {"type": "text"} | text, json_object or json_schema (with json_schema.name and json_schema.schema). |
tools | object[] | No | — | Function tools. |
tool_choice | string or object | No | auto | auto and none are applied. required and a named function are accepted but do not force a call. |
stream | boolean | No | false | Stream server-sent events. |
stream_options.include_usage | boolean | No | false | Adds the final usage chunk. |
prompt_cache_options | object | No | {"mode": "implicit", "ttl": "5m"} | mode: implicit. ttl: 5m or 1h. |
prompt_cache_key, safety_identifier, prediction | — | No | — | Accepted. |
logprobs, top_logprobs | — | No | — | Accepted (top_logprobs 0–20), but no log probabilities are returned. |
temperature, top_p, n, presence_penalty, frequency_penalty | — | No | 1.0, 0.95, 1, 0, 0 | Fixed. Any other value returns 400, so leave them out. |
Reasoning and effort
Kimi K3 always reasons; there is no way to turn it off. reasoning_effort sets how much: max (the default) for the hardest work, high for most tasks, low for fast, simple steps. The reasoning comes back in reasoning_content, next to content. Reasoning tokens are billed as output tokens and count toward max_completion_tokens.
In multi-turn conversations and tool calls, send each assistant message back unchanged, including its reasoning_content.
Image input
Kimi K3 takes images as base64 data URIs. A public image URL is not accepted and returns 400, as on Kimi's own API.
{"role": "user", "content": [
{"type": "image_url", "image_url": {"url": "data:image/png;base64,<BASE64_DATA>"}},
{"type": "text", "text": "Describe this image."}
]}Context caching
Caching is automatic: a repeated prompt prefix is read from the cache at the lower cached-input rate. prompt_cache_options.ttl chooses how long a written prefix stays cached, 5m (the default) or 1h; choose 1h when your requests are more than five minutes apart. usage.prompt_tokens_details.cached_tokens reports the tokens read from the cache and cache_write_tokens the cache writes billed for the request.
Billing dimensions
See the current rates on the model page. A request is billed by the tokens it uses:
- input tokens,
- cached input tokens (
cached_tokens), - cache-write tokens (
cache_write_tokens), - output tokens, including reasoning.
Prices do not change with context length. The charge is taken from the usage reported with the finished response. A request that fails is not charged. Your account's usage records show the exact charge for every request.
Output
A non-streaming Chat Completions request returns:
{
"id": "chatcmpl-...",
"object": "chat.completion",
"created": 1790585961,
"model": "kimi-k3",
"choices": [{
"index": 0,
"finish_reason": "stop",
"message": {"role": "assistant", "reasoning_content": "...", "content": "..."}
}],
"usage": {
"prompt_tokens": 90,
"completion_tokens": 57,
"total_tokens": 147,
"cached_tokens": 90,
"prompt_tokens_details": {"cached_tokens": 90, "cache_write_tokens": 0}
}
}With "stream": true each chunk carries a delta with reasoning_content or content. With stream_options.include_usage a last chunk with an empty choices array carries the usage before data: [DONE].
Responses API and Codex
POST /v1/responses takes the Responses body: input, instructions, max_output_tokens, reasoning.effort (low, high, max), text.format (json_schema), tools (function and the apply_patch custom tool), tool_choice, stream, prompt_cache_options, prompt_cache_key and safety_identifier. The reasoning comes back as a reasoning item with a summary_text part, and a stream carries numbered events from response.created to response.completed. The API is stateless: previous_response_id and conversation are ignored, so send the whole conversation in input. The web_search tool is ignored.
To use Kimi K3 in Codex, add a provider to ~/.codex/config.toml and set SEEDROUTER_API_KEY:
model = "kimi-k3"
model_provider = "seedrouter"
model_context_window = 1048576
[model_providers.seedrouter]
name = "SeedRouter"
base_url = "https://api.seedrouter.ai/v1"
env_key = "SEEDROUTER_API_KEY"
wire_api = "responses"Anthropic Messages format
Code written for the Anthropic Messages API can call Kimi K3 too: send the Messages body to /v1/messages with "model": "kimi-k3". system, max_tokens, tools, tool_choice (auto, none) and output_config.effort (low, high, max) are applied, and metadata.user_id and cache_control are accepted. stop_sequences (up to 5), tool_choice any and output_config.format are accepted but have no effect. The reasoning comes back as thinking blocks. Images go in as base64 sources.
Errors
Errors use {"error": {"code": ..., "message": "..."}} (the Messages endpoint uses Anthropic's error shape). The code is a code from the common error catalog. Failed requests are not charged.
Tips
- Start with
higheffort and move tomaxonly for the hardest problems;lowsuits quick, simple steps. - Set
max_completion_tokenshigh enough for the reasoning as well as the answer: it is one budget for both. - Keep long, reused context at the start of the prompt so later requests read it from the cache, and use the
1hTTL when requests are spread out.
