# Kimi K3

Kimi K3 is Moonshot AI's flagship model for long-horizon coding, agents and knowledge work. It always reasons before it answers, and you choose how hard with `reasoning_effort`. Send the official Kimi request to SeedRouter: change the base URL and the API key, keep the body.

## Model ID

| Model ID  | Context window   | Max output                         | Reasoning effort     | Default effort |
| --------- | ---------------- | ---------------------------------- | -------------------- | -------------- |
| `kimi-k3` | 1,048,576 tokens | 1,048,576 tokens (default 131,072) | `low`, `high`, `max` | `max`          |

Input: text and images. Output: text. See the [model page](https://seedrouter.ai/models/kimi-k3#pricing) for current prices.

## Quick example

```bash
curl https://api.seedrouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $SEEDROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "kimi-k3",
    "messages": [{"role": "user", "content": "Explain context caching in one sentence."}]
  }'
```

```python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["SEEDROUTER_API_KEY"],
    base_url="https://api.seedrouter.ai/v1",
)

completion = client.chat.completions.create(
    model="kimi-k3",
    messages=[{"role": "user", "content": "Explain context caching in one sentence."}],
)
print(completion.choices[0].message.content)
```

```javascript
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.SEEDROUTER_API_KEY,
  baseURL: "https://api.seedrouter.ai/v1",
});

const completion = await client.chat.completions.create({
  model: "kimi-k3",
  messages: [{ role: "user", content: "Explain context caching in one sentence." }],
});
console.log(completion.choices[0].message.content);
```

```go
package main

import (
	"bytes"
	"fmt"
	"io"
	"net/http"
	"os"
)

func main() {
	body := []byte(`{"model": "kimi-k3",
		"messages": [{"role": "user", "content": "Explain context caching in one sentence."}]}`)
	req, _ := http.NewRequest("POST", "https://api.seedrouter.ai/v1/chat/completions", bytes.NewReader(body))
	req.Header.Set("Authorization", "Bearer "+os.Getenv("SEEDROUTER_API_KEY"))
	req.Header.Set("Content-Type", "application/json")
	resp, err := http.DefaultClient.Do(req)
	if err != nil {
		panic(err)
	}
	defer resp.Body.Close()
	out, _ := io.ReadAll(resp.Body)
	fmt.Println(string(out))
}
```

## Endpoints

| Format             | Method and path                                      | Authentication                                                                |
| ------------------ | ---------------------------------------------------- | ----------------------------------------------------------------------------- |
| Chat Completions   | `POST https://api.seedrouter.ai/v1/chat/completions` | `Authorization: Bearer <key>`                                                 |
| Responses          | `POST https://api.seedrouter.ai/v1/responses`        | `Authorization: Bearer <key>`                                                 |
| Anthropic Messages | `POST https://api.seedrouter.ai/v1/messages`         | `x-api-key: <key>` or `Authorization: Bearer <key>`, plus `anthropic-version` |

All three return Kimi's official response format, streaming or not. Keep the API key in server-side code.

## Parameters

Chat Completions fields:

| Name                                                                 | Type                | Required | Default                             | Notes                                                                                                |
| -------------------------------------------------------------------- | ------------------- | -------- | ----------------------------------- | ---------------------------------------------------------------------------------------------------- |
| `model`                                                              | string              | Yes      | —                                   | `kimi-k3`.                                                                                           |
| `messages`                                                           | object\[]           | Yes      | —                                   | Text messages; images as `image_url` parts (see [Image input](#image-input)).                        |
| `max_completion_tokens`                                              | integer             | No       | 131072                              | Up to 1048576. Includes reasoning tokens. `max_tokens` is the deprecated name for the same limit.    |
| `reasoning_effort`                                                   | enum                | No       | `max`                               | `low`, `high` or `max`. Any other value returns 400.                                                 |
| `stop`                                                               | string or string\[] | No       | —                                   | Up to 5 sequences.                                                                                   |
| `response_format`                                                    | object              | No       | `{"type": "text"}`                  | `text`, `json_object` or `json_schema` (with `json_schema.name` and `json_schema.schema`).           |
| `tools`                                                              | object\[]           | No       | —                                   | Function tools.                                                                                      |
| `tool_choice`                                                        | string or object    | No       | `auto`                              | `auto` and `none` are applied. `required` and a named function are accepted but do not force a call. |
| `stream`                                                             | boolean             | No       | `false`                             | Stream server-sent events.                                                                           |
| `stream_options.include_usage`                                       | boolean             | No       | `false`                             | Adds the final usage chunk.                                                                          |
| `prompt_cache_options`                                               | object              | No       | `{"mode": "implicit", "ttl": "5m"}` | `mode`: `implicit`. `ttl`: `5m` or `1h`.                                                             |
| `prompt_cache_key`, `safety_identifier`, `prediction`                | —                   | No       | —                                   | Accepted.                                                                                            |
| `logprobs`, `top_logprobs`                                           | —                   | No       | —                                   | Accepted (`top_logprobs` 0–20), but no log probabilities are returned.                               |
| `temperature`, `top_p`, `n`, `presence_penalty`, `frequency_penalty` | —                   | No       | 1.0, 0.95, 1, 0, 0                  | Fixed. Any other value returns 400, so leave them out.                                               |

## Reasoning and effort

Kimi K3 always reasons; there is no way to turn it off. `reasoning_effort` sets how much: `max` (the default) for the hardest work, `high` for most tasks, `low` for fast, simple steps. The reasoning comes back in `reasoning_content`, next to `content`. Reasoning tokens are billed as output tokens and count toward `max_completion_tokens`.

In multi-turn conversations and tool calls, send each assistant message back unchanged, including its `reasoning_content`.

## Image input

Kimi K3 takes images as base64 data URIs. A public image URL is not accepted and returns 400, as on Kimi's own API.

```json
{"role": "user", "content": [
  {"type": "image_url", "image_url": {"url": "data:image/png;base64,<BASE64_DATA>"}},
  {"type": "text", "text": "Describe this image."}
]}
```

## Context caching

Caching is automatic: a repeated prompt prefix is read from the cache at the lower cached-input rate. `prompt_cache_options.ttl` chooses how long a written prefix stays cached, `5m` (the default) or `1h`; choose `1h` when your requests are more than five minutes apart. `usage.prompt_tokens_details.cached_tokens` reports the tokens read from the cache and `cache_write_tokens` the cache writes billed for the request.

## Billing dimensions

See the [current rates on the model page](https://seedrouter.ai/models/kimi-k3#pricing). A request is billed by the tokens it uses:

* input tokens,
* cached input tokens (`cached_tokens`),
* cache-write tokens (`cache_write_tokens`),
* output tokens, including reasoning.

Prices do not change with context length. The charge is taken from the `usage` reported with the finished response. A request that fails is not charged. Your account's usage records show the exact charge for every request.

## Output

A non-streaming Chat Completions request returns:

```json
{
  "id": "chatcmpl-...",
  "object": "chat.completion",
  "created": 1790585961,
  "model": "kimi-k3",
  "choices": [{
    "index": 0,
    "finish_reason": "stop",
    "message": {"role": "assistant", "reasoning_content": "...", "content": "..."}
  }],
  "usage": {
    "prompt_tokens": 90,
    "completion_tokens": 57,
    "total_tokens": 147,
    "cached_tokens": 90,
    "prompt_tokens_details": {"cached_tokens": 90, "cache_write_tokens": 0}
  }
}
```

With `"stream": true` each chunk carries a `delta` with `reasoning_content` or `content`. With `stream_options.include_usage` a last chunk with an empty `choices` array carries the usage before `data: [DONE]`.

## Responses API and Codex

`POST /v1/responses` takes the Responses body: `input`, `instructions`, `max_output_tokens`, `reasoning.effort` (`low`, `high`, `max`), `text.format` (`json_schema`), `tools` (`function` and the `apply_patch` custom tool), `tool_choice`, `stream`, `prompt_cache_options`, `prompt_cache_key` and `safety_identifier`. The reasoning comes back as a `reasoning` item with a `summary_text` part, and a stream carries numbered events from `response.created` to `response.completed`. The API is stateless: `previous_response_id` and `conversation` are ignored, so send the whole conversation in `input`. The `web_search` tool is ignored.

To use Kimi K3 in Codex, add a provider to `~/.codex/config.toml` and set `SEEDROUTER_API_KEY`:

```toml
model = "kimi-k3"
model_provider = "seedrouter"
model_context_window = 1048576

[model_providers.seedrouter]
name = "SeedRouter"
base_url = "https://api.seedrouter.ai/v1"
env_key = "SEEDROUTER_API_KEY"
wire_api = "responses"
```

## Anthropic Messages format

Code written for the Anthropic Messages API can call Kimi K3 too: send the Messages body to `/v1/messages` with `"model": "kimi-k3"`. `system`, `max_tokens`, `tools`, `tool_choice` (`auto`, `none`) and `output_config.effort` (`low`, `high`, `max`) are applied, and `metadata.user_id` and `cache_control` are accepted. `stop_sequences` (up to 5), `tool_choice` `any` and `output_config.format` are accepted but have no effect. The reasoning comes back as `thinking` blocks. Images go in as `base64` sources.

## Errors

Errors use `{"error": {"code": ..., "message": "..."}}` (the Messages endpoint uses Anthropic's error shape). The `code` is a code from the [common error catalog](https://seedrouter.ai/docs/api/errors). Failed requests are not charged.

## Tips

* Start with `high` effort and move to `max` only for the hardest problems; `low` suits quick, simple steps.
* Set `max_completion_tokens` high enough for the reasoning as well as the answer: it is one budget for both.
* Keep long, reused context at the start of the prompt so later requests read it from the cache, and use the `1h` TTL when requests are spread out.

## Related

* [Kimi K3 pricing](https://seedrouter.ai/models/kimi-k3)
* [Download OpenAPI](https://seedrouter.ai/docs/kimi-k3.openapi.json)
* [Copyable Markdown](https://seedrouter.ai/docs/kimi-k3.md)
