# Kimi K3 parameters, context window and max output tokens

By SeedRouter · Published 2026-09-29 · Updated 2026-09-29

Kimi K3 has 2.8 trillion parameters and activates 16 of its 896 experts for each token. Its context window is 1,048,576 tokens, and a single reply can be up to 1,048,576 tokens long, although the default limit is 131,072. Moonshot AI calls it "the world's first open-source model in the 3-trillion-parameter class" and has released the full weights. It always reasons, reads text and images, and writes text.

## Kimi K3 specs at a glance

| Spec                   | Value                                                                |
| ---------------------- | -------------------------------------------------------------------- |
| Developer              | Moonshot AI                                                          |
| Total parameters       | 2.8 trillion                                                         |
| Experts                | 896, with 16 active per token                                        |
| Architecture           | Mixture of experts with Kimi Delta Attention and Attention Residuals |
| Context window         | 1,048,576 tokens                                                     |
| Max output             | 1,048,576 tokens, reasoning included                                 |
| Default output limit   | 131,072 tokens                                                       |
| Reasoning              | Always on; `low`, `high` or `max` (default `max`)                    |
| Input                  | Text and images (images as base64 only)                              |
| Output                 | Text                                                                 |
| Weights                | Released by Moonshot AI                                              |
| Model ID on SeedRouter | `kimi-k3`                                                            |

## How many parameters does Kimi K3 have?

2.8 trillion. Moonshot AI's quickstart calls Kimi K3 its "most capable flagship model to date, with 2.8 trillion parameters" and "the first open-source model to reach 2.8 trillion parameters".

## How many active parameters does Kimi K3 have?

Moonshot AI's quickstart does not give an active parameter count. What it does say is that the model "efficiently activates 16 out of 896 experts" through its Stable LatentMoE framework. So only a small share of the 2.8 trillion parameters does the work for each token, which keeps the compute per token down. All the weights still have to be loaded, which is why running Kimi K3 yourself takes multi-GPU, data-center hardware.

## What architecture does Kimi K3 use?

A mixture-of-experts model built on two attention changes. In Moonshot AI's words, it "is built on Kimi Delta Attention (KDA), a hybrid linear attention mechanism, and Attention Residuals". Moonshot AI says these changes, together with the sparser expert layout and new training methods, give Kimi K3 "roughly 2.5x the overall scaling efficiency of K2". The full details are due in the Kimi K3 technical report.

## What is the Kimi K3 context window?

1,048,576 tokens, usually written as 1M. Prompt, images, conversation history and the reply all share this one window.

## What is the Kimi K3 max output?

Up to 1,048,576 tokens per request, set with `max_completion_tokens`. The default is 131,072. Moonshot AI's quickstart: "`max_completion_tokens` defaults to 131072 and can be set up to 1048576."

Two things to know:

* **Reasoning counts toward the limit.** Kimi K3 always reasons, and its reasoning tokens use the same budget as the answer.
* **`max_tokens` is the deprecated name** for the same limit. Use `max_completion_tokens`.

If a reply is cut short, raise `max_completion_tokens` or lower `reasoning_effort` so less of the budget goes to thinking. [Context window vs max output tokens](https://seedrouter.ai/blog/context-window-vs-max-output-tokens) explains how to spot a truncated reply in each API format.

## Can you turn off Kimi K3's reasoning?

No. Kimi K3 always reasons. You can set `reasoning_effort` to `low`, `high` or `max`; the default is `max`, which is the slowest and uses the most output tokens. `high` suits most tasks and `low` suits quick, simple steps. The reasoning comes back in `reasoning_content`, next to `content`, and is billed as output tokens.

## What settings are fixed?

Moonshot AI fixes Kimi K3's sampling: `temperature` 1.0, `top_p` 0.95, `n` 1, `presence_penalty` 0 and `frequency_penalty` 0. Any other value returns an error, so leave these fields out of the request.

## Does Kimi K3 read images?

Yes, but only as base64 data URIs. A public image URL is not accepted and returns 400, as on Kimi's own API. The output is always text.

## How do I call Kimi K3?

On SeedRouter, Kimi K3 takes the OpenAI Chat Completions, OpenAI Responses and Anthropic Messages formats with the same key and the model ID `kimi-k3`. The [Kimi K3 API guide](https://seedrouter.ai/blog/kimi-k3-api) walks through the first request, the [Claude Code and Codex setup](https://seedrouter.ai/blog/kimi-k3-claude-code-codex) has tested configs, and the [Kimi K3 API reference](https://seedrouter.ai/docs/kimi-k3) lists every parameter. Live prices are on the [Kimi K3 page](https://seedrouter.ai/models/kimi-k3#pricing).

## Frequently asked questions

### How big is Kimi K3?

2.8 trillion parameters in total, with 16 of 896 experts active per token.

### Is Kimi K3 open source?

Moonshot AI has released the full model weights. [Is Kimi K3 free?](https://seedrouter.ai/blog/is-kimi-k3-free) covers what running them yourself takes.

### Why does Kimi K3 stop at 131,072 tokens?

That is the default output limit. Set `max_completion_tokens` up to 1048576 for longer replies.

### Who makes Kimi K3?

Moonshot AI, the Chinese company behind the Kimi assistant.
