Kimi K3 parameters, context window and max output tokens
Kimi K3 specs: 2.8 trillion parameters, 16 of 896 experts active per token, a 1M-token context window and up to 1,048,576 output tokens.
Read as MarkdownKimi K3 has 2.8 trillion parameters and activates 16 of its 896 experts for each token. Its context window is 1,048,576 tokens, and a single reply can be up to 1,048,576 tokens long, although the default limit is 131,072. Moonshot AI calls it "the world's first open-source model in the 3-trillion-parameter class" and has released the full weights. It always reasons, reads text and images, and writes text.
Kimi K3 specs at a glance
| Spec | Value |
|---|---|
| Developer | Moonshot AI |
| Total parameters | 2.8 trillion |
| Experts | 896, with 16 active per token |
| Architecture | Mixture of experts with Kimi Delta Attention and Attention Residuals |
| Context window | 1,048,576 tokens |
| Max output | 1,048,576 tokens, reasoning included |
| Default output limit | 131,072 tokens |
| Reasoning | Always on; low, high or max (default max) |
| Input | Text and images (images as base64 only) |
| Output | Text |
| Weights | Released by Moonshot AI |
| Model ID on SeedRouter | kimi-k3 |
How many parameters does Kimi K3 have?
2.8 trillion. Moonshot AI's quickstart calls Kimi K3 its "most capable flagship model to date, with 2.8 trillion parameters" and "the first open-source model to reach 2.8 trillion parameters".
How many active parameters does Kimi K3 have?
Moonshot AI's quickstart does not give an active parameter count. What it does say is that the model "efficiently activates 16 out of 896 experts" through its Stable LatentMoE framework. So only a small share of the 2.8 trillion parameters does the work for each token, which keeps the compute per token down. All the weights still have to be loaded, which is why running Kimi K3 yourself takes multi-GPU, data-center hardware.
What architecture does Kimi K3 use?
A mixture-of-experts model built on two attention changes. In Moonshot AI's words, it "is built on Kimi Delta Attention (KDA), a hybrid linear attention mechanism, and Attention Residuals". Moonshot AI says these changes, together with the sparser expert layout and new training methods, give Kimi K3 "roughly 2.5x the overall scaling efficiency of K2". The full details are due in the Kimi K3 technical report.
What is the Kimi K3 context window?
1,048,576 tokens, usually written as 1M. Prompt, images, conversation history and the reply all share this one window.
What is the Kimi K3 max output?
Up to 1,048,576 tokens per request, set with max_completion_tokens. The default is 131,072. Moonshot AI's quickstart: "max_completion_tokens defaults to 131072 and can be set up to 1048576."
Two things to know:
- Reasoning counts toward the limit. Kimi K3 always reasons, and its reasoning tokens use the same budget as the answer.
max_tokensis the deprecated name for the same limit. Usemax_completion_tokens.
If a reply is cut short, raise max_completion_tokens or lower reasoning_effort so less of the budget goes to thinking. Context window vs max output tokens explains how to spot a truncated reply in each API format.
Can you turn off Kimi K3's reasoning?
No. Kimi K3 always reasons. You can set reasoning_effort to low, high or max; the default is max, which is the slowest and uses the most output tokens. high suits most tasks and low suits quick, simple steps. The reasoning comes back in reasoning_content, next to content, and is billed as output tokens.
What settings are fixed?
Moonshot AI fixes Kimi K3's sampling: temperature 1.0, top_p 0.95, n 1, presence_penalty 0 and frequency_penalty 0. Any other value returns an error, so leave these fields out of the request.
Does Kimi K3 read images?
Yes, but only as base64 data URIs. A public image URL is not accepted and returns 400, as on Kimi's own API. The output is always text.
How do I call Kimi K3?
On SeedRouter, Kimi K3 takes the OpenAI Chat Completions, OpenAI Responses and Anthropic Messages formats with the same key and the model ID kimi-k3. The Kimi K3 API guide walks through the first request, the Claude Code and Codex setup has tested configs, and the Kimi K3 API reference lists every parameter. Live prices are on the Kimi K3 page.
Frequently asked questions
How big is Kimi K3?
2.8 trillion parameters in total, with 16 of 896 experts active per token.
Is Kimi K3 open source?
Moonshot AI has released the full model weights. Is Kimi K3 free? covers what running them yourself takes.
Why does Kimi K3 stop at 131,072 tokens?
That is the default output limit. Set max_completion_tokens up to 1048576 for longer replies.
Who makes Kimi K3?
Moonshot AI, the Chinese company behind the Kimi assistant.



