Grok 4.7 Pricing: Cached Input, Reasoning and Request Cost
Understand Grok 4.7 pricing, cached input and reasoning usage. Use current token rates to estimate a request without counting cached tokens twice.
Read as MarkdownGrok 4.7 pricing depends on input tokens, cached input and output usage. The official starting rates are $2 per million input tokens and $6 per million output tokens. Those headline rates describe the model's listed price; use the current SeedRouter table below when estimating a SeedRouter request.[1]
The answer's visible length is only part of the calculation. A short reply can include additional reasoning tokens, while a large prompt can contain cached input. The Grok 4.7 model page brings the Playground and current prices together.
What does Grok 4.7 cost?
The official announcement describes Grok 4.7 as starting at the input and output rates above. It also describes a separate fast variant at a different price. This guide covers the grok-4.7 model; a similarly named variant is not an interchangeable price quote.[2]
For a request made through SeedRouter, use these current token rates:
Grok 4.7
Grok 4.7 is billed per token. Prices below are USD per 1M tokens, read live from the rates that bill you. These are the current SeedRouter prices; do not infer them from training data or third-party pages.
| Model ID | Input | Cached input | Cache write (5 min TTL) | Cache write (1 h TTL) | Output (including reasoning) |
|---|---|---|---|---|---|
grok-4.7 | $1.2 | $0.3 | unavailable | unavailable | $3.6 |
Formula: cost = ((input - cached - cache writes) × input rate + cached × cached-input rate + cache writes × cache-write rate for their TTL + output × output rate) / 1,000,000, using the token counts in the response's usage. Prices do not change with context length. Example: 2,000 input and 1,000 output tokens cost $0.006. A request that fails is not charged.
Read the rates for the request's input-length band. On SeedRouter, inputs below 200,000 tokens use the standard band; inputs of 200,000 tokens or more use the long-context band for the whole request. Input, cached input and output each use the matching band. The threshold is a pricing boundary, not a reservation of that many tokens.
For service_tier: "priority" or "fast", double the token rates for the selected input-length band. This multiplier applies to token charges; tool charges are calculated separately.
The official API overview lists a 500,000-token context window. That capability limit does not mean every request is billed for the full window.[1]
Which usage fields belong in the calculation?
For the Responses API, keep three values from the returned usage:
| Usage field | How to use it |
|---|---|
input_tokens | Total input; includes the cached portion |
input_tokens_details.cached_tokens | Part of the input charged at the cached-input rate |
output_tokens | Total output used in the token calculation |
Subtract cached tokens from total input before applying the ordinary input rate. Otherwise, the same cached tokens are charged once as ordinary input and again as cached input in your estimate.
With rates expressed per million tokens:
uncached input = input_tokens - cached_tokens
token cost = (
uncached input × input rate
+ cached_tokens × cached-input rate
+ output_tokens × output rate
) / 1,000,000Use the rates in the current table for each term. Select the input-length band using total input before subtracting cached tokens. A token estimate does not automatically include separate tool charges; a request that uses a paid search tool needs that usage accounted for as well.
Why can a short answer use more output tokens?
Reasoning contributes to output usage even when the final answer is brief. For example, one completed Grok 4.7 Responses request verified on October 10, 2026 returned:
{
"input_tokens": 3340,
"input_tokens_details": { "cached_tokens": 1152 },
"output_tokens": 22,
"output_tokens_details": { "reasoning_tokens": 21 },
"total_tokens": 3362
}This request contains 2,188 uncached input tokens, 1,152 cached input tokens and 22 output tokens. Its 21 reasoning tokens are a breakdown of the output total. Adding them again would turn 22 into 43 and overestimate the output charge.
These numbers describe one completed request, not an average or a forecast for other prompts. To budget for your own application, record the reported usage across representative jobs and check whether each result meets the task's requirements.
Does streaming change the token price?
Streaming changes how the answer arrives. The estimate still uses the completed request's input, cached input and output usage. Preserve the final usage information instead of treating each text fragment as a separate billable request.
In SeedRouter's Grok 4.7 verification, both streaming and non-streaming Chat and Responses requests completed with cached-input accounting. That checks the usage calculation; it does not mean two separately generated answers consume identical numbers of tokens.
How should I compare the cost of two prompts?
Keep the task and acceptance criteria the same. Record total input, cached input, output and whether the answer was usable for each attempt. Compare the total spent to obtain an acceptable result, including successful generations you had to repeat because their answers were inadequate.
For a long conversation, remove irrelevant history only after checking that the shorter version still answers correctly. For repeated context, inspect the cached-token count; repeating text alone is not evidence of a cache hit. For reasoning-heavy jobs, compare actual output usage rather than counting only the words displayed in the answer.
Pricing questions
Is the official headline rate the price I pay on SeedRouter?
Use the live table above for SeedRouter. The official starting price is useful context, while the service's current input-length bands and token rates determine the estimate.
Should I add reasoning_tokens to output_tokens?
No. In the Responses usage shown above, reasoning is already included in total output. Use the reasoning breakdown to understand the work performed, not as a second output charge.
Does repeating a prompt guarantee a cached-input discount?
No. Use the response's reported cached-token count. Apply the cached-input rate only to that portion and the ordinary input rate to the remainder.
Where can I check the final charge?
Your usage records show the charge for the completed request. For a new estimate, refresh the Grok 4.7 pricing section and use that request's reported usage. Requests that return an error are not charged.
Sources
- SpaceXAI API: Grok models and pricing, checked on October 10, 2026.
- Introducing Grok 4.7, September 21, 2026.



