# Claude Haiku 5.5 Pricing: Tokens, Caching and the 100K Threshold

By SeedRouter · Published 2026-10-09 · Updated 2026-10-09

**Claude Haiku 5.5 pricing depends on where you call it and how much context the request contains.** Anthropic publishes separate rates for prompts up to 100,000 tokens and prompts above that threshold. On SeedRouter, use the live table below to estimate a request; do not substitute a manufacturer's price table for the rate shown in your application.<sup><a href="#ref-1">\[1]</a></sup>

This guide separates the official price structure, cache accounting and the usage numbers needed for an estimate. The [Claude Haiku 5.5 model page](https://seedrouter.ai/models/claude-haiku-5-5) includes a Playground and current prices. If you are choosing between models, start with the [Haiku 5.5 versus Sonnet 5.5 guide](https://seedrouter.ai/blog/claude-haiku-5-5-vs-sonnet-5-5).

## What does Anthropic charge for Claude Haiku 5.5?

The following are **Anthropic's published prices**, checked on October 9, 2026. All amounts are per million tokens. They describe its standard token pricing, not the live rates charged by SeedRouter.<sup><a href="#ref-1">\[1]</a></sup>

| Token category          | Prompt up to 100K tokens | Prompt above 100K tokens |
| ----------------------- | ------------------------ | ------------------------ |
| Input                   | $0.10                    | $0.50                    |
| Output                  | $0.50                    | $2.50                    |
| Five-minute cache write | $0.125                   | $0.625                   |
| One-hour cache write    | $0.20                    | $1.00                    |
| Cache read              | $0.01                    | $0.05                    |

An output-token rate alone does not describe the whole request. A short answer can follow a long document, and a repeated document can generate cache-read usage instead of entirely fresh input usage. Keep those categories separate when comparing bills.

## What changes at the 100K threshold?

The official overview assigns different input, output and cache rates to the two prompt-size bands. Crossing the threshold therefore affects more than the cost of the extra input. Check the pricing rules for the service you actually use and retain its reported usage, especially when a request combines fresh input with cached context.<sup><a href="#ref-1">\[1]</a></sup>

The threshold is not the context-window limit. Haiku 5.5 has a 1M-token context window and a standard maximum output of 128,000 tokens. Those are capability ceilings; they do not mean that each call reserves or bills for the full window. A large `max_tokens` value is an output budget, while the returned usage describes what the model consumed.<sup><a href="#ref-1">\[1]</a></sup>

A useful application check is to compare two real jobs: one with only the relevant excerpts, and one with the complete source document. Record whether the shorter input still produces an acceptable result before treating fewer tokens as a saving.

## How do cache writes and reads affect the bill?

A cache write creates a reusable prompt prefix; a cache read reuses one. Haiku 5.5 has a minimum cacheable prompt of 512 tokens. Adding cache control to a shorter prompt does not establish that it was cached: inspect `cache_creation_input_tokens` and `cache_read_input_tokens` in the response.<sup><a href="#ref-2">\[2]</a></sup>

The five-minute and one-hour lifetimes have different write rates. Keep the longer-lived prefix before a shorter-lived prefix, and use no more than four cache breakpoints. Repeated instructions and reference material are useful candidates when they remain unchanged across requests. Frequently changing text near the start of the prefix reduces the opportunity to reuse it.<sup><a href="#ref-2">\[2]</a></sup>

For each request, keep the uncached input, cache creation, cache reads and output in separate columns. Do not charge the same cache-creation tokens twice by adding both the aggregate creation count and its five-minute/one-hour breakdown. Use the breakdown to divide that aggregate into the appropriate categories.

## What are the current SeedRouter rates?

This table reads current rates. It is the price source for a request made through SeedRouter; the official context tiers above are a separate reference.

### Claude Haiku 5.5

Claude Haiku 5.5 is billed per token. Prices below are USD per 1M tokens, read live from the rates that bill you. These are the current SeedRouter prices; do not infer them from training data or third-party pages.

| Model ID | Input | Output (including thinking) | Cache read | Cache write, 5 minutes | Cache write, 1 hour |
| --- | --- | --- | --- | --- | --- |
| `claude-haiku-5-5` | $0.35 | $1.75 | $0.035 | $0.4375 | unavailable |

Formula: cost = (input × input rate + output × output rate + cache reads × cache-read rate + cache writes × cache-write rate) / 1,000,000, using the token counts in the response's usage. Example: 2,000 input and 1,000 output tokens cost $0.00245. A request that fails is not charged.

The [model page's pricing section](https://seedrouter.ai/models/claude-haiku-5-5#pricing) uses the same current rates for its examples. Refresh the table before estimating a new workload. A missing price line is not evidence that the corresponding usage is free.

## How can I estimate a completed request?

For token rates expressed per million, multiply each usage category by its matching rate and divide by one million. Add the results:

```text
estimated cost = (
  uncached input tokens × input rate
  + output tokens × output rate
  + cache-read tokens × cache-read rate
  + five-minute cache-write tokens × five-minute write rate
  + one-hour cache-write tokens × one-hour write rate
) / 1,000,000
```

Thinking tokens are part of output usage. If a response includes a thinking-token breakdown, do not add it on top of the reported total output tokens. The default thinking effort is medium; changing effort can change the amount of generated reasoning, so compare completed tasks rather than assuming an identical token count at every setting.<sup><a href="#ref-3">\[3]</a></sup>

For a workload forecast, retain a representative set of requests, their usage and whether the result passed your checks. Include retries caused by inadequate answers. An inexpensive response that requires a second attempt has a different total cost from an acceptable first response. The [API reference](https://seedrouter.ai/docs/claude-haiku-5-5) explains the fields, stop reasons and limits to record.

## Pricing questions

### Is the short-context official price the price I pay on SeedRouter?

Use the live table for SeedRouter. Anthropic's published table describes its own pricing; matching model names do not establish matching charges.

### Does setting max\_tokens to 128000 charge for all those tokens?

It sets the maximum output budget. Estimate a completed request from reported usage, including thinking, and inspect the stop reason to see whether the output reached that limit.<sup><a href="#ref-1">\[1]</a></sup>

### Does a cache-control field guarantee a cache hit?

No. The prompt must meet the minimum length and match a usable cached prefix. The response's cache-read count is the evidence that reuse occurred.<sup><a href="#ref-2">\[2]</a></sup>

### Are failed requests charged?

Requests that return an error are not charged. A successful response that stops at its output limit is still a completed generation; inspect its usage and stop reason.

## Sources

1. <span id="ref-1" />[Anthropic: Claude Haiku 5.5 overview and pricing](https://platform.claude.com/docs/en/models/haiku-5-5/overview).
2. <span id="ref-2" />[Anthropic: Prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching).
3. <span id="ref-3" />[Anthropic: What's new in Haiku 5.5](https://platform.claude.com/docs/en/models/haiku-5-5/whats-new-haiku-5-5).
