# Claude Sonnet 5.5

Use `claude-sonnet-5-5` with the Anthropic Messages format. This document distinguishes the official request contract from the behavior observed in compatibility tests. Some advanced options do not yet behave as specified; review the limitations before depending on them.

See the [model page](https://seedrouter.ai/models/claude-sonnet-5-5#pricing) for current input, output and cache prices.

## Quick start

```bash
curl https://api.seedrouter.ai/v1/messages \
  -H "x-api-key: $SEEDROUTER_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-5-5",
    "max_tokens": 1024,
    "messages": [{"role": "user", "content": "Explain how a rainbow forms in three sentences."}]
  }'
```

```python
import os
import anthropic

client = anthropic.Anthropic(
    api_key=os.environ["SEEDROUTER_API_KEY"],
    base_url="https://api.seedrouter.ai",
)

message = client.messages.create(
    model="claude-sonnet-5-5",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Explain how a rainbow forms in three sentences."}],
)
print("".join(block.text for block in message.content if block.type == "text"))
```

```javascript
import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic({
  apiKey: process.env.SEEDROUTER_API_KEY,
  baseURL: "https://api.seedrouter.ai",
});

const message = await client.messages.create({
  model: "claude-sonnet-5-5",
  max_tokens: 1024,
  messages: [{ role: "user", content: "Explain how a rainbow forms in three sentences." }],
});
for (const block of message.content) if (block.type === "text") console.log(block.text);
```

```go
package main

import (
	"bytes"
	"fmt"
	"io"
	"net/http"
	"os"
)

func main() {
	body := []byte(`{"model": "claude-sonnet-5-5", "max_tokens": 1024,
		"messages": [{"role": "user", "content": "Explain how a rainbow forms in three sentences."}]}`)
	req, _ := http.NewRequest("POST", "https://api.seedrouter.ai/v1/messages", bytes.NewReader(body))
	req.Header.Set("x-api-key", os.Getenv("SEEDROUTER_API_KEY"))
	req.Header.Set("anthropic-version", "2023-06-01")
	req.Header.Set("Content-Type", "application/json")
	resp, err := http.DefaultClient.Do(req)
	if err != nil {
		panic(err)
	}
	defer resp.Body.Close()
	out, _ := io.ReadAll(resp.Body)
	fmt.Println(string(out))
}
```

`POST /v1/messages` accepts `x-api-key` or Bearer authentication and `anthropic-version: 2023-06-01`. Send an `anthropic-beta` header for features whose official documentation requires one. Keep credentials in server-side code.

The model also accepts basic OpenAI Chat Completions (`POST /v1/chat/completions`) and Responses (`POST /v1/responses`) requests. Use Messages for the native parameters described below; OpenAI format conversion does not provide every Anthropic feature.

## Official parameter contract

Sonnet 5.5 has a 1M-token context window and a synchronous output limit of 128000 tokens. Batch-specific output limits do not apply to this endpoint. Optional properties have no application-supplied default unless the table says otherwise.

| Parameter                         | Type / required                    | Official constraints and defaults                                                                                                                                                                                                                                    |
| --------------------------------- | ---------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `model`                           | string, required                   | `claude-sonnet-5-5`.                                                                                                                                                                                                                                                 |
| `max_tokens`                      | integer, required                  | 0–128000, including thinking tokens. Officially, 0 populates the prompt cache without generating output; see the current limitation below.                                                                                                                           |
| `messages`                        | object array, required             | At least one conversation message, at most 100000. Each has `role` and `content`; content is a string or content-block array. Ordinary turns use `user`/`assistant`. Mid-conversation `system` messages follow the official placement rules.                         |
| `system`                          | string or text-block array         | Top-level instructions. Text blocks can include cache breakpoints.                                                                                                                                                                                                   |
| `thinking`                        | object                             | Default: `{"type":"adaptive"}`. The other supported mode is `{"type":"between_tools"}`. Manual budgets and `disabled` are rejected.                                                                                                                                  |
| `thinking.display`                | enum                               | Adaptive mode only: `omitted` (default) or `summarized`. An omitted summary does not mean thinking is disabled.                                                                                                                                                      |
| `thinking.block_binding`          | object, beta                       | Adaptive mode only. Requires `thinking-binding-controls-2026-08-01`; follow the official preserved-thinking contract.                                                                                                                                                |
| `output_config.effort`            | enum or null                       | `low`, `medium`, `high`, `xhigh`, `max`; default `high`. Null leaves the default in effect.                                                                                                                                                                          |
| `output_config.format`            | object or null                     | Structured JSON output: `{"type":"json_schema","schema":{...}}`. Use the supported JSON Schema subset.                                                                                                                                                               |
| `stream`                          | boolean                            | Default `false`; `true` returns SSE events.                                                                                                                                                                                                                          |
| `stop_sequences`                  | string array                       | Officially stops generation at a matching string. The current compatibility test did not enforce this behavior.                                                                                                                                                      |
| `temperature`                     | number or null                     | Only 1 is accepted for compatibility; omit it. Other non-null values are rejected.                                                                                                                                                                                   |
| `top_p`                           | number or null                     | Only 0.99–1 is accepted for compatibility; omit it.                                                                                                                                                                                                                  |
| `top_k`                           | no non-null value                  | Sampling is unsupported; omit this property.                                                                                                                                                                                                                         |
| `tools`                           | object array                       | Client tools have `name`, `input_schema`, and optional description / strictness settings. Server tools use their versioned official definitions.                                                                                                                     |
| `tool_choice`                     | object                             | `auto` (default) or `none`. `any` and named forced `tool` are rejected. `auto` may include `disable_parallel_tool_use`.                                                                                                                                              |
| `metadata.user_id`                | string or null                     | At most 512 characters; use an opaque identifier.                                                                                                                                                                                                                    |
| `cache_control`                   | object or null                     | `type: "ephemeral"`; `ttl: "5m"` (default) or `"1h"`. Sonnet 5.5 requires at least 512 cacheable tokens. Block-level cache breakpoints are also supported by the official API.                                                                                       |
| `diagnostics`                     | object or null                     | `previous_message_id`: string of at most 256 characters or null. Requests cache-divergence diagnostics.                                                                                                                                                              |
| `service_tier`                    | enum                               | `auto` (default) or `standard_only`.                                                                                                                                                                                                                                 |
| `speed`                           | enum or null                       | Omit or use `standard` / null. Sonnet 5.5 does not support `fast`.                                                                                                                                                                                                   |
| `inference_geo`                   | string or null                     | The official default comes from the account settings. An accepted request alone does not verify geographic processing.                                                                                                                                               |
| `fallbacks`                       | string, object array or null, beta | `"default"` or up to three fallback entries. Each requires `model`; optional overrides are `max_tokens`, `thinking`, `output_config` and `speed`. See the fallback rules below.                                                                                      |
| `fallback_credit_token`           | string, object or null             | A token from a prior refusal, or `{"token":"...","mode":"strict"}`. Object form requires `fallback-credit-2026-07-01`; mode is `strict` (default) or `best_effort`. Cannot accompany a non-null `fallbacks` value.                                                   |
| `container`                       | string, object or null             | Container ID, or container configuration with optional `id` and `skills` (at most 20). Skills use the official type, identifier and version fields.                                                                                                                  |
| `context_management`              | object or null                     | Official context-edit configuration, including `edits`; null omits the setting. Model-specific edit compatibility still applies.                                                                                                                                     |
| `mcp_servers`                     | object array                       | Official MCP server definitions, subject to the required beta version and server authentication. An empty-array probe does not verify remote MCP execution.                                                                                                          |
| `compaction`                      | object or null, beta               | `{"type":"summarize"}`, with `compact-2026-09-04`; null omits compaction. Enabled compaction cannot be combined with non-null `context_management`, stop sequences or structured-output format. The signed compaction behavior did not pass the current test.        |
| `messages[].output_config.effort` | enum, beta                         | Per-message effort on a system message; requires `mid-conversation-output-config-2026-07-01`. Effort-only system messages may appear anywhere; content-bearing system groups follow the official placement rules. It must not change effort in `between_tools` mode. |

`between_tools` accepts only its `type` property and effort `low`, `medium` or `high`. Do not send `display`, `budget_tokens` or `block_binding` with it. Example:

```json
{
  "model": "claude-sonnet-5-5",
  "max_tokens": 1024,
  "thinking": {"type": "between_tools"},
  "output_config": {"effort": "medium"},
  "messages": [{"role": "user", "content": "Explain this concept briefly."}]
}
```

Assistant prefilling is unsupported. To continue a `pause_turn`, replay the returned server-tool assistant content unchanged. Compaction summarizes existing history and is not an assistant prefill. Preserve thinking blocks and signatures exactly; do not move them between models or edit earlier history without following the official binding rules.

On the native Claude API, computer use requires `computer_toolset_20260801`; `computer_20251124` is rejected. Advisor configurations using `claude-opus-4-8`, `claude-opus-4-7` or `claude-sonnet-5` are also rejected for this executor.

## Fallback request fields

The official `fallbacks` beta retries eligible classifier refusals. It does not retry rate limits, overloads or server errors, and a refusal may remain unresolved. Send `server-side-fallback-2026-07-01` for `"default"` or an explicit list; `server-side-fallback-2026-06-01` supports the list only. Other dated versions are rejected.

An explicit list contains at most three entries with distinct models, none equal to the requested model. Allowed targets come from the beta Models API's `allowed_fallback_models`. Only `model`, `max_tokens`, `thinking`, `output_config` and `speed` are permitted per entry; overrides must be valid for that target. The July beta converts Sonnet 5.5 `between_tools` to Sonnet 5 `disabled` with omitted display when that fallback occurs. With the June beta, provide the Sonnet 5 thinking override yourself.

`fallback_credit_token` is for a separate retry after a refusal. A string selects strict redemption; an object adds `mode`. In `strict` mode, failed redemption rejects the retry. In `best_effort` mode, a token-layer failure can proceed at the normal price and is recorded in `usage.fallback_credit`; malformed tokens and combining credit with `fallbacks` still fail. Redemption also requires the eligible request, account, workspace, platform and five-minute window described in the [official credit guide](https://platform.claude.com/docs/en/build-with-claude/fallback-credit).

A benign request with `fallbacks: "default"`, the July beta header and `speed: "standard"` returned the expected text. This establishes acceptance only: fallback execution and credit redemption have **not been verified end to end** here.

## Media and tool inputs

Images use `image` blocks and PDFs use `document` blocks in a user message. The official source types include public URLs and base64 with the corresponding MIME type. The compatibility checks used a base64 PNG and a one-page base64 PDF, and verified the contents of the answers. They did not exercise every URL, file-size, image-resolution or PDF-page boundary.

Client-side tools use the standard `tool_use` → `tool_result` exchange. Keep tool-use IDs unchanged and return the result in a user message. A successful strict-tool example verifies that example's arguments, not every supported JSON Schema keyword.

## Responses

A non-streaming response contains `id`, `type: "message"`, `role: "assistant"`, `model`, `content`, `stop_reason`, `stop_sequence` and `usage`, plus optional official fields such as `container`, `diagnostics`, `context_management`, `stop_details` and beta response fields. Content can include text, thinking, tool calls, tool results or other official block types; do not assume the first block is text.

With streaming, handle `message_start`, `content_block_start`, `content_block_delta`, `content_block_stop`, `message_delta` and `message_stop`. Errors can also occur inside a stream. Usage can include ordinary input/output tokens, thinking-token details, cache reads, and separate 5-minute / 1-hour cache creation counts.

An official classifier refusal is a normal response with `stop_reason: "refusal"` and `stop_details`, rather than an HTTP error. In a fallback response, `model` identifies the model that answered, `fallback` content blocks mark transitions, and `usage.iterations` describes the attempts. Inspect these fields instead of assuming the requested model served the response. These response behaviors remain unverified here.

Messages errors use `{"type":"error","error":{"type":"...","message":"..."}}`. Failed requests are not charged.

## Compatibility verification: 2026-10-01

| Result                                      | Checked behavior                                                                                                                                                                                                                                                                   |
| ------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Observed working                            | Basic text, ordinary multi-turn recall, non-conflicting string/block system instructions, streaming, adaptive and between-tools requests, JSON output, auto/none tools, a strict tool call, tool-result replay, base64 image/PDF inputs, and 5m/1h cache-write/read usage.         |
| Rejected according to the model contract    | Invalid output-token budgets, removed sampling settings, manual/disabled thinking, invalid between-tools combinations, forced tools, assistant prefill, legacy computer tools, and oversized metadata IDs.                                                                         |
| Accepted, effect not established            | All five effort levels, summarized thinking, binding settings, metadata, service tier, region selection, diagnostics, an empty context-edit list, null container, empty MCP list and computer-toolset declaration. Toolset declaration does not establish successful computer use. |
| Known mismatch                              | `max_tokens: 0` returned 400. A stop-sequence request returned the stop string and trailing text. On-demand compaction returned ordinary text rather than a signed compaction block.                                                                                               |
| Additional behavior requiring investigation | A per-message effort request did not recall the earlier value; a conflicting system/user instruction probe followed the user instruction. These results do not establish that every system prompt or multi-turn request fails.                                                     |

Beta tool execution, real MCP connections, Files API references, geographic residency, thinking-signature replay, full context/output ceilings, media boundary cases and refusal/fallback behavior have not been validated end to end. HTTP 200 and a returned model name do not authenticate which model ran or prove all supplied options took effect.

## References

* [Official Sonnet 5.5 overview](https://platform.claude.com/docs/en/models/sonnet-5-5/overview)
* [Sonnet 5.5 changes and restrictions](https://platform.claude.com/docs/en/models/sonnet-5-5/whats-new-sonnet-5-5)
* [Messages parameter reference](https://platform.claude.com/docs/en/api/messages/create)
* [Per-message effort](https://platform.claude.com/docs/en/build-with-claude/effort)
* [On-demand compaction](https://platform.claude.com/docs/en/build-with-claude/compaction-on-demand)
* [Refusals and fallback](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback)
* [Fallback credit](https://platform.claude.com/docs/en/build-with-claude/fallback-credit)
* [OpenAPI schema](https://seedrouter.ai/docs/claude-sonnet-5-5.openapi.json)
* [Copyable Markdown](https://seedrouter.ai/docs/claude-sonnet-5-5.md)
