Claude Sonnet 5.5
Claude Sonnet 5.5 Messages reference: official parameters, adaptive and between-tools thinking, cache usage, response fields and tested compatibility limits.
Use claude-sonnet-5-5 with the Anthropic Messages format. This document distinguishes the official request contract from the behavior observed in compatibility tests. Some advanced options do not yet behave as specified; review the limitations before depending on them.
See the model page for current input, output and cache prices.
Quick start
curl https://api.seedrouter.ai/v1/messages \
-H "x-api-key: $SEEDROUTER_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-5-5",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "Explain how a rainbow forms in three sentences."}]
}'POST /v1/messages accepts x-api-key or Bearer authentication and anthropic-version: 2023-06-01. Send an anthropic-beta header for features whose official documentation requires one. Keep credentials in server-side code.
The model also accepts basic OpenAI Chat Completions (POST /v1/chat/completions) and Responses (POST /v1/responses) requests. Use Messages for the native parameters described below; OpenAI format conversion does not provide every Anthropic feature.
Official parameter contract
Sonnet 5.5 has a 1M-token context window and a synchronous output limit of 128000 tokens. Batch-specific output limits do not apply to this endpoint. Optional properties have no application-supplied default unless the table says otherwise.
| Parameter | Type / required | Official constraints and defaults |
|---|---|---|
model | string, required | claude-sonnet-5-5. |
max_tokens | integer, required | 0–128000, including thinking tokens. Officially, 0 populates the prompt cache without generating output; see the current limitation below. |
messages | object array, required | At least one conversation message, at most 100000. Each has role and content; content is a string or content-block array. Ordinary turns use user/assistant. Mid-conversation system messages follow the official placement rules. |
system | string or text-block array | Top-level instructions. Text blocks can include cache breakpoints. |
thinking | object | Default: {"type":"adaptive"}. The other supported mode is {"type":"between_tools"}. Manual budgets and disabled are rejected. |
thinking.display | enum | Adaptive mode only: omitted (default) or summarized. An omitted summary does not mean thinking is disabled. |
thinking.block_binding | object, beta | Adaptive mode only. Requires thinking-binding-controls-2026-08-01; follow the official preserved-thinking contract. |
output_config.effort | enum or null | low, medium, high, xhigh, max; default high. Null leaves the default in effect. |
output_config.format | object or null | Structured JSON output: {"type":"json_schema","schema":{...}}. Use the supported JSON Schema subset. |
stream | boolean | Default false; true returns SSE events. |
stop_sequences | string array | Officially stops generation at a matching string. The current compatibility test did not enforce this behavior. |
temperature | number or null | Only 1 is accepted for compatibility; omit it. Other non-null values are rejected. |
top_p | number or null | Only 0.99–1 is accepted for compatibility; omit it. |
top_k | no non-null value | Sampling is unsupported; omit this property. |
tools | object array | Client tools have name, input_schema, and optional description / strictness settings. Server tools use their versioned official definitions. |
tool_choice | object | auto (default) or none. any and named forced tool are rejected. auto may include disable_parallel_tool_use. |
metadata.user_id | string or null | At most 512 characters; use an opaque identifier. |
cache_control | object or null | type: "ephemeral"; ttl: "5m" (default) or "1h". Sonnet 5.5 requires at least 512 cacheable tokens. Block-level cache breakpoints are also supported by the official API. |
diagnostics | object or null | previous_message_id: string of at most 256 characters or null. Requests cache-divergence diagnostics. |
service_tier | enum | auto (default) or standard_only. |
speed | enum or null | Omit or use standard / null. Sonnet 5.5 does not support fast. |
inference_geo | string or null | The official default comes from the account settings. An accepted request alone does not verify geographic processing. |
fallbacks | string, object array or null, beta | "default" or up to three fallback entries. Each requires model; optional overrides are max_tokens, thinking, output_config and speed. See the fallback rules below. |
fallback_credit_token | string, object or null | A token from a prior refusal, or {"token":"...","mode":"strict"}. Object form requires fallback-credit-2026-07-01; mode is strict (default) or best_effort. Cannot accompany a non-null fallbacks value. |
container | string, object or null | Container ID, or container configuration with optional id and skills (at most 20). Skills use the official type, identifier and version fields. |
context_management | object or null | Official context-edit configuration, including edits; null omits the setting. Model-specific edit compatibility still applies. |
mcp_servers | object array | Official MCP server definitions, subject to the required beta version and server authentication. An empty-array probe does not verify remote MCP execution. |
compaction | object or null, beta | {"type":"summarize"}, with compact-2026-09-04; null omits compaction. Enabled compaction cannot be combined with non-null context_management, stop sequences or structured-output format. The signed compaction behavior did not pass the current test. |
messages[].output_config.effort | enum, beta | Per-message effort on a system message; requires mid-conversation-output-config-2026-07-01. Effort-only system messages may appear anywhere; content-bearing system groups follow the official placement rules. It must not change effort in between_tools mode. |
between_tools accepts only its type property and effort low, medium or high. Do not send display, budget_tokens or block_binding with it. Example:
{
"model": "claude-sonnet-5-5",
"max_tokens": 1024,
"thinking": {"type": "between_tools"},
"output_config": {"effort": "medium"},
"messages": [{"role": "user", "content": "Explain this concept briefly."}]
}Assistant prefilling is unsupported. To continue a pause_turn, replay the returned server-tool assistant content unchanged. Compaction summarizes existing history and is not an assistant prefill. Preserve thinking blocks and signatures exactly; do not move them between models or edit earlier history without following the official binding rules.
On the native Claude API, computer use requires computer_toolset_20260801; computer_20251124 is rejected. Advisor configurations using claude-opus-4-8, claude-opus-4-7 or claude-sonnet-5 are also rejected for this executor.
Fallback request fields
The official fallbacks beta retries eligible classifier refusals. It does not retry rate limits, overloads or server errors, and a refusal may remain unresolved. Send server-side-fallback-2026-07-01 for "default" or an explicit list; server-side-fallback-2026-06-01 supports the list only. Other dated versions are rejected.
An explicit list contains at most three entries with distinct models, none equal to the requested model. Allowed targets come from the beta Models API's allowed_fallback_models. Only model, max_tokens, thinking, output_config and speed are permitted per entry; overrides must be valid for that target. The July beta converts Sonnet 5.5 between_tools to Sonnet 5 disabled with omitted display when that fallback occurs. With the June beta, provide the Sonnet 5 thinking override yourself.
fallback_credit_token is for a separate retry after a refusal. A string selects strict redemption; an object adds mode. In strict mode, failed redemption rejects the retry. In best_effort mode, a token-layer failure can proceed at the normal price and is recorded in usage.fallback_credit; malformed tokens and combining credit with fallbacks still fail. Redemption also requires the eligible request, account, workspace, platform and five-minute window described in the official credit guide.
A benign request with fallbacks: "default", the July beta header and speed: "standard" returned the expected text. This establishes acceptance only: fallback execution and credit redemption have not been verified end to end here.
Media and tool inputs
Images use image blocks and PDFs use document blocks in a user message. The official source types include public URLs and base64 with the corresponding MIME type. The compatibility checks used a base64 PNG and a one-page base64 PDF, and verified the contents of the answers. They did not exercise every URL, file-size, image-resolution or PDF-page boundary.
Client-side tools use the standard tool_use → tool_result exchange. Keep tool-use IDs unchanged and return the result in a user message. A successful strict-tool example verifies that example's arguments, not every supported JSON Schema keyword.
Responses
A non-streaming response contains id, type: "message", role: "assistant", model, content, stop_reason, stop_sequence and usage, plus optional official fields such as container, diagnostics, context_management, stop_details and beta response fields. Content can include text, thinking, tool calls, tool results or other official block types; do not assume the first block is text.
With streaming, handle message_start, content_block_start, content_block_delta, content_block_stop, message_delta and message_stop. Errors can also occur inside a stream. Usage can include ordinary input/output tokens, thinking-token details, cache reads, and separate 5-minute / 1-hour cache creation counts.
An official classifier refusal is a normal response with stop_reason: "refusal" and stop_details, rather than an HTTP error. In a fallback response, model identifies the model that answered, fallback content blocks mark transitions, and usage.iterations describes the attempts. Inspect these fields instead of assuming the requested model served the response. These response behaviors remain unverified here.
Messages errors use {"type":"error","error":{"type":"...","message":"..."}}. Failed requests are not charged.
Compatibility verification: 2026-10-01
| Result | Checked behavior |
|---|---|
| Observed working | Basic text, ordinary multi-turn recall, non-conflicting string/block system instructions, streaming, adaptive and between-tools requests, JSON output, auto/none tools, a strict tool call, tool-result replay, base64 image/PDF inputs, and 5m/1h cache-write/read usage. |
| Rejected according to the model contract | Invalid output-token budgets, removed sampling settings, manual/disabled thinking, invalid between-tools combinations, forced tools, assistant prefill, legacy computer tools, and oversized metadata IDs. |
| Accepted, effect not established | All five effort levels, summarized thinking, binding settings, metadata, service tier, region selection, diagnostics, an empty context-edit list, null container, empty MCP list and computer-toolset declaration. Toolset declaration does not establish successful computer use. |
| Known mismatch | max_tokens: 0 returned 400. A stop-sequence request returned the stop string and trailing text. On-demand compaction returned ordinary text rather than a signed compaction block. |
| Additional behavior requiring investigation | A per-message effort request did not recall the earlier value; a conflicting system/user instruction probe followed the user instruction. These results do not establish that every system prompt or multi-turn request fails. |
Beta tool execution, real MCP connections, Files API references, geographic residency, thinking-signature replay, full context/output ceilings, media boundary cases and refusal/fallback behavior have not been validated end to end. HTTP 200 and a returned model name do not authenticate which model ran or prove all supplied options took effect.
