Text
Claude Haiku 5.5
Claude Haiku 5.5 Messages reference: parameters, thinking, forced tools, cache usage, beta features, streaming and response handling.
Use claude-haiku-5-5 with POST https://api.seedrouter.ai/v1/messages. The model accepts text, images and documents and returns text or tool requests. The model page shows current token rates.
The contract below follows Anthropic's model-specific documentation, checked October 9, 2026. Official capability limits and end-to-end verification are separate: a field being accepted does not prove that its intended effect occurred. See the compatibility results below before using advanced options.
Quick start
curl https://api.seedrouter.ai/v1/messages \
-H "x-api-key: $SEEDROUTER_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-haiku-5-5",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "Classify this request as billing, technical or account: I was charged twice. Return only the label."}]
}'import os
import anthropic
client = anthropic.Anthropic(
api_key=os.environ["SEEDROUTER_API_KEY"],
base_url="https://api.seedrouter.ai",
)
message = client.messages.create(
model="claude-haiku-5-5",
max_tokens=1024,
messages=[{"role": "user", "content": "Summarize the purpose of a database index."}],
)
for block in message.content:
if block.type == "text":
print(block.text)import Anthropic from '@anthropic-ai/sdk';
const client = new Anthropic({
apiKey: process.env.SEEDROUTER_API_KEY,
baseURL: 'https://api.seedrouter.ai',
});
const message = await client.messages.create({
model: 'claude-haiku-5-5',
max_tokens: 1024,
messages: [{ role: 'user', content: 'Summarize the purpose of a database index.' }],
});
for (const block of message.content) {
if (block.type === 'text') console.log(block.text);
}Keep your API key on the server. Select response blocks by type; a response can begin with thinking or a tool call.
package main
import (
"bytes"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"time"
)
func main() {
body, err := json.Marshal(map[string]any{
"model": "claude-haiku-5-5",
"max_tokens": 1024,
"messages": []map[string]string{
{"role": "user", "content": "Summarize the purpose of a database index."},
},
})
if err != nil { panic(err) }
req, err := http.NewRequest("POST", "https://api.seedrouter.ai/v1/messages", bytes.NewReader(body))
if err != nil { panic(err) }
req.Header.Set("x-api-key", os.Getenv("SEEDROUTER_API_KEY"))
req.Header.Set("anthropic-version", "2023-06-01")
req.Header.Set("Content-Type", "application/json")
client := &http.Client{Timeout: 2 * time.Minute}
res, err := client.Do(req)
if err != nil { panic(err) }
defer res.Body.Close()
data, err := io.ReadAll(res.Body)
if err != nil { panic(err) }
if res.StatusCode >= 400 { panic(fmt.Sprintf("HTTP %d: %s", res.StatusCode, data)) }
fmt.Println(string(data))
}Request parameters
There are 25 top-level fields in the native contract. Optional does not mean nullable: only the rows explicitly mentioning null accept it. Unknown fields and unsupported sampling fields are discarded before forwarding, following the text-model parameter policy. Invalid values of supported fields return an invalid_request_error before generation.
| Field | Required | Contract |
|---|---|---|
model | Yes | claude-haiku-5-5. |
max_tokens | Yes | Integer 0–128000, including thinking. No API default. The Playground starts at 8192. |
messages | Yes | 1–100000 messages with a role and string or content-block array. See conversation rules below. |
system | No | String or array of text blocks. Not null. |
thinking | No | Adaptive by default, or disabled. No manual budget and no between-tools mode. |
output_config | No | Object with effort, format and optional beta task_budget. |
stop_sequences | No | Array of stopping strings. |
stream | No | Boolean; false by default. |
temperature | No | Omit. Discarded here; the official compatibility value is 1. |
top_p | No | Omit. Discarded here; the official compatibility value is 0.99. |
top_k | No | Unsupported and discarded. |
tools | No | Array of client tools or official server tool declarations. |
tool_choice | No | auto, none, any, or named tool. Forced tools are supported. |
metadata | No | Object; optional user_id is a string of at most 512 characters or null. |
cache_control | No | Null or {"type":"ephemeral","ttl":"5m"}; TTL also accepts 1h. Default TTL is 5m. |
container | No | Null, container ID string, or object with optional ID and up to 20 skills. |
context_management | No | Null or an object containing official context edits; beta headers apply. |
mcp_servers | No | Array of at most 20 URL servers; requires a matching MCP beta header. |
service_tier | No | auto or standard_only. Haiku has no Priority Tier capacity. |
inference_geo | No | global, us, or null. Omission uses the account default; inspect reported usage before assuming a region. |
diagnostics | No | Null or object; previous_message_id is null or a string of at most 256 characters. |
compaction | No | Null or {"type":"summarize","instructions":"..."}. Instructions are optional, nullable and at most 16384 characters. |
fallbacks | No | Null or default with the corresponding beta. Haiku has no automatic fallback models; explicit lists are invalid. |
fallback_credit_token | No | Null, a token string, or {token,mode}. The API must verify eligibility and validity; do not assume any model is an eligible target. |
speed | No | standard or null. Fast mode is unsupported. |
The Playground provides controls for the supported fields, including JSON controls for nested structures. Sampling parameters and the fixed standard speed are omitted from the form. The model ID is fixed to this page. Use the JSON request preview to inspect the submitted body.
Thinking and effort
The default is adaptive thinking with medium effort and omitted thinking text. Effort accepts low, medium, high, xhigh, max, or null to use the default.
{
"thinking": {"type": "adaptive", "display": "summarized"},
"output_config": {"effort": "medium"}
}To disable thinking, use {"type":"disabled"} with low, medium or high effort. Do not include display or block_binding in disabled mode. enabled, budget_tokens, between_tools, and disabled thinking at xhigh/max are invalid.
Adaptive display accepts omitted, summarized, or null. The generic beta value updates requires thinking-display-updates-2026-08-18; Anthropic does not currently establish readable progress updates for Haiku, so do not depend on that output.
Optional thinking.block_binding requires thinking-binding-controls-2026-08-01. It is null or an object whose prefix_mismatch_behavior is error, drop_block, or null. Keep earlier conversation turns and complete thinking blocks unchanged when replaying history. Thinking signatures are bound to the account that produced them or an account linked to it.
output_config.task_budget is null or { "type": "tokens", "total": 20000 } with optional integer/null remaining. It requires task-budgets-2026-03-13; total must be at least 20000. No additional remaining range is imposed here.
Tools and structured output
Client tools require a name of 1–128 letters, digits, underscores or hyphens, and an input_schema with type: "object". Use tool_choice: {"type":"any"} or {"type":"tool","name":"lookup"} to force a declared tool. With adaptive thinking, a forced tool response begins at the tool call without a thinking block.
disable_parallel_tool_use is an optional boolean for auto, any and tool choices; it is not a field of none. Send the tool result back with the original tool_use_id. The Playground displays calls but does not execute your client tools.
{
"tools": [{
"name": "lookup",
"description": "Look up a product by SKU.",
"input_schema": {
"type": "object",
"properties": {"sku": {"type": "string"}},
"required": ["sku"],
"additionalProperties": false
}
}],
"tool_choice": {"type": "tool", "name": "lookup"}
}Structured responses use output_config.format: {"type":"json_schema","schema":{...}}. Follow Anthropic's supported JSON Schema subset, including additionalProperties: false on objects. A valid shape does not guarantee factually correct values. Strict tools and structured output have schema-wide limits; see the official structured-output reference.
Computer use requires computer_toolset_20260801; the old computer tool versions are invalid. Browser use has its own browser_toolset_20260801. Declaring a tool does not verify that a complete server-tool session works. Review the tool's official guide and any beta requirements before using it.
Conversations and context management
Ordinary assistant prefill is unsupported. A paused server-tool continuation is different: resend the complete assistant blocks as directed by the Messages protocol.
A content-bearing system message can appear after a user message or paused server-tool result. It must be followed by an assistant message or be the last message. Consecutive system messages are evaluated as one group. Do not insert one between a client tool call and its required result.
An empty-content system message may change only output_config.effort with mid-conversation-output-config-2026-07-01. It can appear anywhere. While thinking is disabled, it cannot change the effective effort. System clear_at accepts never, next_user_message, or null with mid-conversation-system-clear-at-2026-08-21; turn-scoped messages allow text only, without output config or block caching.
Context edits include:
| Edit | Beta | Main constraints |
|---|---|---|
clear_tool_uses_20250919 | context-management-2025-06-27 | Trigger count at least 1; keep count at least 0. |
clear_thinking_20251015 | context-management-2025-06-27 | Keep all, or at least one thinking turn. Place before tool-use clearing when combining edits. |
compact_20260112 | compact-2026-01-12 | Input-token trigger at least 50000; default 150000. |
On-demand compaction requires compact-2026-09-04. It cannot be combined with context_management, stop_sequences, an output format, forced tools or task_budget.remaining. A signed compaction block also cannot be combined with task_budget.remaining or threshold compaction. Preserve the returned block and signature when continuing.
Images, PDFs and request size
Images accept JPEG, PNG, GIF and WebP by URL, base64, or file reference. PDFs accept URL, base64 or file reference. File references require the relevant Files API beta and valid file access. Text documents can use text or content sources.
The native request limit is 32 MB. Official image limits are up to 600 images, 10 MB of base64-encoded data per image and 8000 pixels on either edge; requests with many images can have tighter platform-specific limits. PDFs must be unencrypted and have at most 600 pages for this model's context size. The API remains responsible for inspecting remote files; local shape checks cannot prove the contents of a URL.
The Playground uploads attachments before submitting URLs. Conversation JSON also supports native media content blocks. Full media-size and context-window boundary probes are separate from a small example request.
Prompt caching and billing
Haiku's minimum cacheable prompt is 512 tokens. Smaller marked prompts can run without creating a cache entry. Use at most four cache breakpoints; automatic top-level cache control consumes one slot. Put longer-lived cache prefixes before shorter-lived ones.
max_tokens: 0 requests cache prewarming without answer generation. It cannot accompany stream: true, structured output or forced tool use. Keep thinking and effort settings consistent between cache preparation and the requests that reuse it.
Read usage.input_tokens, output_tokens, cache_creation_input_tokens, cache_read_input_tokens, and the 5m/1h breakdown under cache_creation. Thinking is included in output tokens; a reported thinking-token breakdown is not an extra charge to add again. Current rates are on the pricing section, with further explanation in the pricing guide.
Responses, streaming and errors
A completed response contains id, type: "message", role: "assistant", model, content, stop_reason, stop_sequence and usage. Optional container, diagnostics, context_management, stop_details and input_transformations are preserved when returned.
Handle end_turn, max_tokens, stop_sequence, tool_use, pause_turn, compaction, refusal and model_context_window_exceeded. A limit stop or refusal is not the same as an HTTP error. Never treat the first content block as guaranteed text.
Streaming uses the Messages SSE events: message_start, content_block_start, content_block_delta, content_block_stop, message_delta and message_stop. Also handle ping and error events. Preserve thinking signatures and tool blocks needed by later turns.
Errors use the Anthropic shape:
{"type":"error","error":{"type":"invalid_request_error","message":"max_tokens must be an integer from 0 to 128000."}}Requests that return an error are not charged. See error handling for shared error types.
OpenAI-compatible formats
The same ID is available with /v1/chat/completions and /v1/responses. Use their native fields: Chat uses messages; Responses uses input. Native Claude options belong to Messages and should not be copied wholesale into an OpenAI-format body.
curl https://api.seedrouter.ai/v1/chat/completions \
-H "Authorization: Bearer $SEEDROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"claude-haiku-5-5","max_tokens":256,"messages":[{"role":"user","content":"Reply with OK."}]}'curl https://api.seedrouter.ai/v1/responses \
-H "Authorization: Bearer $SEEDROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"claude-haiku-5-5","max_output_tokens":256,"input":"Reply with OK."}'Compatibility results
Checked October 9, 2026 in the development environment. These checks establish the observed behavior of specific requests, not every official limit or production deployment.
| Capability | Observed result |
|---|---|
| Native Messages and SSE | Text response and complete event sequence verified. |
| Classification | Returned Billing; 41 input and 5 output tokens. |
| Structured JSON and client tools | JSON values, automatic/none/named/any selection, strict-tool arguments and tool-result continuation verified. |
| Images and PDFs | Returned the expected image color and PDF marker from base64 fixtures. Full media boundaries were not tested. |
| Cache prewarming | max_tokens: 0 returned no generated text and zero output tokens. |
| Five-minute and one-hour caching | Creation and subsequent cache-hit usage verified for both TTLs. |
| Stop sequences | Returned the requested stop reason and stopped before the excluded suffix. |
| Thinking and effort | All five effort values were accepted. Some explicit disabled-thinking requests still returned thinking blocks. Acceptance alone does not verify effort behavior. |
| System instructions and per-message effort | Results were inconsistent; a larger-budget per-message test still returned unrelated text. Test your exact conversation before rollout. |
| On-demand compaction | Returned a signed compaction block and stop_reason: compaction. Complete replay and billing validation remain outstanding. |
| Metadata and inference geography | metadata.user_id returned a permission error; explicit geography returned an account-type restriction. |
| MCP | The current MCP beta returned a credential restriction. A complete MCP session was not verified. |
| OpenAI Chat and Responses | Basic requests and explicit max/none reasoning requests returned the expected answer. Reasoning semantics were not independently established. |
| Other beta fields | Task budget, binding controls and fallback default were accepted; full feature semantics were not established. |
One-hour cache-write and compaction billing have not passed release validation. Maximum-context/output runs, hosted tools with separate charges, Files API access and fallback-credit redemption were not tested. Keep the official request shapes; do not infer support from a successful status alone.
