Claude Opus 5.5 is live on SeedRouter
SeedRouter Docs

Claude Fable 5.1

Call Claude Fable 5.1 with the official Anthropic Messages API, or the OpenAI Chat Completions and Responses formats: adaptive thinking, a 1M-token context window and up to 128K output tokens.

View Markdown

Claude Fable 5.1 is Anthropic's current Fable model for demanding reasoning and long-running agentic work. Send the official Anthropic Messages request to SeedRouter: change the base URL and the API key, keep the body. The same model also answers the OpenAI Chat Completions and Responses formats.

Model ID

Model IDContext windowMax outputThinkingDefault effort
claude-fable-5-11M tokens128K tokensAdaptive, always onhigh

See the model page for current prices.

Quick example

curl https://api.seedrouter.ai/v1/messages \
  -H "x-api-key: $SEEDROUTER_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-fable-5-1",
    "max_tokens": 1024,
    "messages": [{"role": "user", "content": "Summarize the trade-offs of event sourcing in three bullet points."}]
  }'

Endpoints

FormatMethod and pathAuthentication
Anthropic MessagesPOST https://api.seedrouter.ai/v1/messagesx-api-key: <key> or Authorization: Bearer <key>, plus anthropic-version
OpenAI Chat CompletionsPOST https://api.seedrouter.ai/v1/chat/completionsAuthorization: Bearer <key>
OpenAI ResponsesPOST https://api.seedrouter.ai/v1/responsesAuthorization: Bearer <key>

The Messages endpoint forwards your request body as sent, including optional fields, and returns the official response. An anthropic-beta header is passed on as well. Keep the API key in server-side code.

Parameters

NameTypeRequiredDefaultNotes
modelstringYes—claude-fable-5-1.
max_tokensintegerYes—0–128000. Includes the tokens spent on thinking. 0 only pre-warms the prompt cache.
messagesobject[]Yes—Alternating user and assistant turns; content is a string or an array of content blocks. The last turn must be user. Exception: to continue a pause_turn response, send its content back as the last assistant message.
systemstring or object[]No—System prompt.
thinkingobjectNo{"type": "adaptive"}Thinking is adaptive and always on. display: omitted (default) or summarized.
output_config.effortenumNohighlow, medium, high, xhigh, max. Steers how much the model thinks.
output_config.formatobjectNo—A JSON schema for structured output.
stop_sequencesstring[]No—Stop when one of these strings is generated.
streambooleanNofalseStream the response as server-sent events.
temperaturenumberNo—Only 1 (the default) is accepted, for backwards compatibility; any other value returns a 400 error. Leave it out.
top_pnumberNo—Only values from 0.99 to 1 are accepted, for backwards compatibility; any other value returns a 400 error. Leave it out.
top_kintegerNo—Not accepted: any value returns a 400 error. Leave it out.
toolsobject[]No—Tool definitions.
tool_choiceobjectNo—auto or none; forcing a tool (any or tool) is not supported by this model.
metadata.user_idstringNo—An opaque id for your end user, up to 512 characters.
cache_controlobjectNo—Top-level prompt-cache breakpoint.
container, context_management, mcp_servers, diagnostics, service_tier, inference_geo, speed—No—Forwarded as sent.

Thinking and effort

Claude Fable 5.1 always uses adaptive thinking: the model decides how much to think, and output_config.effort steers it. Higher effort usually means more output tokens, a longer wait and a higher cost. With thinking.display set to summarized, the response includes thinking blocks you can show; with omitted, they are left out. Thinking tokens are billed as output tokens.

Media inputs

Images and PDFs go in a user turn's content as image and document blocks, with a url source or, as in the official API, a base64 source:

{"role": "user", "content": [
  {"type": "image", "source": {"type": "url", "url": "https://example.com/chart.png"}},
  {"type": "text", "text": "What does this chart show?"}
]}

Replace the example URL with a publicly reachable file of your own.

Billing dimensions

See the current rates on the model page. A request is billed by the tokens it uses:

  • input tokens,
  • output tokens, including thinking,
  • prompt-cache reads, and
  • prompt-cache writes, with separate 5-minute and 1-hour rates.

The charge is taken from the usage reported with the finished response. A request that fails is not charged. Your account's usage records show the exact charge for every request.

Output

A non-streaming request returns the official message object:

{
  "id": "msg_...",
  "type": "message",
  "role": "assistant",
  "model": "claude-fable-5-1",
  "content": [{"type": "text", "text": "..."}],
  "stop_reason": "end_turn",
  "usage": {"input_tokens": 18, "output_tokens": 4, "cache_read_input_tokens": 0, "cache_creation_input_tokens": 0}
}

With "stream": true the response is a stream of the official events: message_start, content_block_start, content_block_delta, content_block_stop, message_delta and message_stop. The final message_delta carries the output token count.

OpenAI-compatible formats

The same model answers the OpenAI formats, so existing OpenAI code only needs a new base URL and model ID:

from openai import OpenAI

client = OpenAI(api_key="YOUR_SEEDROUTER_KEY", base_url="https://api.seedrouter.ai/v1")

chat = client.chat.completions.create(
    model="claude-fable-5-1",
    messages=[{"role": "user", "content": "Hello"}],
)
response = client.responses.create(model="claude-fable-5-1", input="Hello")

These requests are converted to the Messages format, so a field with no Messages counterpart has no effect. Their responses carry the official OpenAI fields and may include a few extra usage fields; read usage.total_tokens and the official fields.

Errors

Errors on /v1/messages use the Anthropic shape, {"type": "error", "error": {"type": "...", "message": "..."}}; the other formats use {"error": {"code": ..., "message": "..."}}. The code is a code from the common error catalog. Failed requests are not charged.

Tips

  • Start with the default effort and raise it only for tasks that need more thinking; effort changes both quality and cost.
  • Set max_tokens high enough for the thinking as well as the answer: it is one budget for both.
  • Put long, reused context first and mark it with cache_control so later requests read it from the cache at the lower rate.