Claude Opus 5.5 is live on SeedRouter
SeedRouter Docs

GPT-6 Astra

Call GPT-6 Astra with the official OpenAI Responses or Chat Completions API: a 1.05M-token context window, up to 128K output tokens and reasoning effort you choose.

View Markdown

GPT-6 Astra is OpenAI's most capable model, built for the hardest end-to-end work: complex reasoning, coding, research and document creation. Send the official OpenAI Responses or Chat Completions request to SeedRouter: change the base URL and the API key, keep the body.

Model ID

Model IDContext windowMax outputReasoning effortDefault effort
gpt-6-astra1.05M tokens (922K input)128K tokenslow, medium, high, xhigh, maxNot documented

Input: text and images. Output: text. Knowledge cutoff: April 30, 2026. See the model page for current prices.

Quick example

curl https://api.seedrouter.ai/v1/responses \
  -H "Authorization: Bearer $SEEDROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-6-astra",
    "input": "Summarize the trade-offs of event sourcing in three bullet points."
  }'

Endpoints

FormatMethod and pathAuthentication
OpenAI ResponsesPOST https://api.seedrouter.ai/v1/responsesAuthorization: Bearer <key>
OpenAI Chat CompletionsPOST https://api.seedrouter.ai/v1/chat/completionsAuthorization: Bearer <key>
Anthropic MessagesPOST https://api.seedrouter.ai/v1/messagesx-api-key: <key> or Authorization: Bearer <key>, plus anthropic-version

The OpenAI endpoints forward your request body as sent and return the official response. Keep the API key in server-side code.

Parameters

Responses API fields:

NameTypeRequiredDefaultNotes
modelstringYes—gpt-6-astra.
inputstring or object[]Yes—A prompt, or input items: messages whose content holds input_text and input_image parts.
instructionsstringNo—System (developer) instructions.
max_output_tokensintegerNoModel maximum16–128000. Includes reasoning tokens.
reasoning.effortenumNoNot documentedlow, medium, high, xhigh, max.
reasoning.summaryenumNo—auto, concise or detailed: return a summary of the reasoning.
reasoning.modeenumNostandardstandard or pro.
reasoning.contextenumNoautoauto, current_turn or all_turns: which earlier reasoning is shown to the model.
text.verbosityenumNomediumlow, medium or high.
text.formatobjectNo—Structured output, for example {"type": "json_schema", "name": "...", "schema": {...}}.
temperaturenumberNo—0–2. Not accepted: GPT-6 Astra always reasons. Leave it out.
top_pnumberNo—0–1. Not accepted: GPT-6 Astra always reasons. Leave it out.
top_logprobsintegerNo—0–20. Not accepted: GPT-6 Astra always reasons. Leave it out.
streambooleanNofalseStream the response as server-sent events.
tools, tool_choice, parallel_tool_calls, max_tool_calls, include, metadata, store, truncation, prompt_cache_key, prompt_cache_options, prompt_cache_retention, context_management, moderation, prompt, previous_response_id, conversation, user—No—Forwarded as sent.

Not available: background, service_tier (including Batch, Flex and fast mode) and safety_identifier.

Reasoning and effort

GPT-6 Astra always reasons: none is not supported and returns a 400 error. OpenAI does not document a default effort for GPT-6 Astra; leave reasoning.effort out to use the model's own default. Higher effort usually means more output tokens, a longer wait and a higher cost. Reasoning tokens are billed as output tokens and count toward max_output_tokens.

Image input

Images go in a user message's content as input_image parts with an image_url:

{"role": "user", "content": [
  {"type": "input_image", "image_url": "https://example.com/chart.png"},
  {"type": "input_text", "text": "What does this chart show?"}
]}

Replace the example URL with a publicly reachable image of your own.

Billing dimensions

See the current rates on the model page. A request is billed by the tokens it uses:

  • input tokens,
  • output tokens, including reasoning,
  • cached input tokens (usage.input_tokens_details.cached_tokens), and
  • cache-write tokens (usage.input_tokens_details.cache_write_tokens).

When a prompt has more than 272K input tokens, the whole request is billed at the long-context rates. The charge is taken from the usage reported with the finished response. A request that fails is not charged. Your account's usage records show the exact charge for every request.

Output

A non-streaming request returns the official response object:

{
  "id": "resp_...",
  "object": "response",
  "status": "completed",
  "model": "gpt-6-astra",
  "output": [{"type": "message", "role": "assistant", "content": [{"type": "output_text", "text": "..."}]}],
  "usage": {"input_tokens": 27, "input_tokens_details": {"cached_tokens": 0, "cache_write_tokens": 0}, "output_tokens": 102, "output_tokens_details": {"reasoning_tokens": 0}, "total_tokens": 129}
}

With "stream": true the response is a stream of the official events, including response.created, response.output_text.delta and response.completed; the final response.completed event carries the usage.

Chat Completions

Existing Chat Completions code only needs a new base URL and model ID:

from openai import OpenAI

client = OpenAI(api_key="YOUR_SEEDROUTER_KEY", base_url="https://api.seedrouter.ai/v1")

chat = client.chat.completions.create(
    model="gpt-6-astra",
    messages=[{"role": "user", "content": "Hello"}],
    max_completion_tokens=1024,
)
print(chat.choices[0].message.content)

Set the output limit with max_completion_tokens and the effort with reasoning_effort. Chat Completions does not support function calling with GPT-6 Astra; use the Responses API for tools.

Anthropic Messages format

Code written for the Anthropic Messages API can call GPT-6 Astra too: send the Messages body to /v1/messages with "model": "gpt-6-astra". The request is converted to Chat Completions, so a field with no Chat Completions counterpart has no effect. The response carries the official Messages fields and may include a few extra usage fields; read usage.input_tokens, usage.output_tokens and the official fields.

Errors

Errors use {"error": {"code": ..., "message": "..."}}. The code is a code from the common error catalog. Failed requests are not charged.

Tips

  • Start with medium and raise the effort only for tasks that need deeper reasoning; effort changes both quality and cost.
  • Set max_output_tokens high enough for the reasoning as well as the answer: it is one budget for both.
  • Keep long, reused context at the start of the prompt so later requests read it from the cache at the lower rate.