Claude Opus 5.5 is live on SeedRouter
SeedRouter Docs

GPT-6 Luna

Call GPT-6 Luna with the official OpenAI Responses or Chat Completions API: a 1.05M-token context window, up to 128K output tokens and reasoning effort you choose.

View Markdown

GPT-6 Luna is OpenAI's most efficient GPT-6 model, built for focused, high-volume tasks, and the lowest-cost model in the GPT-6 family. Send the official OpenAI Responses or Chat Completions request to SeedRouter: change the base URL and the API key, keep the body.

Model ID

Model IDContext windowMax outputReasoning effortDefault effort
gpt-6-luna1.05M tokens (922K input)128K tokensnone, low, medium, high, xhigh, maxmedium

Input: text and images. Output: text. Knowledge cutoff: May 18, 2026. See the model page for current prices.

Quick example

curl https://api.seedrouter.ai/v1/responses \
  -H "Authorization: Bearer $SEEDROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-6-luna",
    "input": "Summarize the trade-offs of event sourcing in three bullet points."
  }'

Endpoints

FormatMethod and pathAuthentication
OpenAI ResponsesPOST https://api.seedrouter.ai/v1/responsesAuthorization: Bearer <key>
OpenAI Chat CompletionsPOST https://api.seedrouter.ai/v1/chat/completionsAuthorization: Bearer <key>
Anthropic MessagesPOST https://api.seedrouter.ai/v1/messagesx-api-key: <key> or Authorization: Bearer <key>, plus anthropic-version

The OpenAI endpoints forward your request body as sent and return the official response. Keep the API key in server-side code.

Parameters

Responses API fields:

NameTypeRequiredDefaultNotes
modelstringYes—gpt-6-luna.
inputstring or object[]Yes—A prompt, or input items: messages whose content holds input_text and input_image parts.
instructionsstringNo—System (developer) instructions.
max_output_tokensintegerNoModel maximum16–128000. Includes reasoning tokens.
reasoning.effortenumNomediumnone, low, medium, high, xhigh, max.
reasoning.summaryenumNo—auto, concise or detailed: return a summary of the reasoning.
reasoning.modeenumNostandardstandard or pro.
reasoning.contextenumNoautoauto, current_turn or all_turns: which earlier reasoning is shown to the model.
text.verbosityenumNomediumlow, medium or high.
text.formatobjectNo—Structured output, for example {"type": "json_schema", "name": "...", "schema": {...}}.
temperaturenumberNo—0–2. Accepted only with reasoning.effort set to none.
top_pnumberNo—0–1. Accepted only with reasoning.effort set to none.
top_logprobsintegerNo—0–20. Accepted only with reasoning.effort set to none.
streambooleanNofalseStream the response as server-sent events.
tools, tool_choice, parallel_tool_calls, max_tool_calls, include, metadata, store, truncation, prompt_cache_key, prompt_cache_options, prompt_cache_retention, context_management, moderation, prompt, previous_response_id, conversation, user—No—Forwarded as sent.

Not available: background, service_tier (including Batch, Flex and fast mode) and safety_identifier.

Reasoning and effort

GPT-6 Luna defaults to medium. At none the model answers without reasoning, and only then are temperature, top_p and top_logprobs accepted. Higher effort usually means more output tokens, a longer wait and a higher cost. Reasoning tokens are billed as output tokens and count toward max_output_tokens.

Image input

Images go in a user message's content as input_image parts with an image_url:

{"role": "user", "content": [
  {"type": "input_image", "image_url": "https://example.com/chart.png"},
  {"type": "input_text", "text": "What does this chart show?"}
]}

Replace the example URL with a publicly reachable image of your own.

Billing dimensions

See the current rates on the model page. A request is billed by the tokens it uses:

  • input tokens,
  • output tokens, including reasoning,
  • cached input tokens (usage.input_tokens_details.cached_tokens), and
  • cache-write tokens (usage.input_tokens_details.cache_write_tokens).

When a prompt has more than 272K input tokens, the whole request is billed at the long-context rates. The charge is taken from the usage reported with the finished response. A request that fails is not charged. Your account's usage records show the exact charge for every request.

Output

A non-streaming request returns the official response object:

{
  "id": "resp_...",
  "object": "response",
  "status": "completed",
  "model": "gpt-6-luna",
  "output": [{"type": "message", "role": "assistant", "content": [{"type": "output_text", "text": "..."}]}],
  "usage": {"input_tokens": 27, "input_tokens_details": {"cached_tokens": 0, "cache_write_tokens": 0}, "output_tokens": 102, "output_tokens_details": {"reasoning_tokens": 0}, "total_tokens": 129}
}

With "stream": true the response is a stream of the official events, including response.created, response.output_text.delta and response.completed; the final response.completed event carries the usage.

Chat Completions

Existing Chat Completions code only needs a new base URL and model ID:

from openai import OpenAI

client = OpenAI(api_key="YOUR_SEEDROUTER_KEY", base_url="https://api.seedrouter.ai/v1")

chat = client.chat.completions.create(
    model="gpt-6-luna",
    messages=[{"role": "user", "content": "Hello"}],
    max_completion_tokens=1024,
)
print(chat.choices[0].message.content)

Set the output limit with max_completion_tokens and the effort with reasoning_effort. On Chat Completions, GPT-6 Luna supports function calling only with reasoning_effort set to none; use the Responses API for reasoning with tools.

Anthropic Messages format

Code written for the Anthropic Messages API can call GPT-6 Luna too: send the Messages body to /v1/messages with "model": "gpt-6-luna". The request is converted to Chat Completions, so a field with no Chat Completions counterpart has no effect. The response carries the official Messages fields and may include a few extra usage fields; read usage.input_tokens, usage.output_tokens and the official fields.

Errors

Errors use {"error": {"code": ..., "message": "..."}}. The code is a code from the common error catalog. Failed requests are not charged.

Tips

  • For classification and extraction, try none first: it is the fastest and cheapest setting.
  • Set max_output_tokens high enough for the reasoning as well as the answer: it is one budget for both.
  • Keep long, reused context at the start of the prompt so later requests read it from the cache at the lower rate.