Claude Opus 5.5 is live on SeedRouter
OpenAIText generation

GPT-6 Luna — low-cost GPT-6 for high-volume work

GPT-6 Luna is OpenAI's most efficient GPT-6 model, made for focused tasks at scale, with a 1.05M-token context. $0.08 / 1M input, $0.40 / 1M output per 1M tokens.

Available$0.08 / 1M input · $0.40 / 1M output$0.10 / 1M input · $0.50 / 1M output-20%
  • Reasoning effort control
  • 1.05M-token context
  • Image input
  • Prompt caching
  • Streaming
Context window
1.05M tokens
Max output
128K tokens
Effort levels
none to max
Knowledge cutoff
May 18, 2026

Try GPT-6 Luna

Chat with GPT-6 Luna, attach images, and see the tokens and cost of every answer.

Parameters

gpt-6-luna$0.08 in · $0.40 out per 1M tokens

Reasoning effort
Reasoning summary
Reasoning mode
Reasoning context
Verbosity

Temperature and Top P work only with reasoning effort none.

Conversation total$0.000000
Ready

Try one of these prompts

Messages are billed from your balance at the rates below, like an API call. Failed requests are free.

Real GPT-6 Luna output

One prompt, sent to GPT-6 Luna at low reasoning effort. Unedited.

Explain the difference between a mutex and a semaphore in two short paragraphs, and say when to use each.

A mutex is a lock that allows only one thread or process at a time to access a shared resource. The thread that locks it is responsible for unlocking it. Use a mutex to protect shared data or a critical section from simultaneous access. A semaphore maintains a count and lets a limited number of threads or processes access a resource at once; it can also signal that an event has occurred. Use a counting semaphore to manage a pool of limited resources, or a binary semaphore for signaling between threads.

Generated 2026-09-28. Answers vary from run to run.

GPT-6 Luna pricing

GPT-6 Luna is billed per token from your balance. Rates below are live, the same ones your account is charged.

One typical request2,000 input tokens, 1,000 output tokens
$0.0006per request
Reasoning counts as output, so higher effort raises the output side.
What $20.00 buysOutput tokens at the current rate
50,000,000output tokens
Input and cached tokens cost less per token than output.

Example sizes are for illustration. Your bill follows the tokens each request reports.

ItemCounted asRateBilling
InputPrompt tokens not read from or written to the cache$0.08 / 1MPer 1M tokens
OutputAnswer and reasoning tokens$0.40 / 1MPer 1M tokens
Cached inputPrompt tokens served from the prompt cache$0.008 / 1MPer 1M tokens
Cache writePrompt tokens written to the prompt cache$0.10 / 1MPer 1M tokens
Input, over 272KInput when the prompt exceeds 272K tokens$0.16 / 1MPer 1M tokens
Output, over 272KOutput when the prompt exceeds 272K tokens$0.60 / 1MPer 1M tokens
Cached input, over 272KCache reads when the prompt exceeds 272K tokens$0.016 / 1MPer 1M tokens
Cache write, over 272KCache writes when the prompt exceeds 272K tokens$0.20 / 1MPer 1M tokens
Failed requestAny request that returns an errorFreeNot charged

What is GPT-6 Luna?

GPT-6 Luna is OpenAI's most efficient GPT-6 model, built for focused, high-volume tasks. It is the lowest-cost model in the GPT-6 family, below GPT-6 Sol and GPT-6 Astra. GPT-6 Luna reads up to 1.05M tokens of context and writes up to 128K tokens per request.

GPT-6 Luna lets you choose how hard it thinks: reasoning effort runs from none to max, and the default is medium. At none it answers without reasoning and accepts temperature and top_p, which suits classification and extraction.

Model ID: gpt-6-luna · Input: text, images · Output: text · Formats: OpenAI Responses, OpenAI Chat Completions, Anthropic Messages.

Model ID
gpt-6-luna
Context window
1.05M tokens
Max output
128K tokens
Effort levels
none to max
Input
Text, images

Two ways to use GPT-6 Luna

Call it from your code, or let your coding agent do it for you.

API

Call it from your backend

Best for apps and services

Point the official OpenAI SDK at SeedRouter and keep your code. GPT-6 Luna answers both the Responses and Chat Completions formats.

  1. 1Create an API key
  2. 2Set the base URL to https://api.seedrouter.ai/v1
  3. 3Send gpt-6-luna as the model
Agent

Hand it to your coding agent

Best for Claude Code, Codex and other agents

Copy a ready prompt. Your agent shows the request and cost before anything is spent.

  1. 1Copy the prompt
  2. 2Paste it into your agent
  3. 3Approve the request

What GPT-6 Luna is good at

GPT-6 Luna is made for jobs you run thousands of times, where speed and cost per call matter.

Classification

GPT-6 Luna sorts tickets, reviews and messages into the labels you define.

Extraction

Pull names, dates and totals out of documents into clean fields.

Summaries

Turn long threads and reports into short, readable notes.

Images in the prompt

GPT-6 Luna reads screenshots and photos alongside your text.

Effort you control

None for the fastest answers, higher when a task needs more thought.

Cheaper repeat context

Cached input on GPT-6 Luna costs 10% of the input rate.

What teams build with GPT-6 Luna

Ticket triage

Label and route every support request with GPT-6 Luna before a human reads it.

Content tagging

Tag product listings, reviews or posts at catalogue scale.

Structured extraction

Turn invoices and forms into JSON with a text.format schema.

Start using GPT-6 Luna in four steps

  1. 01

    Create a key

    Sign in and create an API key. Add credit whenever you need it; there is no subscription.

  2. 02

    Change the base URL

    Point the OpenAI SDK at https://api.seedrouter.ai/v1. Your code stays the same.

  3. 03

    Pick GPT-6 Luna

    Send gpt-6-luna as the model. Keep reasoning effort low or none for the lowest cost.

  4. 04

    Check the bill

    Every request shows its tokens and charge in your usage records.

Get an API key

GPT-6 Luna on SeedRouter at a glance

What you can use today.

FeatureGPT-6 Luna
OpenAI ResponsesYes, official format
OpenAI Chat CompletionsYes
Anthropic MessagesYes
Image inputYes, by URL
StreamingYes
Prompt cachingYes, cached input and cache writes billed apart
Reasoning effort noneYes
Failed requestsNot charged

Batch, Flex, background mode and service tiers are not offered.

GPT-6 Luna limits

1.05M-token context

On GPT-6 Luna, prompt, images and history share one 1.05M-token window; input is capped at 922K.

128K tokens out

One GPT-6 Luna request can return up to 128K tokens, reasoning included.

Long prompts cost more

Above 272K input tokens, the whole GPT-6 Luna request is billed at the higher long-context rates.

Sampling needs effort none

GPT-6 Luna accepts temperature and top_p only when reasoning effort is none.

Why call GPT-6 Luna through SeedRouter

One key, three formats

Call GPT-6 Luna with OpenAI Responses, Chat Completions or Anthropic Messages, using the same key.

Pay per token

Top up once and spend it on any model. No plan, no monthly fee.

Failures are free

A request that errors is not charged.

Live prices

The rates on this page are the rates your account pays.

Prompt caching

Repeated prompt prefixes are read from the cache at a lower GPT-6 Luna rate.

Image input

Send screenshots and photos by URL with your prompt.

Streaming

Stream GPT-6 Luna tokens as they arrive with stream: true.

Related models

Other models you can call with the same key.

GPT-6 Astra
gpt-6-astra

Call GPT-6 Astra with the OpenAI SDK. OpenAI's most capable model, with a 1.05M-token context, live per-token pricing and a playground you can try right now.

View pricing
GPT-6 Sol
gpt-6-sol

Call GPT-6 Sol with the OpenAI SDK. 1.05M-token context, reasoning effort from none to max, live per-token pricing and a playground you can try right now.

View pricing
claude-fable-5
claude-fable-5

Anthropic

View pricing
claude-fable-5-1
claude-fable-5-1

Anthropic

View pricing
claude-opus-5-5
claude-opus-5-5

Anthropic

View pricing
dreamina-seedance-2-0
dreamina-seedance-2-0

ByteDance

View pricing

GPT-6 Luna FAQ

What is GPT-6 Luna?+

GPT-6 Luna is OpenAI's most efficient GPT-6 model, built for focused, high-volume tasks. It has a 1.05M-token context window, writes up to 128K tokens per request and lets you set reasoning effort from none to max.

Should I use GPT-6 Luna or GPT-6 Sol?+

Use GPT-6 Luna for classification, extraction, summaries and other repeatable jobs where cost matters most. Move to GPT-6 Sol for demanding coding and agent tasks, and to GPT-6 Astra for the hardest work.

How much does GPT-6 Luna cost?+

GPT-6 Luna is billed at $0.08 / 1M input and $0.40 / 1M output per 1M tokens, the lowest rates in the GPT-6 family. Cached input costs 10% of the input rate. Prompts above 272K tokens use the long-context rates.

What is the default reasoning effort of GPT-6 Luna?+

GPT-6 Luna defaults to medium. You can set none, low, medium, high, xhigh or max; none turns reasoning off and is the only level that accepts temperature and top_p.

What is the knowledge cutoff of GPT-6 Luna?+

GPT-6 Luna has a knowledge cutoff of May 18, 2026. For newer facts, pass them in the prompt.

What is the GPT-6 Luna model ID and how do I call it?+

The model ID is gpt-6-luna. Create a SeedRouter API key, point the OpenAI SDK at https://api.seedrouter.ai/v1 and send gpt-6-luna as the model with the Responses or Chat Completions API.

Can GPT-6 Luna read images?+

Yes. GPT-6 Luna accepts images alongside text, sent by URL in the request. It returns text only.

Am I charged if a GPT-6 Luna request fails?+

No. A GPT-6 Luna request that returns an error is not charged. You pay only for the tokens a finished request reports.

Try GPT-6 Luna now

Send your first GPT-6 Luna request from the playground or your own code in minutes.