GPT-6 Luna — low-cost GPT-6 for high-volume work
GPT-6 Luna is OpenAI's most efficient GPT-6 model, made for focused tasks at scale, with a 1.05M-token context. $0.08 / 1M input, $0.40 / 1M output per 1M tokens.
- Reasoning effort control
- 1.05M-token context
- Image input
- Prompt caching
- Streaming
Try GPT-6 Luna
Chat with GPT-6 Luna, attach images, and see the tokens and cost of every answer.
gpt-6-luna$0.08 in · $0.40 out per 1M tokens
Temperature and Top P work only with reasoning effort none.
Try one of these prompts
Messages are billed from your balance at the rates below, like an API call. Failed requests are free.
Real GPT-6 Luna output
One prompt, sent to GPT-6 Luna at low reasoning effort. Unedited.
Explain the difference between a mutex and a semaphore in two short paragraphs, and say when to use each.
A mutex is a lock that allows only one thread or process at a time to access a shared resource. The thread that locks it is responsible for unlocking it. Use a mutex to protect shared data or a critical section from simultaneous access. A semaphore maintains a count and lets a limited number of threads or processes access a resource at once; it can also signal that an event has occurred. Use a counting semaphore to manage a pool of limited resources, or a binary semaphore for signaling between threads.
Generated 2026-09-28. Answers vary from run to run.
GPT-6 Luna pricing
GPT-6 Luna is billed per token from your balance. Rates below are live, the same ones your account is charged.
Example sizes are for illustration. Your bill follows the tokens each request reports.
| Item | Counted as | Rate | Billing |
|---|---|---|---|
| Input | Prompt tokens not read from or written to the cache | $0.08 / 1M | Per 1M tokens |
| Output | Answer and reasoning tokens | $0.40 / 1M | Per 1M tokens |
| Cached input | Prompt tokens served from the prompt cache | $0.008 / 1M | Per 1M tokens |
| Cache write | Prompt tokens written to the prompt cache | $0.10 / 1M | Per 1M tokens |
| Input, over 272K | Input when the prompt exceeds 272K tokens | $0.16 / 1M | Per 1M tokens |
| Output, over 272K | Output when the prompt exceeds 272K tokens | $0.60 / 1M | Per 1M tokens |
| Cached input, over 272K | Cache reads when the prompt exceeds 272K tokens | $0.016 / 1M | Per 1M tokens |
| Cache write, over 272K | Cache writes when the prompt exceeds 272K tokens | $0.20 / 1M | Per 1M tokens |
| Failed request | Any request that returns an error | Free | Not charged |
What is GPT-6 Luna?
GPT-6 Luna is OpenAI's most efficient GPT-6 model, built for focused, high-volume tasks. It is the lowest-cost model in the GPT-6 family, below GPT-6 Sol and GPT-6 Astra. GPT-6 Luna reads up to 1.05M tokens of context and writes up to 128K tokens per request.
GPT-6 Luna lets you choose how hard it thinks: reasoning effort runs from none to max, and the default is medium. At none it answers without reasoning and accepts temperature and top_p, which suits classification and extraction.
Model ID: gpt-6-luna · Input: text, images · Output: text · Formats: OpenAI Responses, OpenAI Chat Completions, Anthropic Messages.
Two ways to use GPT-6 Luna
Call it from your code, or let your coding agent do it for you.
Call it from your backend
Point the official OpenAI SDK at SeedRouter and keep your code. GPT-6 Luna answers both the Responses and Chat Completions formats.
- 1Create an API key
- 2Set the base URL to https://api.seedrouter.ai/v1
- 3Send gpt-6-luna as the model
Hand it to your coding agent
Copy a ready prompt. Your agent shows the request and cost before anything is spent.
- 1Copy the prompt
- 2Paste it into your agent
- 3Approve the request
What GPT-6 Luna is good at
GPT-6 Luna is made for jobs you run thousands of times, where speed and cost per call matter.
Classification
GPT-6 Luna sorts tickets, reviews and messages into the labels you define.
Extraction
Pull names, dates and totals out of documents into clean fields.
Summaries
Turn long threads and reports into short, readable notes.
Images in the prompt
GPT-6 Luna reads screenshots and photos alongside your text.
Effort you control
None for the fastest answers, higher when a task needs more thought.
Cheaper repeat context
Cached input on GPT-6 Luna costs 10% of the input rate.
What teams build with GPT-6 Luna
Ticket triage
Label and route every support request with GPT-6 Luna before a human reads it.
Content tagging
Tag product listings, reviews or posts at catalogue scale.
Structured extraction
Turn invoices and forms into JSON with a text.format schema.
Start using GPT-6 Luna in four steps
- 01
Create a key
Sign in and create an API key. Add credit whenever you need it; there is no subscription.
- 02
Change the base URL
Point the OpenAI SDK at https://api.seedrouter.ai/v1. Your code stays the same.
- 03
Pick GPT-6 Luna
Send gpt-6-luna as the model. Keep reasoning effort low or none for the lowest cost.
- 04
Check the bill
Every request shows its tokens and charge in your usage records.
GPT-6 Luna on SeedRouter at a glance
What you can use today.
| Feature | GPT-6 Luna |
|---|---|
| OpenAI Responses | Yes, official format |
| OpenAI Chat Completions | Yes |
| Anthropic Messages | Yes |
| Image input | Yes, by URL |
| Streaming | Yes |
| Prompt caching | Yes, cached input and cache writes billed apart |
| Reasoning effort none | Yes |
| Failed requests | Not charged |
Batch, Flex, background mode and service tiers are not offered.
GPT-6 Luna limits
1.05M-token context
On GPT-6 Luna, prompt, images and history share one 1.05M-token window; input is capped at 922K.
128K tokens out
One GPT-6 Luna request can return up to 128K tokens, reasoning included.
Long prompts cost more
Above 272K input tokens, the whole GPT-6 Luna request is billed at the higher long-context rates.
Sampling needs effort none
GPT-6 Luna accepts temperature and top_p only when reasoning effort is none.
Why call GPT-6 Luna through SeedRouter
One key, three formats
Call GPT-6 Luna with OpenAI Responses, Chat Completions or Anthropic Messages, using the same key.
Pay per token
Top up once and spend it on any model. No plan, no monthly fee.
Failures are free
A request that errors is not charged.
Live prices
The rates on this page are the rates your account pays.
Prompt caching
Repeated prompt prefixes are read from the cache at a lower GPT-6 Luna rate.
Image input
Send screenshots and photos by URL with your prompt.
Streaming
Stream GPT-6 Luna tokens as they arrive with stream: true.
Related models
Other models you can call with the same key.
Call GPT-6 Astra with the OpenAI SDK. OpenAI's most capable model, with a 1.05M-token context, live per-token pricing and a playground you can try right now.
View pricingCall GPT-6 Sol with the OpenAI SDK. 1.05M-token context, reasoning effort from none to max, live per-token pricing and a playground you can try right now.
View pricingGPT-6 Luna guides
Walkthroughs and comparisons for this model.
How to access the GPT-6 API: get a key, call GPT-6 Astra, Sol or Luna with the OpenAI SDK, set reasoning effort, stream, send images and avoid common errors.
GPT-6 API pricing for Astra, Sol and Luna per million tokens: input, output, cached input, cache writes and the 272K long-context rate, with worked examples.
GPT-6 Astra vs Sol vs Luna compared: specs, prices, OpenAI benchmark results, reasoning effort and a clear pick for coding, agents and high-volume work.
GPT-6 Astra, Sol and Luna vs Claude Opus 5.5: specs, token prices, the benchmark results both labs published, and which model fits coding, agents and cost.
GPT-6 vs GPT-5.6 compared: Sol and Luna at half the price, the same 1.05M context, newer knowledge, OpenAI benchmark gains and what to change when you migrate.
GPT-6 Luna FAQ
What is GPT-6 Luna?+
GPT-6 Luna is OpenAI's most efficient GPT-6 model, built for focused, high-volume tasks. It has a 1.05M-token context window, writes up to 128K tokens per request and lets you set reasoning effort from none to max.
Should I use GPT-6 Luna or GPT-6 Sol?+
Use GPT-6 Luna for classification, extraction, summaries and other repeatable jobs where cost matters most. Move to GPT-6 Sol for demanding coding and agent tasks, and to GPT-6 Astra for the hardest work.
How much does GPT-6 Luna cost?+
GPT-6 Luna is billed at $0.08 / 1M input and $0.40 / 1M output per 1M tokens, the lowest rates in the GPT-6 family. Cached input costs 10% of the input rate. Prompts above 272K tokens use the long-context rates.
What is the default reasoning effort of GPT-6 Luna?+
GPT-6 Luna defaults to medium. You can set none, low, medium, high, xhigh or max; none turns reasoning off and is the only level that accepts temperature and top_p.
What is the knowledge cutoff of GPT-6 Luna?+
GPT-6 Luna has a knowledge cutoff of May 18, 2026. For newer facts, pass them in the prompt.
What is the GPT-6 Luna model ID and how do I call it?+
The model ID is gpt-6-luna. Create a SeedRouter API key, point the OpenAI SDK at https://api.seedrouter.ai/v1 and send gpt-6-luna as the model with the Responses or Chat Completions API.
Can GPT-6 Luna read images?+
Yes. GPT-6 Luna accepts images alongside text, sent by URL in the request. It returns text only.
Am I charged if a GPT-6 Luna request fails?+
No. A GPT-6 Luna request that returns an error is not charged. You pay only for the tokens a finished request reports.
Try GPT-6 Luna now
Send your first GPT-6 Luna request from the playground or your own code in minutes.
