Claude Opus 5.5 is live on SeedRouter

GPT-6.1 Sol reasoning effort: low, medium, high, xhigh or max?

Which GPT-6.1 Sol reasoning effort to use: real timings, token counts and costs at low, medium, high, xhigh and max, why it feels slow, and where to start.

Read as Markdown

Start GPT-6.1 Sol at medium, its default, and use low when you want the answer quickly. Raise the effort to high, xhigh or max only for tasks that fail at medium: long refactors, hard debugging, multi-step agent work. On the short tasks in the tests below, every effort level gave the right answer, but max took about ten times as long to show its first word and used two and a half to five times the tokens of low.

That is also the main reason GPT-6.1 Sol feels slow. It always reasons before it answers, and you wait for that reasoning before the first visible text.

Which reasoning efforts does GPT-6.1 Sol support?

Five: low, medium, high, xhigh and max. The default is medium. According to OpenAI's GPT-6.1 Sol model page, "the none and minimal reasoning efforts are not supported", so unlike GPT-6 Sol there is no way to switch reasoning off. Without none, sampling settings such as temperature and top_p have no effect either.

Set the effort per request. In the Responses API it is reasoning.effort:

from openai import OpenAI

client = OpenAI(api_key="YOUR_SEEDROUTER_KEY", base_url="https://api.seedrouter.ai/v1")

response = client.responses.create(
    model="gpt-6.1-sol",
    input="Find the race condition in this handler and suggest a fix: ...",
    reasoning={"effort": "low"},
)
print(response.output_text)
print(response.usage.output_tokens_details.reasoning_tokens)

In Chat Completions the field is reasoning_effort. Reasoning tokens are counted in output_tokens, billed as output, and share the max_output_tokens budget with the answer.

How do the five efforts compare in practice?

The test sent three prompts with known answers to GPT-6.1 Sol through SeedRouter on October 5, 2026, once per effort level, all five levels at the same time. Times are seconds from request to complete answer. Cost is at OpenAI's list prices of $2 input and $10 output per million tokens, checked on October 5, 2026.

Counting problem — how many integers from 1 to 1000 are divisible by 3 or 5 but not by 7 (answer: 401).

EffortSecondsReasoning tokensOutput tokensCostCorrect
low14.051225$0.0023Yes
medium13.446187$0.0020Yes
high14.6134272$0.0028Yes
xhigh16.2303362$0.0037Yes
max20.4516574$0.0058Yes

Harder counting problem — how many integers from 1 to 99,999 are multiples of 7 and contain no digit 7 (answer: 8,434).

EffortSecondsReasoning tokensOutput tokensCostCorrect
low17.5473711$0.0072Yes
medium27.61,0311,273$0.0128Yes
high28.31,0341,255$0.0126Yes
xhigh31.21,5521,763$0.0177Yes
max42.62,0702,259$0.0227Yes

Code review — find the bugs in a four-line Python median function that sorts its input in place and mishandles even-length lists.

EffortSecondsReasoning tokensOutput tokensCostFound both bugs
low14.976221$0.0023Yes, with a fixed version
medium15.491194$0.0020Yes
high19.5296396$0.0040Yes
xhigh19.5516606$0.0061Yes
max28.41,0341,114$0.0112Yes

These are single runs on short tasks, so treat them as a sense of scale, not a benchmark. Reasoning-token counts vary between runs, and harder tasks are where higher effort earns its cost.

Why is GPT-6.1 Sol slow?

Most of the wait is reasoning, and reasoning comes first. A second test streamed the same 300-word explanation at three effort levels and measured when the first visible text arrived:

EffortFirst text afterTotal timeReasoning tokensAnswer speed
low3.0 s20.5 s5625 tokens/s
medium13.6 s24.7 s82440 tokens/s
max30.6 s43.5 s2,51734 tokens/s

Once the answer starts, it streams at a similar pace at every level. What changes is the silent part before it: at max, the model spent half a minute reasoning before writing anything. If your users watch a spinner, low with streaming is the setting that feels fastest.

OpenAI has also announced GPT-6.1 Sol Ultrafast, with "up to 8x faster token generation compared to its standard speed in Codex", in its launch post.

Can xhigh or max be worse than medium?

On a given task, yes: higher effort is not a guarantee of a better answer. In the runs above, max never beat low on correctness, it only cost more. OpenAI's own results point the same way for routine work: GPT-6.1 Sol beat GPT-6 Sol's best DeepSWE score "at a lower reasoning effort", and its largest factuality gain over GPT-6 Sol came at low effort.

Where higher effort pays off in OpenAI's numbers is long, hard work: its OSWorld 2.0 and Terminal-Bench Science results for GPT-6.1 Sol are reported at maximum effort. The practical rule is to measure on your own tasks. Run a sample at medium and at one level higher, and keep the higher level only where it fixes failures.

Which effort should you use?

  • low — chat replies, classification, extraction, simple code edits, anything a user waits for.
  • medium — the default for coding help, reviews and most agent steps.
  • high — multi-file changes and debugging that fails at medium.
  • xhigh — long refactors and planning across many steps.
  • max — the hardest problems, where a wrong answer costs more than the extra tokens.

For a whole agent, mix levels: plan at a higher effort, run routine tool steps at low or medium. Prices for every level are on the GPT-6.1 Sol page, and the GPT-6.1 Sol pricing guide explains how reasoning tokens show up on the bill.

Frequently asked questions

What is the default reasoning effort of GPT-6.1 Sol?

medium. If you leave the field out, GPT-6.1 Sol reasons at medium.

Should I use low or medium for GPT-6.1 Sol?

Use medium as the default and low where speed matters more than depth. In the test above, low answered every question correctly and showed its first text in about three seconds, against about fourteen at medium.

Why is GPT-6.1 Sol so slow?

It reasons before it answers, and the reasoning is not streamed as text. At medium and above, that silent phase takes most of the wait. Lower the effort or stream the response to start showing text sooner.

Is xhigh better than high on GPT-6.1 Sol?

Only on tasks that need it. On short tasks both levels gave the same answers, and xhigh used more tokens. Test both on your own work before paying for xhigh.

Can I turn off reasoning in GPT-6.1 Sol?

No. GPT-6.1 Sol does not support none or minimal. If you need answers without reasoning, GPT-6 Sol still supports none. See the GPT-6.1 Sol vs GPT-6 Astra vs GPT-6 Sol comparison.

Related guides