
GPT-6.1 Sol’s $2 Tokens Come With a Benchmark-Shaped Asterisk
Every time I see “the same intelligence for less,” I look for the footnote. Years of covering technology pricing have taught me that the discount usually arrives before the definition of “same.” OpenAI’s GPT-6.1 Sol announcement follows that tradition, although there’s something genuinely useful underneath the confetti.
Released at DevDay on September 29, just seven days after GPT-6 Sol, the new model targets coding and agentic workflows. Another week, another migration review. Somewhere, an enterprise validation team is quietly screaming.
The fifth-price claim survives arithmetic
Standard API pricing is straightforward: $2 per million input tokens and $10 per million output tokens, against GPT-6 Astra’s $10 and $50. That is one-fifth. No interpretive dance required.
But those ordinary rates are unchanged from GPT-6 Sol. The upgrade is better reported performance at the existing price, plus cached input falling from $0.20 to $0.10 per million tokens.
<> One-fifth the token price is not a promise of one-fifth the application bill./>
OpenAI’s own documentation supplies the less photogenic details:
- Cache writes cost $2.50 per million tokens.
- Requests above 272,000 input tokens incur double input/cache rates and 1.5× output rates for the entire request.
- Fast mode costs 2× Standard; Batch and Flex cost 50% less.
- Regional processing adds a 10% premium where available.
A million-token context window sounds magnificent until someone fills it and discovers the pricing threshold. Capacity and affordability are different features.
“Near-Astra” has a postcode
OpenAI reports that Sol matches Astra on DeepSWE v1.1 at roughly one-fifth the task cost, improving on its predecessor’s best score by 6.4 percentage points. On AutomationBench, medium-effort Sol beats Anthropic’s Opus 5.5 by 2.2 points at about one-third the cost.
That is worth investigating. Seriously.
But these are vendor-reported evaluations, with specific effort settings and environments. OpenAI also notes differences between research/API evaluations and production ChatGPT, while competitor figures come from public reports. This is not a uniformly controlled cage match.
On Terminal-Bench Science, maximum-effort Sol averages $5.47 per task, versus $23.80 for Astra. Astra still leads the scoring at 68.1%, and OpenAI still recommends it for the hardest scientific research.
So the sensible interpretation is “potentially excellent worker model,” not “premium intelligence has become a commodity everywhere.”
Hacker News commenter MintsJohn asks for evidence about usable context, language coverage, and respect for existing code architecture. Those are better purchasing questions than whether a marketing chart labels something intelligent. A model that rewrites your architecture unnecessarily has not saved you money.
The migration checklist beats the leaderboard
The API identifier is gpt-6.1-sol. Tool calling requires the Responses API; Chat Completions works without tools. There’s no fine-tuning, and Fast mode is unavailable with EU data residency. These constraints matter more than another decimal place on a benchmark.
My evaluation order would be:
1. Replay representative repository and workflow tasks, including ugly failures.
2. Measure cost per accepted result, counting retries, tools, review time, and latency.
3. Route difficult cases to Astra rather than making every task pay premium rent.
Keep permissions narrow. Lower prices do not soften security consequences.
OpenAI rates Sol Critical in cybersecurity and High in biological and chemical capability under its own framework. Separately, it halted the planned GPT-6.1 Astra rollout over safety concerns. Cheaper deployment and withholding a more advanced model are distinct decisions—not proof that either model is harmless.
My Bet
Sol becomes the default worker for many teams that validate OpenAI’s results on their own workloads. Astra becomes an escalation path, not the everyday hammer. The winners will be developers who build reliable routing, evaluations, and permission boundaries—not whoever boasts the largest context window. Cheaper tokens help. Cheaper mistakes would help more.

