OpenAI launched GPT-6.1 Sol on September 22, 2026 at $2 per million input tokens, one-fifth of GPT-6 Astra’s rates, and reports it matches Astra on DeepSWE v1.1 coding at that price. The practical answer for routing teams: move prefix-heavy agentic coding, document, and business-workflow agents to gpt-6.1-sol now, keep Astra for the hardest scientific research, and gate full migration on your own eval run, because every capability number so far comes from OpenAI itself.
What OpenAI claims Sol does at one-fifth the price
OpenAI’s launch post describes GPT-6.1 Sol as an upgrade to GPT-6 Sol that “nearly matches GPT-6 Astra’s intelligence on agentic coding, computer use, and professional work at one-fifth of Astra’s standard input and output token prices.” Standard API pricing is $2 per million input tokens, $0.10 per million cached input tokens, and $10 per million output tokens. Cached input sits 95% below standard input and 50% below GPT-6 Sol’s cached price.
The reported benchmark deltas, all vendor-reported:
- DeepSWE v1.1 (agentic coding): Sol matches Astra at roughly one-fifth the cost, and beats GPT-6 Sol’s best score by 6.4 percentage points at lower reasoning effort and cost.
- GDP.pdf (document QA over complex PDFs): Sol scores higher than Opus 5.5 with fallbacks at less than half the cost per task, and approaches Astra’s state-of-the-art at roughly one-fifth the cost per task.
- AutomationBench 1.0.6 (multi-step business workflows): Sol scores 2.2 percentage points above Opus 5.5 at medium reasoning effort, at roughly a third of the cost, up 4.8 points from GPT-6 Sol at the same setting.
- OSWorld 2.0 offline set (computer use): Sol comes within 2.1 percentage points of Astra at maximum reasoning effort at roughly one-seventh the cost per task.
- Terminal-Bench Science 0.1 (scientific research): Sol costs $5.47 per task on average, versus $23.21 for Opus 5.5 and $23.80 for Astra, but Astra still posts the highest score at 68.1%.
None of these numbers has an independent reproduction yet. That does not make them useless; vendor evals are usually directionally informative. It does mean the correct posture is the one Groundy applied when Microsoft made Copilot model changes without documentation: treat the claims as vendor-asserted until independent evals or your own runs confirm them.
What the three headline evals actually measure
The routing decision depends on what each benchmark rewards, and the independent papers behind two of them complicate the headline deltas.
DeepSWE is the strongest signal in the table. The independent DeepSWE paper describes 113 original long-horizon software engineering tasks run under a fixed harness: every model uses mini-swe-agent with a single bash tool and a shared prompt, so the leaderboard reflects model capability rather than per-vendor scaffolding. That design matters because the same paper warns that harness choice is itself a confound for cross-model leaderboards, and its confidence intervals caution against reading close scores as strict rankings. Two caveats attach. First, the paper describes v1 (May 2026) while OpenAI cites v1.1, so task composition may differ. Second, “matches Astra” under a confidence-interval regime means the scores are statistically indistinguishable, which supports the routing argument but not a claim of equivalence in every repository.
GDP.pdf measures how accurately models answer professional questions over complex PDFs, including tables, charts, diagrams, and fine print. There is no independent paper in the evidence describing its construction, so the “higher than Opus 5.5 with fallbacks at less than half the cost per task” claim rests entirely on OpenAI’s post. Directionally it fits the pattern (Sol is strong on professional document work), but it is the least externally anchored row in the table.
AutomationBench grades cross-application REST workflow orchestration, end-state only. OpenAI’s version 1.0.6 tested agents on end-to-end workflows using 47 tools across sales, marketing, operations, support, finance, and HR. The independent AutomationBench paper (the lineage OpenAI’s 1.0.6 descends from) reports that state-of-the-art models all score below 10%, with Opus 4.7 topping the leaderboard at 9.9%. That context changes how the 2.2-point delta reads: it is a large relative gap in a regime where no model reliably completes these workflows. An agent that fails roughly nine of ten cross-app business tasks is not yet a default for unattended automation regardless of which model wins the comparison. The same paper also documents that cost per task is driven by token volume, not headline price: Sonnet 4.6 cost more per task than Opus 4.7 at every reasoning effort despite lower pass rates. And OpenAI itself flags that its Claude Fable 5.1 cost datapoint understates actual cost because it omits fallbacks, which occurred on roughly 40% of tasks. Any fallback-based routing comparison inherits that measurement problem.
Where Astra still wins
OpenAI’s own post concedes the ceiling: “GPT-6 Astra still achieves the highest score among the models tested at 68.1%, and should be used for the most difficult scientific research tasks.” That is Terminal-Bench Science 0.1, and it is the one eval where the cost-per-task math does not rescue Sol, because a wrong answer at $5.47 is still wrong.
Independent evidence supports treating Astra as a distinct tier on the hardest problems. The RoboDojo evaluation measured GPT-6 Astra at a 22.48% average success rate and 28.97 Score over 2,100 trials, ranking above all 40 public robot policies, though the authors flag single-seed evaluation as a limitation. A September 2026 exposition also attributes the short counting proof of the Erdős, Sós conjecture to a pre-release GPT-6 Astra, cited to the FrontierMath Erdős report. One attributed proof is not a benchmark, but it is a concrete instance of the research-tier capability OpenAI says Sol does not match.
Sol does close most of the gap on OSWorld 2.0 computer use (within 2.1 points of Astra at one-seventh the cost) and on OpenAI’s factuality check, where it reduces the share of responses containing a factual error from 11.4% to 7.7% at low reasoning effort and stays within 1.9 points of Astra across tested settings. OpenAI cautions that the factuality prompts come from deliberately difficult flagged conversations and are not representative of typical usage, so read that improvement as a stress-test result, not a prediction for production traffic.
A routing table keyed to what each eval measures
| Workload | Signal (all vendor-reported) | Reported cost ratio | Route to |
|---|---|---|---|
| Agentic coding in real codebases | Sol matches Astra on DeepSWE v1.1 | ~1/5 of Astra | Sol |
| PDF/document QA, tables, fine print | Sol beats Opus 5.5 w/ fallbacks; approaches Astra | <1/2 vs Opus, ~1/5 vs Astra | Sol |
| Cross-app business workflows | Sol +2.2 pts over Opus 5.5; SOTA regime under 10% | ~1/3 vs Opus | Sol, with human review gates |
| Computer-use automation | Sol within 2.1 pts of Astra | ~1/7 vs Astra | Sol for volume; Astra if the 2.1 pts matter |
| Hardest scientific research | Astra leads at 68.1% on Terminal-Bench Science 0.1 | Astra $23.80 vs Sol $5.47 per task | Astra |
| Open-ended chat | Sol not available in Chat | n/a | Whatever you use today |
The $0.10 cached-input line is the real story for agent loops
Token price cuts matter most where token volume is structural, and agentic loops have exactly that shape: every step resends the system prompt, tool definitions, and accumulated context, with only a small delta of new information. A 95% discount on cached input attacks the dominant cost center of always-on agents directly.
The economics are not automatic, though. The SoL-Pi harness paper (independent, single research group, measured on a GPT-5.6 Sol backend rather than GPT-6.1 Sol, so treat its dollar figures as structural rather than current) documents three things that belong in your model of cache economics. Cached reads still bill. Shortening context can reduce prompt-cache reuse when it changes a previously cached prefix. And preserving a long prefix is not always the cheapest choice over an entire task. Their full stack cut cache-read traffic from 2.1326B to 1.0605B tokens while cache-write traffic rose from 0.0141B to 0.0316B, and total model cost still fell from $1,339 to $894. The lesson transfers even though the backend prices do not: cache-hit rate is a tuning target, and sometimes trading reuse for less repeated input wins.
A complementary lever appears in production agent design: Fin-Analyst’s pipeline caches specialist verdicts by MD5 hash of the prompt, so unchanged inputs reuse a prior verdict instead of issuing a new API call at all. Provider prompt caching and application-level result caching stack.
A cost-per-task worksheet for agents that resend context
Token price is not cost per task, as the AutomationBench paper’s Sonnet-versus-Opus inversion shows. Compute it for your loop:
Cost per task = (cache-write tokens × $2) + (cache-read tokens × $0.10) + (uncached input × $2) + (output tokens × $10), all per million.
A hypothetical example, flagged as such: a 40-step coding agent with a stable 150k-token prefix, 2k new input tokens per step, and 4k output tokens per step. Step one writes the prefix at $0.30. The remaining 39 steps read it at $0.10/M for $0.585. New input adds $0.16 and output adds $1.60. Total: roughly $2.65. Without caching, the same loop re-bills 6.08M input tokens at list price for about $13.76. Note where the money sits in the cached version: output tokens at $10/M become the largest line. Once caching works, the next optimization is shortening agent outputs, not further prefix engineering.
Two checks before trusting that arithmetic in production. Measure your actual cache-hit rate, because SoL-Pi’s data shows context compaction and prefix changes quietly destroy reuse. And measure output token volume per task, because Sol’s reported cost ratios (one-fifth, one-seventh) are per-benchmark settings with particular reasoning efforts, not promises about your workload.
What you can actually use today
Sol is available to Plus, Pro, Business, Enterprise, and Edu users in ChatGPT Work and Codex, and through the OpenAI API as gpt-6.1-sol. It is not available in Chat. OpenAI announced a Sol Ultrafast variant with up to 8x faster token generation in Codex as coming in the following days; it was not shipping when this was checked. OpenAI’s own ChatGPT documentation now steers users toward Sol for complex work at lower cost than Astra, which tells you how the vendor wants the default routed.
If your agents run inside an editor, remember that the editor’s model routing is a layer you do not control: Cursor’s routing is the vendor’s layer, so verify which model ID actually serves your traffic rather than trusting a UI label. The same applies to any gateway or aggregator: confirm gpt-6.1-sol is listed and billed as such before rerouting production load.
A verification checklist before you migrate
Every capability number in this article is vendor-reported, and the evidence includes a template for why that matters: a 2026 independent reproduction of LeWorldModel found the evaluation protocol described in the paper failed to reproduce the paper’s own reported result on the paper’s own released weights, while the repository’s evaluation defaults did. Protocol choices determined the headline. Before moving production traffic to Sol:
- Replay 50-100 representative tasks from your own distribution through both Sol and your current default, with the harness held fixed, exactly as DeepSWE holds mini-swe-agent fixed across models.
- Report confidence intervals, not point scores. DeepSWE’s own authors caution against strict rankings on close results; a 2-point Sol deficit inside the interval is parity for routing purposes.
- Log cost per completed task, including fallbacks and retries. OpenAI’s own Fable 5.1 caveat (fallbacks on ~40% of tasks) shows how omission flatters a comparison.
- Instrument cache-hit rate and output token volume before and after any context-engineering change.
- Keep Astra in the routing table for scientific and research-tier tasks, where OpenAI itself says to use it.
This is the same posture that makes cheap-specialist routing work elsewhere: as the Terminus-4B analysis showed, a cheap model absorbing one well-measured category of work is a sound bet precisely because the measurement is narrow and repeatable.
Routing verdict
Move prefix-heavy agentic coding, PDF and document QA, and multi-step business-workflow agents from Astra or Opus 5.5 fallbacks to gpt-6.1-sol, where OpenAI’s reported numbers show near-parity at one-fifth the token prices and cached input at $0.10 per million tokens plausibly restructures the economics of loops that resend long shared context every step. Keep Astra for the hardest scientific research, where its 68.1% Terminal-Bench Science lead is the one gap OpenAI itself calls decisive. The strategic shift is from strongest-model-by-default to eval-matched tiering, and it survives even if Sol’s reported parity softens under independent testing, because the cost structure does not depend on the last benchmark point.
The strongest limitation stands: no source outside OpenAI has reproduced a single GPT-6.1 Sol score as of September 30, 2026, parity is benchmark-specific, the DeepSWE version OpenAI cites (v1.1) postdates the independent paper (v1), and AutomationBench’s deltas live in a regime where every tested model fails most tasks. Treat the migration as a measured experiment with your own eval gate, not a settled conclusion.
Frequently Asked Questions
Where is GPT-6 Astra still recommended for use?
GPT-6 Astra still achieves the highest score among the models tested at 68.1%, and should be used for the most difficult scientific research tasks.
Is GPT-6.1 Sol available in ChatGPT Chat?
Sol is available to Plus, Pro, Business, Enterprise, and Edu users in ChatGPT Work and Codex, and through the OpenAI API as gpt-6.1-sol. It is not available in Chat.

Join the discussion
Share a useful perspective or ask a question about this article.