Models

OpenAI Launches GPT-6.1 Sol at One-Fifth of Astra's Price

GPT-6.1 Sol nearly matches GPT-6 Astra on coding, computer use and research benchmarks at one-fifth the price, with cached input at $0.10 per million tokens.

Introducing GPT-6.1 Sol
Introducing GPT-6.1 SolAI-generated
By Marcus Bennett5 min read

Updated

Why it matters

  • GPT-6.1 Sol costs one-fifth of GPT-6 Astra's standard input and output token prices, with cached input at $0.10 per million tokens.
  • On Terminal-Bench Science 0.1 at maximum reasoning effort, GPT-6.1 Sol averages $5.47 per task versus $23.21 for Opus 5.5 and $23.80 for Astra.
  • Available today in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu users, and via the API as gpt-6.1-sol; an Ultrafast variant with up to 8x faster token generation arrives in the coming days.

OpenAI has released GPT-6.1 Sol, an upgrade to GPT-6 Sol that the company says nearly matches GPT-6 Astra's intelligence on agentic coding, computer use, and professional work—at one-fifth of Astra's standard input and output token prices. The announcement lands as competition over the price-to-performance ratio of frontier-adjacent models intensifies, with enterprises weighing per-task costs as heavily as benchmark scores.

The pricing structure is the sharpest signal of intent. Cached input costs $0.10 per million tokens, which OpenAI describes as 95% less than its standard input pricing and 50% less than GPT-6 Sol's cached input rate. The company frames this as giving "developers more room to build and run capable agents that reuse context across requests"—a direct pitch at agent builders, where repeated context injection drives most of the bill. Standard API pricing sits at $2 per million input tokens and $10 per million output tokens.

Coding benchmarks

On DeepSWE v1.1, a benchmark that evaluates complex software-engineering tasks in real codebases, GPT-6.1 Sol matches GPT-6 Astra at roughly one-fifth of the cost. It also beats GPT-6 Sol's best score by 6.4 percentage points while running at a lower reasoning effort and cost.

Professional work

The model makes its case on document-heavy and workflow-heavy tasks too. On GDP.pdf, which measures how accurately models answer professional questions using complex PDFs—including tables, charts, diagrams and fine print—GPT-6.1 Sol scores higher than Opus 5.5 with fallbacks at less than half the cost per task across tested reasoning settings. It approaches GPT-6 Astra's state-of-the-art performance at roughly one-fifth the cost per task. The benchmark draws prompts from professional workflows in finance, healthcare, legal, and seven other domains.

On AutomationBench, which tests whether agents correctly complete multi-step business workflows using 47 tools across sales, marketing, operations, support, finance and HR, GPT-6.1 Sol scores 2.2 percentage points above Opus 5.5 at medium reasoning effort—at roughly a third of the cost. That result is up 4.8 percentage points from GPT-6 Sol at the same setting.

OpenAI notes a caveat on competitor comparisons: the datapoint for Claude Fable 5.1 understates its actual cost because it omits the cost of fallbacks, which occurred on roughly 40% of tasks.

Computer use and research

On OSWorld 2.0's offline set, which evaluates agents on long-horizon computer-use workflows spanning everyday and professional tasks, GPT-6.1 Sol outperforms GPT-6 Sol by seven percentage points at maximum reasoning effort, at less than half the cost. It comes within 2.1 percentage points of Astra's score at maximum reasoning effort at roughly one-seventh the cost per task. OpenAI reports the partial reward on the offline set from the v2026.08.08 release.

The cost gap is widest on scientific work. On Terminal-Bench Science 0.1, which evaluates scientific workflows including data analysis, simulation and theorem proving, GPT-6.1 Sol more than doubles GPT-6 Sol's score at maximum reasoning effort while costing less than half as much per task. At maximum effort, GPT-6.1 Sol costs $5.47 per task on average, compared with $23.21 for Opus 5.5 and $23.80 for Astra—what OpenAI calls "substantial scientific capability at over 75% lower cost than either model."

GPT-6 Astra still holds the highest score among models tested on that benchmark at 68.1%, and OpenAI says it "should be used for the most difficult scientific research tasks."

Factuality

GPT-6.1 Sol also reduces factual errors on deliberately difficult prompts. Its largest improvement over GPT-6 Sol comes at low reasoning effort, where the share of responses containing a factual error drops from 11.4% to 7.7%—a reduction of approximately 32%. Across tested reasoning settings, its error rate stays within 1.9 percentage points of GPT-6 Astra's, at less than one-fifth the cost per task. The evaluation measures answers containing at least one factual error on de-identified ChatGPT conversations where users flagged an earlier model's mistake; OpenAI cautions these error-inducing prompts are not representative of typical usage, where factual errors are rarer.

Safety evaluations

On alignment, OpenAI reports that GPT-6.1 Sol "shows substantial improvements over GPT-6 Sol in our alignment evaluations, bringing it closer to GPT-6 Astra." The model is more transparent about its limitations and more reliable at respecting user intent and safety constraints, according to the company. In challenging evaluations, it posts lower failure rates than GPT-6 Sol on transparency about broken search tools, respecting explicit restrictions, and avoiding unauthorized outcomes during agentic tasks. OpenAI says it "observed no attempts to bypass an automated safety reviewer, matching GPT-6 Astra and GPT-6 Sol."

One disclosed test checks whether agents tell users when their search tool is broken instead of offering a best guess. GPT-6.1 Sol fails to disclose the problem in 2.1% of cases, compared with 4.9% for GPT-6 Sol, 1.5% for GPT-6 Astra, and 28.7% for GPT-6 Luna. OpenAI notes the tasks are selected to elicit failures and do not represent typical usage; effort was set to maximum. Full details appear in the GPT-6.1 Sol system card addendum on OpenAI's deployment safety site.

Availability

GPT-6.1 Sol is available starting today to all Plus, Pro, Business, Enterprise and Edu users in ChatGPT Work and Codex. It is not yet available in Chat. Developers can access it through the OpenAI API as gpt-6.1-sol. In the coming days, OpenAI will also offer GPT-6.1 Sol Ultrafast, with up to 8x faster token generation compared to standard speed in Codex.

The release tightens the squeeze on mid-tier model pricing: if a one-fifth-cost model now sits within a few points of the flagship on most agentic benchmarks, the burden shifts to Astra-class models to justify their premium on the hardest tasks alone.

Original: surgehq.ai

Share this article:

More from Marcus Bennett

Marcus Bennett

Show full bio

Senior reporter covering consumer brands and retail at AI In Context.

126 articles

Related articles

  1. OpenAI Ships GPT-6.1 Sol at a Fifth of Astra's Price
  2. OpenAI's WebSocket Overhaul Makes Agents 40% Faster
  3. OpenAI's GPT-6.1 Sol Nears Flagship Astra at One-Fifth the Price
  4. OpenAI Launches GPT-6 Sol and Luna at Half the Price
  5. OpenAI Ships GPT-5.1: Faster Reasoning, Better Coding, Same Price

« Previous article