GPT-6.1 Sol Matches Astra-Class Coding at One-Fifth the Price
OpenAI's GPT-6.1 Sol, released September 29, 2026, matches Astra on DeepSWE v1.1 coding at one-fifth the cost, with $0.10 per million cached input tokens.
Updated
Why it matters
- OpenAI released GPT-6.1 Sol on September 29, 2026 at $2 input / $10 output / $0.10 cached input per million tokens.
- Sol matches GPT-6 Astra on DeepSWE v1.1 at roughly one-fifth the cost and lands within 2.1 points of Astra on OSWorld 2.0 at one-seventh the cost per task.
- Terminal-Bench Science cost per task: $5.47 for Sol versus $23.21 for Claude Opus 5.5 and $23.80 for Astra.
- Context window is 1,050,000 tokens with 128,000 max output; prompts above 272K tokens cost 2x input and 1.5x output.
- The model is closed-weight, API-only (no fine-tuning), available in the OpenAI API, ChatGPT Work, and Codex.
OpenAI released GPT-6.1 Sol on September 29, 2026, claiming near-Astra results on agentic coding, computer use, and professional work at one-fifth of GPT-6 Astra's standard input and output rates. The model is available today through the OpenAI API as gpt-6.1-sol, and in ChatGPT Work and Codex. It upgrades GPT-6 Sol, the mid-tier model in the GPT-6 family, and cached input drops to $0.10 per million tokens — 50% below GPT-6 Sol.
The release matters because it compresses the price gap between frontier and mid-tier performance. If a mid-tier model matches OpenAI's most expensive tier on coding benchmarks at a fifth of the cost, the economics of running autonomous agents change for every team currently paying Astra or Claude Opus 5.5 rates. All benchmark figures below are vendor-reported in OpenAI's launch post; OpenAI states competitor numbers came from public reports.
Where does Sol sit in the GPT-6 lineup?
OpenAI now ships three GPT-6 tiers, and the pricing ladder is steep:
- GPT-6 Astra: $10 input, $50 output, $1 cached input per million tokens
- GPT-6.1 Sol: $2 input, $10 output, $0.10 cached input per million tokens
- GPT-6 Luna: $0.10 input, $0.50 output, $0.01 cached input per million tokens
The cached input price is the number that matters most for agents. Agents resend the same system prompt, tool schemas, and history on every step. Cached reads now cost 5% of the uncached input rate, down from 10% on GPT-6 Sol. For long-running agent workloads, that halving compounds across thousands of steps.
What do the benchmarks show?
The vendor-reported results span coding, professional document work, computer use, science, and factuality.
Coding. On DeepSWE v1.1, GPT-6.1 Sol matches GPT-6 Astra at roughly one-fifth of the cost. It beats GPT-6 Sol's best score by 6.4 percentage points, at a lower reasoning effort.
Professional work. On GDP.pdf, which tests answers over complex professional PDFs, Sol beats Claude Opus 5.5 with fallbacks at less than half the cost per task. On AutomationBench 1.0.6, Sol scores 2.2 points above Opus 5.5 at medium effort, at roughly one-third the cost.
Computer use. On the OSWorld 2.0 offline set, Sol gains 7 points over GPT-6 Sol at maximum effort. It lands within 2.1 points of Astra at roughly one-seventh the cost per task.
Science. On Terminal-Bench Science 0.1, Sol more than doubles GPT-6 Sol's score at max effort. Average cost per task is $5.47, versus $23.21 for Opus 5.5 and $23.80 for Astra. Astra still leads at 68.1%, and OpenAI recommends it for the hardest research. This is the one domain where the flagship keeps a clear edge.
Factuality. At low effort, the share of responses with a factual error falls from 11.4% to 7.7% — a reduction of about 32%. The eval uses deliberately difficult conversations where users had flagged earlier model errors.
What API details do developers need?
The model page specifies the constraints that will shape deployment decisions:
- Context window of 1,050,000 tokens, 128,000 max output tokens, and an April 30, 2026 knowledge cutoff
- Text and image input; text output
reasoning.effortaccepts low, medium (default), high, xhigh, and max; the none and minimal settings are not supported- Tool calling requires the Responses API; Chat Completions works without tool calling
- Prompts above 272K input tokens cost 2x input and cache rates, and 1.5x output, for the full request
- Batch and Flex are 50% cheaper; Fast mode costs 2x standard
- US and EU data residency are supported; Fast mode is unavailable with EU residency
- Fine-tuning is not supported
OpenAI also plans a GPT-6.1 Sol Ultrafast option in Codex within days, promising up to 8x faster token generation than standard speed.
How does GPT-6.1 Sol compare with rivals?
Against the current field of flagship-adjacent models, Sol's pricing is aggressive on cache reads. Standard first-party API list prices, verified September 30, 2026:
| Feature | GPT-6.1 Sol | Claude Opus 5.5 | Claude Sonnet 5.5 | Gemini 3.1 Pro Preview |
|---|---|---|---|---|
| Status | Generally available | Generally available | Generally available | Preview |
| Input (per 1M) | $2.00 | $4.00 | $2.00 | $2.00 (≤200K) |
| Output (per 1M) | $10.00 | $20.00 | $10.00 | $12.00 (≤200K) |
| Cached input read | $0.10 | $0.20 | $0.20 | $0.20 + $4.50/1M tokens/hr storage |
| Context window | 1,050,000 | 1M | 1M | 1,048,576 |
| Max output tokens | 128,000 | 128K | 128K | 65,536 |
| Long-prompt surcharge | >272K: 2x input/cache, 1.5x output | None; 1M at standard rates | None; 1M at standard rates | >200K: $4 input, $18 output |
| Input modalities | Text, image | Text, image | Text, image | Text, image, video, audio, PDF |
| Reasoning control | low to max (5 levels), default medium | Adaptive, default medium | Adaptive, default high | Thinking supported |
| Batch pricing | 50% off standard | $2 / $10 | $1 / $5 | $1 / $6 |
| Knowledge cutoff | Apr 30, 2026 | Jun 2026 | Jun 2026 | Not listed |
Three caveats apply. Claude Sonnet 5.5 matches Sol's $2 and $10 list price, but Sol's cached input costs half as much — and OpenAI's benchmarks compare Sol against Opus 5.5, not Sonnet 5.5. Gemini 3.1 Pro matches Sol on input, charges $12 for output, and remains in preview. And list prices are not direct cost comparisons: Anthropic notes its newer tokenizer produces roughly 30% more tokens for the same text.
Can you deploy it, and where?
Yes, but only as a hosted model. GPT-6.1 Sol is live in the OpenAI API as gpt-6.1-sol, and in ChatGPT Work and Codex. Self-hosting is not an option; the weights are closed, as they are for every model in the comparison table above. Fine-tuning is off the table too, which pushes teams toward prompting and reasoning-effort tuning for behavior control.
Why this release shifts the market
The pattern across OpenAI's benchmarks is consistent: near-flagship scores at a third to a seventh of the per-task cost. On Terminal-Bench Science 0.1, Sol's $5.47 average per task undercuts both Opus 5.5 ($23.21) and Astra ($23.80) by more than four times. On OSWorld 2.0, the gap to Astra is 2.1 points at one-seventh the cost.
The $0.10 cached input rate is the sharpest competitive edge in the table. For agent workloads that re-send context on every step, cache economics can dominate total spend. Sol halves Anthropic's $0.20 cache-read price while matching Sonnet 5.5's base rates.
The limits are equally clear. Astra retains the lead on the hardest science tasks at 68.1% on Terminal-Bench Science. The 272K-token surcharge threshold means very long prompts carry a 2x input and 1.5x output penalty that competitors with flat 1M-token pricing avoid. EU-residency deployments lose Fast mode, and the promised Ultrafast option in Codex — up to 8x faster token generation — arrives within days, giving teams running latency-sensitive agent pipelines a reason to wait before benchmarking Sol against their production workloads.
Original: openai.com
More from James Calloway
Show full bio
News editor covering industry trends and analytics at AI In Context.
200 articles