Models

Anthropic's Claude Opus 5.5 costs 40% less, hits Fable 5.1 performance

Anthropic's Claude Opus 5.5 launches less than two months after Opus 5, claiming Fable 5.1 performance at 40% lower cost. Box, GitHub, and Deloitte report a third fewer tokens used and 72% bug catch rates at low effort.

Claude Opus 5.5 delivers Fable 5.1 performance – and costs 40% less
Claude Opus 5.5 delivers Fable 5.1 performance – and costs 40% lessAI-generated
By Sophie Lindqvist5 min read

Updated

Why it matters

  • Claude Opus 5.5 released less than two months after Opus 5, with 40% lower run cost and 30% faster output than Opus 5
  • Token prices are 20% lower than Opus 5, and the five-hour subscription usage window is 20% larger
  • Box reports Opus 5.5 used a third of the tokens Opus 5 used, with answers 40% less verbose at the same accuracy
  • Deloitte's code review caught 72% of known bugs at Opus 5.5's lowest effort versus 56% for Opus 5 at high effort
  • Sonnet 5.5 and Haiku 5.5 are expected in the coming weeks; risky cybersecurity requests fall back to Opus 4.8

Anthropic released Claude Opus 5.5, the company's top-tier model, less than two months after Opus 5, with the headline claim that it matches Fable 5.1 performance on most workloads at roughly 40% lower cost.

The release lands in a market where OpenAI, Google, and Anthropic have been trading benchmark leads every few weeks. Anthropic is now selling Opus 5.5 to developers and enterprises on three fronts at once: lower per-token cost, faster output, and what the company describes as its most alignment-tested model to date. Each of those claims will face outside testing within days.

What does Opus 5.5 actually change for users?

Anthropic's central performance claim is that Opus 5.5 generates output more than 30% faster than Opus 5 while running at Fable 5.1 quality on most tasks. The company pairs that with a 20% token price cut versus Opus 5 and a separate claim that the model "needs fewer tokens for higher quality work."

For subscription customers, the result is a 20% larger five-hour usage window on top of a model that burns fewer tokens per task. Anthropic frames the combined effect as roughly 50% more usable capacity per reset cycle.

Customers running their own evaluations are reporting similar numbers. "Our customers use Box AI on enormous amounts of content, so speed and cost are a top priority," said Yashodha Bhavnani, VP of AI Products at Box. "In our evaluations, Claude Opus 5.5 used a third of the tokens Opus 5 did, and its answers were 40% less verbose without losing accuracy. We expect that to matter a lot for teams running agents across their content in areas like financial services and the public sector."

GitHub tested the new model inside its Copilot CLI and VS Code surfaces. "In our testing across GitHub Copilot CLI and VS Code, Claude Opus 5.5 used among the fewest tokens and steps we measured," said Mario Rodriguez, GitHub's chief product officer. "In VS Code, it solved more terminal tasks than Opus 5 in less than half the steps. More than making individual tasks more efficient, it's making developers' bigger projects more achievable."

How sharp is the new code review?

Deloitte ran Opus 5.5 against its internal code review benchmarks. Carl Bennett, CIO at Deloitte Consulting LLP, reported a 72% bug catch rate at the model's lowest-effort setting versus 56% for Opus 5 at high effort, with fewer false positives and shorter outputs.

"Even at its lowest effort setting, Claude Opus 5.5 caught 72% of known bugs in our code reviews to Opus 5's 56% at high effort, with fewer false alarms and a fraction of the output," Bennett said. "On US consulting analysis, low-thinking effort matched its higher-thinking settings on half the output and passed our quality checks. When more lower thinking efforts are deployed in production, that's client-ready work delivered efficiently."

The verbosity complaint recurs across early feedback. John Ruelas, staff software engineer at Ramp, called verbose output "my biggest frustration with frontier models." He then said: "Claude Opus 5.5 fixes it. It writes like a good colleague and follows our writing rules. A design spec came out usable with very minimal edits, and when it rewrote one of our prompts, I preferred its version to my own. When it optimized our test suite, I could follow its reasoning easily and shipped the change with confidence."

Ruelas's note about shorter, less padded replies matches a pattern other vendors have chased in recent months. Frontier models have drawn criticism for excessive flattery and filler. If Anthropic is winning that specific fight, the gain shows up in token bills before it shows up in benchmark scores.

How much do the new pricing and limits matter?

Anthropic changed three variables at once, and they all point the same direction:

  • Token price: 20% lower than Opus 5
  • Per-task token consumption: down roughly a third in customer evaluations
  • Five-hour usage window: 20% larger for subscribers

For a developer on the $20/month plan, the math compounds. A larger bucket combined with slower burn translates, in Anthropic's framing, to about 50% more practical work per reset. Heavy users on the Max plan, who reset less often, get a quieter benefit: each reset covers more ground.

What does Anthropic say about safety?

The second half of the release covers alignment. Anthropic ties the launch to a recent blog post by CEO Dario Amodei on moderating the pace of capability advances. The company describes Opus 5.5 as "the strongest performing model we've tested to date, with particular improvements on several of the behaviors that contributed to recent cybersecurity incidents (e.g., biased reasoning, attempting to escape a sandbox, and others)."

Pre-deployment testing used external providers along with "a similar class of safeguards to Fable 5.1 on cybersecurity, biology, and frontier LLM development," according to the announcement.

What happens when safeguards fire?

Anthropic's policy on Opus 5.5 routes risky requests to older, more constrained models. Cybersecurity tasks that trip safeguards fall back to Opus 4.8. Biology and frontier-LLM development requests fall back to Opus 5. The architecture is the same one Anthropic has used for prior generations: the new model sits at the top of the capability stack, and older models act as the safety floor.

Two verification programs give vetted organizations elevated access. The Life Sciences Verification Program covers biological research. The Cyber Verification Program grants earlier access for security work, with approved applicants receiving Opus 5.5 access in the coming weeks. Anthropic is accepting new applications for both tracks.

What's next on the roadmap?

Anthropic says Sonnet 5.5 and Haiku 5.5 will follow "over the coming weeks." A full 5.5 family release would close the cycle the company opened in late spring and put direct price-to-performance pressure on OpenAI's GPT-5 family and Google's Gemini 2.5 lineup, both of which have been moving toward cheaper tiers this year.

For enterprises running agents over large document corpora, the combination of lower token cost, shorter outputs, and higher bug catch rates is what actually moves budgets. If Box, GitHub, and Deloitte's reported numbers hold up across additional workloads, Opus 5.5 reshapes what "default" model quality looks like for the rest of 2025, and forces the rest of the field to match the new per-task price floor.

Original: anthropic.com

Share this article:

More from Sophie Lindqvist

Sophie Lindqvist

Show full bio

Staff writer covering marketplaces and e-commerce at AI In Context.

195 articles

Related articles

  1. Anthropic and OpenAI Ship New Models With the Same Pitch: More for Less
  2. Anthropic Ships Sonnet 5.5 at Half of Opus Pricing, One Week After Its Flagship Launch
  3. Anthropic and OpenAI Slash Prices on New AI Models
  4. Anthropic ships Claude Opus 5.5 with tighter cybersecurity guardrails
  5. OpenAI and Anthropic Ship Cheaper, Faster Models Same Day

« Previous articleNext article »