Research

Multi-Agent AI Teams Cost Up to 5.1x More for Marginal Gains

Teams of AI agents cost up to 5.1x more than solo agents while only one of four tests with GPT-6 Sol and Claude Opus 5.5 showed measurable gains, Vals AI finds.

AI agent teams waste massive tokens for barely measurable quality gains, research finds
AI agent teams waste massive tokens for barely measurable quality gains, research findsAI-generated
By Marcus Bennett3 min read

Updated

Why it matters

  • Agent teams cost up to 5.1x more in tokens than solo agents, according to Vals AI.
  • Only one of four tests with GPT-6 Sol and Claude Opus 5.5 showed a measurable quality gain.
  • Anthropic's own data shows quality plateaus beyond ten agents while token costs keep climbing.

Teams of AI agents cost up to 5.1 times more than a single agent while delivering barely measurable quality gains, according to new benchmark research from Vals AI.

The finding lands as multi-agent architectures have become one of the most heavily marketed patterns in enterprise AI. Vendors and consultants routinely promise that swarms of cooperating models — one planning, several executing, another checking — will outperform a lone agent on complex work. Vals AI's numbers suggest the opposite in most cases.

Out of four tests run with OpenAI's GPT-6 Sol and Anthropic's Claude Opus 5.5, only one showed a measurable gain from running agents as a team rather than solo. In the other three, the extra agents added cost without adding demonstrable quality.

Anthropic's own internal data points in the same direction. According to figures cited alongside the Vals AI results, quality plateaus once a team grows beyond ten agents — while token consumption continues to climb.

Why does this matter for AI buyers?

The economics of agent deployment are increasingly decided by token spend, not model quality. If a team of agents burns several times the tokens of a solo agent and produces roughly equivalent output, the multi-agent premium buys very little.

For enterprises budgeting inference costs, the research implies a straightforward test: before committing to multi-agent pipelines, teams should measure whether the added agents produce measurable gains on their own tasks. In three of the four Vals AI test cases, that measurement would have said no.

The 5.1x cost ceiling is the worst case in the study, not the average. But the direction of the result is consistent: more agents meant more tokens, and usually not more quality.

What did the tests actually measure?

Vals AI compared solo agents against agent teams using two frontier models: GPT-6 Sol and Claude Opus 5.5. The evaluation covered four test scenarios.

  • One of four scenarios showed a measurable quality gain from teaming agents.
  • Three of four scenarios showed no measurable gain.
  • Up to 5.1x higher token cost for agent teams versus solo agents.

The pattern held across both models, indicating the issue lies with the multi-agent pattern itself rather than with any single vendor's models.

Does Anthropic's data confirm the plateau?

Yes. Anthropic — which sells the Claude models used in the tests — has published its own data showing that quality gains flatten out beyond ten agents. After that point, adding agents increases token consumption without meaningfully improving results.

That a model provider is publishing numbers that undercut the case for very large agent teams lends the finding weight. It aligns with the independent benchmark results from Vals AI rather than contradicting them.

What is the practical takeaway?

The research does not say multi-agent systems are useless — one test in four did show a gain. It says the gains are not reliable enough to justify the cost as a default architecture choice.

For engineering teams, that reframes the decision. Multi-agent setups need to earn their token budget with measured results on the specific task at hand, not assumed improvements from adding more cooperating models. Where a solo agent matches team performance, the solo agent wins on cost by default.

As agent deployments scale across enterprises, benchmarks like Vals AI's are likely to push buyers toward smaller, measured agent configurations — and to put vendors under pressure to show that the swarms they sell actually outperform the single agents they replace.

Original: vals.ai

Share this article:

More from Marcus Bennett

Marcus Bennett

Show full bio

Senior reporter covering consumer brands and retail at AI In Context.

192 articles

Related articles

  1. OpenAI Signs Big Four Consultancies to Deploy Its Frontier Agents
  2. Anthropic and OpenAI Ship New Models With the Same Pitch: More for Less
  3. OpenAI Models Are Coming to Amazon Bedrock via a Stateful Agent Runtime
  4. OpenAI Launches Frontier, an Enterprise Platform for AI Agents
  5. Cloudflare Brings OpenAI's GPT-5.4 to Agent Cloud for Enterprises

« Previous article