Models

OpenAI Says Two API Settings Tripled GPT-5.6's ARC-AGI-3 Score

OpenAI says retained reasoning and compaction tripled GPT-5.6 Sol's 7.8% ARC-AGI-3 score and cut output tokens sixfold, arguing harness design, not model capability, drove the gap.

How enabling two settings tripled our scores on the ARC-AGI-3 benchmark
How enabling two settings tripled our scores on the ARC-AGI-3 benchmarkPeter Blanchard / Openverse
By James Calloway5 min read

Updated

Why it matters

  • Enabling retained reasoning and compaction roughly tripled GPT-5.6 Sol's ARC-AGI-3 score while cutting output tokens by 6x
  • Initial scores were 7.8% for GPT-5.6 Sol and 0.4% for GPT-5.5 on the 2D puzzle benchmark
  • The ARC-AGI-3 harness discarded private reasoning after each action and truncated history beyond 175,000 characters

OpenAI reports that enabling two API settings — retained reasoning and compaction — tripled GPT-5.6 Sol's score on the ARC-AGI-3 benchmark while cutting output tokens by six times on the public task set.

The company's initial numbers were low. GPT-5.6 Sol scored just 7.8% on ARC-AGI-3, a benchmark of unfamiliar 2D puzzle games. GPT-5.5 could barely play the games at all, managing 0.4%. OpenAI described itself as puzzled by the results: the same models had solved longstanding open problems in mathematics such as the cycle double cover conjecture, and beaten commercial games including Pokémon FireRed, Slay the Spire and the first stages of Baba Is You.

The discrepancy points to a persistent problem in AI evaluation. Benchmarks rarely measure models in isolation. They also measure less visible choices about API settings, harness design, and prompting — a bundle of engineering decisions that can dominate headline scores. The ARC-AGI-3 episode is a concrete case study in how much those choices matter, and it arrives as labs, researchers and policymakers increasingly lean on benchmark numbers to compare frontier models.

What ARC-AGI-3 measures

ARC-AGI-3, run by the Arc Prize Foundation, is designed to measure how well AI agents learn and reason. Agents explore unfamiliar 2D games and infer how they work without explicit instructions. Twenty-five demo games are publicly playable at arcprize.org/tasks.

The benchmark deliberately uses a generic harness, without tools or special features. ARC's stated reasoning: a simple harness makes model shortcomings more visible and model comparisons more fair. Commercial developers, by contrast, optimize harnesses for each model's specific features and quirks. That philosophical gap is what produced the score divergence.

Two harness choices, one broken agent

Inspired by ARC's own analysis of GPT-5.5's shortcomings, OpenAI examined GPT-5.6 Sol's game attempts. Like ARC, the team observed that the model dwelled a long time on each action and struggled to make progress.

But closer inspection revealed that much of the model's confusion came from harness settings, not the model itself.

First, after each game action, the harness discarded all private reasoning. The model had to figure out the game anew on every turn. It could still see a record of past moves and brief accompanying notes, but not the plans, insights or reasoning that produced them.

Second, the harness used a rolling truncation window. When the conversation context exceeded 175,000 characters, the oldest messages were discarded. As history grew, older actions became invisible. The model lost not only its past thinking but its memory of its own past actions.

Together, OpenAI concluded, these two features explained why GPT-5.6 Sol failed to learn over time.

The fix: match the production setup

OpenAI's models are trained to think in private reasoning messages before outputting replies or tool calls. Those thinking messages are retained as part of the conversation history in production. If a conversation grows too long, the system summarizes it and continues. This is how the models are trained and how they run in ChatGPT and Codex.

To match that setup, OpenAI reimplemented the ARC-AGI-3 harness with its Responses API, where passing the previous response ID automatically retains reasoning across tool calls and turns.

The change produced two effects. GPT-5.6 Sol spent less time thinking before each action, because it no longer had to interpret the game from scratch every turn. And with access to its past thoughts, the model became much better at learning over time and deploying coherent strategies.

The second change replaced rolling truncation with compaction, another Responses API feature. Rolling truncation has two drawbacks, OpenAI wrote: the model loses earlier observations and actions, and it spends much of each task operating with a fuller context window, which can slightly impair performance. With compaction enabled, GPT-5.6 Sol preserved what it had learned about each game across longer runs and achieved higher scores with fewer output tokens.

The combined effect: GPT-5.6 Sol (max) reached roughly three times the score with six times fewer output tokens. OpenAI published an animation showing the model's 175K context window solving a series of ARC-AGI-3 puzzles under each harness configuration.

Why this matters beyond one benchmark

The takeaway is a caution about evaluation methodology, not just a score correction. "Evals rarely measure models in isolation — they also measure a bundle of less visible choices about API settings, harness design, and prompting," OpenAI wrote, noting this isn't the first time a low public benchmark score traced back to a generic harness that dropped reasoning messages.

That argument cuts both ways, and benchmark authors know it. ARC's generic harness exists precisely to avoid per-model optimization and to keep comparisons fair across vendors. OpenAI's counter is that a harness which strips reasoning doesn't reflect real-world use in ChatGPT or Codex, and therefore understates deployed capability. The tension between neutrality and fidelity to production settings is unlikely to be resolved soon, and it affects every downstream consumer of benchmark numbers — from procurement teams to regulators.

For API developers trying to maximize performance, OpenAI's recommendations are explicit:

  • Use the Responses API, not the legacy Chat Completions API
  • Retain reasoning
  • Use compaction

For anyone comparing models, OpenAI recommends relying on evals that use those settings, which it says best match real-world use in its products. The company also thanked ARC for its years of work on AGI evaluation and for the analysis that prompted the closer look — and invited readers to test themselves against frontier models on the public games at arcprize.org/tasks.

The episode leaves benchmark consumers with a practical lesson: when a frontier model posts a surprisingly low score, the harness deserves as much scrutiny as the model.

Source: OpenAI News

Share this article:

More from James Calloway

James Calloway

Show full bio

News editor covering industry trends and analytics at AI In Context.

119 articles

Related articles

  1. OpenAI Says Two API Settings Tripled GPT-5.6's ARC-AGI-3 Score
  2. OpenAI says GPT-5.2 sets new state of the art on FrontierMath
  3. First Look at GPT-5: Leading Developers Test OpenAI's New Model
  4. OpenAI says GPT-5.6 Sol cut its own serving costs by 20 percent

« Previous articleNext article »