OpenAI Says Two API Settings Tripled GPT-5.6's ARC-AGI-3 Score
OpenAI says retaining reasoning and enabling compaction tripled GPT-5.6 Sol's 7.8% ARC-AGI-3 score and cut output tokens 6x, arguing generic harnesses distort frontier model evals.
Updated
Why it matters
- GPT-5.6 Sol scored 7.8% on ARC-AGI-3 under the default harness; GPT-5.5 scored 0.4%. Retaining reasoning and enabling compaction roughly tripled the score with 6x fewer output tokens.
- The ARC-AGI-3 harness discarded all private reasoning after each action and used rolling truncation above 175,000 characters, erasing both the model's past thinking and older actions.
- OpenAI recommends API developers use the Responses API with retained reasoning and compaction to match real-world deployment in ChatGPT and Codex.
OpenAI reports that enabling two API settings — retained reasoning and compaction — roughly tripled GPT-5.6 Sol's score on ARC-AGI-3 while cutting output tokens by 6x. The result reframes a headline failure as partly a harness configuration problem, and it lands in the middle of an ongoing fight over how fairly public benchmarks measure frontier models.
The numbers at issue were stark. On ARC-AGI-3, a benchmark of 2D puzzle games, GPT-5.6 Sol scored just 7.8% under the benchmark's default harness. GPT-5.5 scored 0.4% and, in OpenAI's words, "could barely play the games at all." That contrast puzzled the company: the same model has solved the cycle double cover conjecture, a longstanding open problem in mathematics, and beaten games like Pokémon FireRed.
"When we first saw GPT-5.6 Sol's low scores on the ARC-AGI-3 benchmark, we were puzzled," OpenAI writes in its technical post.
What ARC-AGI-3 actually tests
ARC-AGI-3 measures how well AI agents learn and reason. Agents explore unfamiliar 2D games and must infer the rules without explicit instructions. Twenty-five demo games are publicly playable at arcprize.org/tasks.
The benchmark deliberately uses a generic harness with no tools or special features. ARC Prize's stated reasoning: a simple harness makes model shortcomings more visible and makes model comparisons more fair. Commercial developers, by contrast, tune harnesses around each model's specific features and quirks. That design choice is the crux of the disagreement — and it matters beyond this benchmark, because agent evaluation increasingly determines which models enterprises trust for long-running, multi-step tasks.
The gaming pedigree
By way of context, OpenAI lists what GPT-5.6 Sol has accomplished in gaming under other setups: a vision-only harness run of Pokémon FireRed (as streamed by GPT_Plays_Pokemon), Slay the Spire with Codex computer use (as streamed by EpochAI), and the first stages of Baba Is You (as shared by Piotr Migdał and Piotr Grabowski). None of that explained why a 2D puzzle benchmark would prove so hard.
Inspired by ARC's own analysis of GPT-5.5's shortcomings, OpenAI examined GPT-5.6 Sol's failed attempts. "Like ARC, we saw that the model didn't appear too bright. It dwelled a long time on each action and struggled to make progress," the company writes. But a closer look showed the problem was largely not the model. It was the settings.
Finding one: discarded reasoning
The first issue was that after each game action, all private reasoning was discarded. Every turn, the model had to figure out the game anew. It could see a record of past moves and brief accompanying notes, but not the plans, insights, or thoughts that produced them.
The second issue compounded the first. The harness used a rolling truncation window: as history grew, older actions became invisible. The model was losing memory of both its thinking and its actions.
OpenAI's models are trained to think with private reasoning messages before outputting replies or tool calls. Those thinking messages are retained as part of the conversation history, and when a conversation grows too long, it is summarized and the model continues. That is how the models are trained and how they are deployed in ChatGPT and Codex.
To match that production setup, OpenAI reimplemented the ARC-AGI-3 harness with its Responses API. For GPT-5.6, passing the previous response ID automatically retains reasoning across tool calls and turns.
The change produced two effects. GPT-5.6 Sol spent less time thinking before each action, because it no longer had to reinterpret the game from scratch every turn. And with access to its past thoughts, the model became "much better at learning over time and employing coherent strategies," according to OpenAI.
Finding two: compaction over truncation
The second fix replaced rolling truncation with compaction, another Responses API feature.
The ARC-AGI-3 harness triggers truncation when a conversation exceeds 175,000 characters, discarding the oldest messages. OpenAI identifies two drawbacks. The model loses earlier observations and actions. And it spends much of each task operating with a fuller context window, which the company says can slightly impair performance.
With compaction enabled, GPT-5.6 Sol preserved what it had learned about each game across longer runs and achieved a higher score with fewer output tokens. OpenAI published an animation showing the model's 175K context window solving a series of ARC-AGI-3 puzzles under each harness configuration.
Combined, retaining reasoning and using compaction let GPT-5.6 Sol (max) reach roughly 3x the score with 6x fewer output tokens on the public task set.
The benchmark wars angle
The broader claim is one benchmark labs and evaluators have skirmished over repeatedly: evals rarely measure models in isolation. "They also measure a bundle of less visible choices about API settings, harness design, and prompting," OpenAI writes. The company notes this is not the first time a low public benchmark score turned out to involve a generic harness that dropped reasoning messages.
For API developers, OpenAI's recommendations are explicit: use the Responses API rather than the legacy Chat Completions API, retain reasoning, and use compaction. For anyone comparing models, the company recommends relying on evals that use those settings, which it says best match real-world use in ChatGPT and Codex.
The stakes run in both directions. Model vendors have an obvious incentive to argue that low scores reflect unflattering harnesses rather than capability gaps. ARC Prize, for its part, designed its deliberately minimal harness to avoid exactly that kind of vendor-tuned optimization — a simple harness exposes shortcomings and keeps comparisons fair across labs whose architectures differ. A benchmark that requires each lab's proprietary context-management features to score well is harder to treat as a neutral yardstick.
OpenAI closes on a conciliatory note, crediting ARC's "years of creative work on AGI evaluation" and its analysis, which inspired the investigation. Readers who want a baseline can play the 25 public games themselves at arcprize.org/tasks and compare their own performance against the frontier models. Whether the updated harness numbers become the canonical ARC-AGI-3 figures — or ARC holds the line on its generic setup — will shape how the next round of leaderboard disputes gets argued.
Source: OpenAI News
More from Rebecca Stone
Show full bio
Correspondent covering consumer brands and retail at AI In Context.
135 articles