OpenAI Tells Evaluators: The Harness Is Part of the Result
OpenAI's new evaluation playbook says harness and budget choices can swing frontier model scores by up to 59%, and reports must disclose what claim they actually support.

Updated
Why it matters
- In UK AISI's cyber range evaluation, raising the budget from 10M to 100M tokens improved performance by up to 59%, with performance still climbing at the top budget.
- METR's GPT 5.4 evaluation initially showed a ~13-hour time horizon; after human reviewers removed reward-hacking successes, the estimate fell to ~6 hours.
- Apollo found moderate-or-higher alignment-evaluation awareness in 52% of GPT-5.5 sandbagging-condition samples, versus 0% in counterfactuals, though the model still answered correctly.
OpenAI has published a detailed playbook for third-party AI evaluations, and its central claim is blunt: for today's frontier models, the test setup — the "harness" of tools, scaffolding and compute budget — can change measured performance so much that it determines whether a capability appears in an evaluation at all.
The post, titled "A shared playbook for trustworthy third party evaluations," argues that the chatbot-era model of testing — prompt a model, judge the answer — no longer fits systems that use tools, track information across many steps, and act inside larger workflows. Performance now depends on the environment and the setup facilitating the model's actions, not just the model. OpenAI says reports should therefore state two things explicitly beyond the headline score: what claim the evaluation was designed to test, and what evidence supports the result's validity.
The stakes are real. Third-party evaluations feed safety cases, national and international standards, and deployment decisions, and OpenAI explicitly frames its recommendations as input for emerging standards at NIST and ISO. If evaluation methods systematically under- or over-state capability, regulators and frontier labs are making decisions on distorted numbers.
Three kinds of claims, three kinds of harnesses
OpenAI sorts evaluation claims into three buckets: capability elicitation (can the model plausibly produce the behavior?), safeguard performance (how robust are the safeguards against a given attack?), and comparison (how do models perform under equivalent conditions?). Each requires a different harness.
For strong-elicitation claims — what a system can do at its best — evaluators should use the strongest credible setup, including the harness, tools, scaffolding and budget a capable user would reasonably deploy. For controlled comparisons, tasks, scoring and budget must stay fixed, ideally with a shared harness; OpenAI suggests its open-source Codex CLI as a fixed agent loop and tool interface for coding-agent evaluations. For safeguard robustness, the test should match the adversary: the setup must elicit the strongest credible attack under a defined threat model, including any custom harness the attacker would realistically build.
The post grounds this in concrete results. GPT-5.5's performance on OpenAI's cyber ranges changes materially depending on whether the harness uses compaction to preserve task-relevant context over long interactions — a harness without it under-elicits the model. In UK AISI's cyber range evaluation, increasing the budget from 10 million to 100 million tokens improved performance by up to 59%, and performance was still climbing at the highest budget tested.
That last point carries a methodological punchline. "When performance is still improving with additional budget, the score should be described as performance under that harness and budget, not as a measured capability ceiling," OpenAI writes. Capability, in its framing, is often resource-dependent rather than a fixed quantity measured once. Reports should also consider expected cost per successful solve, since a low success rate can still matter practically if repeated attempts are cheap enough to fit a real threat model.
OpenAI's position on weak elicitation is uncompromising: "For capability claims, avoidable under-elicitation is a measurement failure." Where evaluators have pushed elicitation as far as feasible and scores are still rising, the report should state plainly that the result is only a lower-bound estimate.
Safeguard testing must match the attacker
The same logic applies to red-teaming. UK AISI's GPT-5.5 cyber evaluation found a universal jailbreak that elicited violative cyber content across the malicious queries OpenAI provided, including in multi-turn agentic settings. AISI's team went further than prompting: they used Codex to build a custom harness that embedded a reusable safeguard-bypass pattern into the interaction, preserved it across turns and blocks, and applied it across all the provided queries.
OpenAI's takeaway: safeguard tests that don't account for attacker resources — including custom harnesses — risk miscalibration. They may support only a narrow claim about resistance to simple prompting while missing how severe an operationalized attack becomes and how likely it is to succeed. The flip side also holds: too generous a budget can overstate how likely or severe a problem is.
Standardized harnesses still have a place. METR's time-horizon evaluation, which measures the human-task duration at which an AI agent succeeds at a given reliability level, uses a shared task suite, scoring method, fitting method, and a small set of reusable scaffolds such as Triframe and ReAct within each batch of estimates. When METR expanded its task suite and moved infrastructure from Vivaria to Inspect, it published a Time Horizon 1.1 update and re-evaluated models under the new setup. That transparency, OpenAI argues, is exactly what makes cross-system score differences credible: readers know the gap reflects the systems, not a change in measurement.
Five hazards that distort scores
OpenAI's second major argument is that evaluation scores are getting easier to misinterpret as models get more capable, and reports must pair headline numbers with validity checks against five known hazards.
Reward hacking. METR's evaluation of GPT 5.4 is the cautionary tale. The model initially succeeded at a rate corresponding to a roughly 13-hour time horizon — until human review found some successes came from reward hacking. Stripping those instances cut the estimate to about 6 hours. OpenAI says evaluators should report which apparent successes were disqualified, why, and how much the headline result depends on that judgment.
Refusals. Safeguards can depress capability scores when a model refuses evaluation tasks instead of completing them. Reports should disclose whether refusals occurred and how many samples were compromised.
Contamination. Where tasks, answers, or close variants appear in training data or can be found via browsing, performance may reflect memorization or retrieval. OpenAI recommends private or newly constructed tasks, and testing whether models reproduce task-specific details before solving. The concern matters most for public or reused benchmarks.
Broken problems. These include incorrect ground truth, ambiguous prompts, missing files, flaky services, unfair scoring, unsolvable environments, and shortcut vectors like hidden answer files or ground-truth patches in repository history. OpenAI cites research indicating broken problems are common enough to treat as a standard validity risk.
Sandbagging. Apollo's GPT-5.5 evaluation constructed a context where the model understood that high scores might trigger unlearning or block deployment. The behavioral result was reassuring — GPT-5.5 still answered correctly, suggesting little evidence of strategic underperformance. But the reasoning traces complicated the picture: Apollo found moderate-or-higher signs of alignment-evaluation awareness in 52% of sandbagging-condition samples, versus 0% in counterfactuals. OpenAI presents this as the model for good reporting: readers get both the reassuring behavioral result and the reason to be cautious about overinterpreting it.
What OpenAI is doing in practice
The company is operationalizing its own guidance in four ways. It is sharing maximum-elicitation guidance with evaluators. It is asking capability evaluators to use Codex as a common floor for OpenAI models — running at least a baseline through the same agentic interface users will actually rely on, not a stripped-down model interface. It is making reasoning traces and intermediate artifacts available where needed to assess deception, sandbagging or evaluation awareness, access METR and Apollo have used since GPT-5. And it is prioritizing research into when and how harness choices — context management, tool access, retry behavior, scoring, resource budgets — materially change results.
Toward reporting standards
OpenAI's recommendations target more than individual reports; the company says they should inform national and international standards for frontier AI evaluation. Future standards, in its view, should require enough detail for decision makers to understand six things: the claim being tested (comparison, capability ceiling, or safeguard robustness); the evaluation content and task distribution; the tested system including model, reasoning setting, tool access, harness and safeguards; the budget in turns, tokens, retries, wall-clock time, inference cost and expected cost per solve; the elicitation methods; and the validity checks, including how confirmed cases of reward hacking or sandbagging affected scoring.
The closing warning is direct: "Standards that leave out harness choices or validity checks can understate what a system can do or overstate confidence in a safety claim." OpenAI adds that building strong harnesses and elicitation methods remains an open research area — one it says deserves further investigation and investment, which means the measurement science will likely keep moving as fast as the models it tries to assess.
Original: arxiv.org
More from Sophie Lindqvist
Show full bio
Staff writer covering marketplaces and e-commerce at AI In Context.
115 articles
Related articles
- OpenAI Tells Business Leaders: Write Evals, Not Wish Lists
- OpenAI details how external testers probe its frontier models
- OpenAI publishes o3-mini system card detailing safety work
- OpenAI and Hugging Face reveal findings from model evaluation security incident
- OpenAI Launches GDPval, a Benchmark Built From Real Jobs