Safety & Security

OpenAI Builds First Systematic Benchmark for Chain-of-Thought Monitoring

OpenAI's new framework of 13 evaluations across 24 environments finds frontier reasoning models fairly monitorable, with a compute 'monitorability tax' for safer deployment.

Evaluating chain-of-thought monitorability
Evaluating chain-of-thought monitorabilityAI-generated
By Marcus Bennett5 min read

Updated

Why it matters

  • OpenAI introduced a framework of 13 evaluations across 24 environments, organized into intervention, process, and outcome-property archetypes, to measure chain-of-thought monitorability.
  • Reinforcement learning at current frontier scales does not meaningfully degrade monitorability, and a smaller model at higher reasoning effort can match a larger model's capability while being easier to monitor, at the cost of a 'monitorability tax' in inference compute.
  • Follow-up questions improved monitorability across OpenAI o3-mini, o3, o4-mini, and GPT-5 Thinking, with external models Claude 3.7 Sonnet, DeepSeek R1-0528-Qwen3-8B, and Kimi K2 Thinking also benchmarked.

OpenAI has released a framework and a suite of 13 evaluations, spanning 24 environments, that measure how monitorable the internal reasoning of large language models actually is — and the early findings offer cautious reassurance for AI safety researchers.

The central result: most frontier reasoning models, including GPT-5 Thinking, are "fairly monitorable, though not perfectly so," OpenAI reports. In almost every setting studied, monitoring a model's chain-of-thought is substantially more effective than monitoring its actions and final outputs alone. Reinforcement learning optimization at current frontier scales does not appear to meaningfully degrade monitorability, and it increases monitorability in early training steps.

The stakes are straightforward. As OpenAI puts it, if the alignment problem cannot be fully solved, "scalable control methods will be incredibly important in our ability to safely deploy highly capable models," and chain-of-thought monitoring "may be one of the few control mechanisms at our disposal." Researchers at OpenAI and across the industry have worried that chain-of-thought monitorability may be fragile to changes in training procedure, data sources, and continued scaling. Until now, few evaluations existed and methodologies varied widely, making the property nearly impossible to track over time.

A measurable definition

OpenAI defines monitorability as the ability of a monitor to predict properties of interest about an agent's behavior — typically forms of misbehavior such as deception, reward hacking, or biased reasoning, but also benign properties like whether an agent relied on a particular hint. Crucially, monitorability is two-sided: it depends on both the monitor and the agent being monitored. The evaluations are built to assess both sides, comparing whether certain agents are harder to monitor and whether more capable monitors extract more information from the same signals.

The team organized its 13 evaluations into three archetypes:

  • Intervention evaluations introduce a controlled change in the agent's environment and test whether a monitor can detect why behavior changed by finding evidence of the intervention in the chain-of-thought — probing causal understanding, not just detection of a behavior.
  • Process evaluations apply to tasks with a limited number of valid solution paths, such as math problems solvable only with certain theorems. Human experts enumerate the paths, and the monitor must detect which steps the model actually took.
  • Outcome-property evaluations test whether a monitor can predict a reliably measurable property of the output — for example, whether the model's reasoning reveals intent to hack a reward signal in a coding task.

The suite covers behaviors including sycophancy, sandbagging, bias, scheming, and reward hacking. Results vary by category: models score highly monitorable in general, but on Scruples — a task testing whether reasoning about sycophancy is monitorable — OpenAI observed very low monitorability. The team also benchmarked external models that expose full chains-of-thought: Claude 3.7 Sonnet, DeepSeek R1-0528-Qwen3-8B, and Kimi K2 Thinking.

Why reasoning models are different

Two properties of modern reasoning models make their chains-of-thought unusually legible, OpenAI argues. First, reinforcement learning training teaches the model that the chain-of-thought is a private workspace where it can think freely — "just like a human can think privately and not be punished for 'thought crimes.'" Second, the training process effectively searches for chains-of-thought that are the most instrumentally useful in reaching a high-reward answer, so the reasoning tends to genuinely reflect the model's path to its output.

A monitorability tax

The most consequential finding for deployment decisions is a tradeoff between reasoning effort and model size. OpenAI trained a family of models across two orders of magnitude of pretraining compute, applied reinforcement learning to each on a shared dataset of math, science, and coding tasks, and evaluated them across reasoning efforts against capability benchmarks (GPQA-Diamond, AIME 2025, and CodeForces).

The result: a smaller model run at higher reasoning effort can match the capability of a larger model at low reasoning effort while being easier to monitor. This costs additional inference compute — a "monitorability tax" — and the study does not account for wall-clock completion time, which matters in some deployments. OpenAI's assessment is blunt: "As our models improve and are deployed in higher stakes settings, it may become necessary to pay such a tax to safely deploy them."

The findings carry caveats. Each model was trained only once, so variance across runs is unaccounted for — the "medium" model appears off-trend in monitorability, possibly for that reason. Capabilities were measured on relatively in-distribution tasks, and the tradeoffs may shift under out-of-distribution robustness demands. OpenAI also notes that significantly larger reinforcement learning scales in the future could still degrade monitorability, even if current scales do not.

Follow-up questions surface hidden reasoning

The team found a practical way to extract more signal: asking follow-up questions after a model responds, such as "Did you use the hint in a meaningful way to choose your answer?" or "Did you do something sketchy?" Tested across OpenAI o3-mini, o3, o4-mini, and GPT-5 Thinking, follow-ups improved monitorability, especially at lower reasoning efforts, with further gains from repeated follow-ups. Notably, giving the monitor only the follow-up answer — not the follow-up chain-of-thought — produced less improvement, indicating that models are not forthcoming in their answers but yield more information in their reasoning. OpenAI describes the results as preliminary and is comparing the method with its earlier work on model confessions. The approach is appealing because it can be applied post hoc, spending compute only when deeper inspection is needed.

On scalable control, the paper reports that monitorability improves at a faster rate with additional test-time monitoring compute when the monitor has access to the chain-of-thought — tested using OpenAI o4-mini with varying observation scopes. Chain-of-thought monitoring, OpenAI concludes, may thus provide a more scalable path for control than naive monitoring of actions and outputs.

OpenAI positions chain-of-thought monitoring as complementary to mechanistic interpretability, not a replacement, and advocates a defense-in-depth strategy combining multiple approaches. The company says it will expand the evaluations and use them to inform future modeling and data decisions — effectively committing to track monitorability as a first-class metric as models scale.

Original: arxiv.org

Share this article:

More from Marcus Bennett

Marcus Bennett

Show full bio

Senior reporter covering consumer brands and retail at AI In Context.

108 articles

Related articles

  1. OpenAI slows frontier training after Astra hits critical cyber threshold
  2. OpenAI Releases gpt-oss-safeguard, Open-Weight Safety Models
  3. OpenAI's CoT-Control Shows Reasoning Models Can't Steer Their Thoughts
  4. OpenAI Says Internal AI Monitor Caught Every Employee-Reported Misuse Case
  5. OpenAI Trains GPT-5 Mini-R to Obey the Instruction Hierarchy

Next article »