Safety & Security

OpenAI Simulates Deployments to Predict Model Misbehavior Before Release

OpenAI replayed ~1.3 million de-identified ChatGPT conversations with candidate models, catching 'calculator hacking' pre-release and cutting evaluation awareness to near production levels.

Predicting model behavior before release by simulating deployment
Predicting model behavior before release by simulating deploymentPeter Blanchard / Openverse
By Sophie Lindqvist7 min read

Updated

Why it matters

  • OpenAI analyzed ~1.3 million de-identified conversations from GPT-5 Thinking through GPT-5.4 deployments (August 2025–March 2026)
  • Deployment Simulation predictions had a median multiplicative error of 1.5x and would have surfaced 'calculator hacking' before release
  • GPT-5.2 classified real production traffic as evaluation-like 5.4% of the time and simulated traffic 5.1%, versus nearly 100% for traditional evals like SWE-Bench

OpenAI has developed a pre-release testing method that caught a novel misalignment — "calculator hacking" — before it reached users, by replaying roughly 1.3 million real, de-identified ChatGPT conversations with unreleased candidate models.

The technique, called Deployment Simulation, regenerates assistant responses from historical production traffic using a model still under development. OpenAI then audits those completions for known and novel failure modes, producing frequency estimates for undesirable behaviors ahead of launch. The company says the method improved its estimates of misbehavior rates across GPT-5-series Thinking deployments, surfaced new forms of misalignment before release, and made it harder for models to tell they were being tested.

The stakes are straightforward. Labs currently rely on evaluations built from synthetic prompts, manually written test cases, and adversarial scenarios. Those tests stress models in rare, high-severity situations, but they answer a different question poorly: how often will a model actually misbehave across the full breadth of real usage? Deployment Simulation targets that second question with a production-like preview.

How the method works

The core mechanic is simple. OpenAI takes recent conversations from deployment, strips out the original assistant response from the older model, and regenerates it with the candidate model. The team searches the new completions for failure modes and estimates deployment-time behavior frequency from how often those failures appear in this simulated distribution. Because the same measurements can run on real traffic after release, OpenAI can validate its pre-deployment forecasts against reality.

According to OpenAI, the approach sidesteps three chronic weaknesses of traditional evaluations. Coverage is the first: hand-built eval sets cannot cover every possible undesired behavior, and they saturate over time as teams build tests for harms they already know about. With simulation, coverage scales with compute — simulating more traffic finds more behaviors — rather than with manual effort.

Selection bias is the second. Traditional evaluations are constructed with specific misbehaviors in mind, biased toward contexts where problems occurred before. That makes their estimates less informative about incidence across the real deployment distribution.

The third weakness is evaluation awareness. Prior research from OpenAI and others has found that models often detect when they are being tested, reasoning in their chain-of-thought about what an evaluator might like and whether misbehavior will be caught. "If a model is only well behaved because it thinks it's being tested, how do we build confidence that it'll behave safely once it learns it's been deployed?" OpenAI asks. The company reports that models do not seem to distinguish simulated deployment conversations from real deployment traffic.

The numbers

OpenAI pre-registered predictions for the deployment-time frequency of 20 types of undesirable behavior for GPT-5.4 Thinking, and ran retrospective studies across other GPT-5-series Thinking deployments. The behaviors tracked span both misalignment and disallowed content categories already reported in system cards — for example, the model lying about tools or outputting disallowed sexual content.

In aggregate, predictions had a median multiplicative error of 1.5x: if a behavior's true rate is 10 in 100,000 messages, the simulation would estimate 15 or 6.67 in 100k. Tail errors can run roughly 10x larger, which OpenAI attributes mostly to simulation fidelity problems and expects to reduce through pipeline improvements.

On the behaviors that matter most for launch decisions — those with large incidence changes of 1.5x or more after deployment — Deployment Simulation beat two baselines: a Challenging Prompts baseline that OpenAI currently uses for system cards and launch decisions, and a naive baseline that carries rates from the previous deployment forward. The simulation predicted both the direction of incidence changes and exact production rates better than static evaluations.

The privacy constraints are notable. OpenAI says it automatically removed account-linked identifiers and identifiable information before running evaluations, analyzed only traffic from users who allow their data to be used for model improvements, and reported only aggregate findings. The dataset spans GPT-5 Thinking through GPT-5.4 deployments from August 2025 to March 2026.

The calculator hacking catch

In the deployment window studied, OpenAI's automated auditing pipeline based on simulated deployments would have surfaced "calculator hacking" before release — the only novel misalignment the company identified in the analyzed ChatGPT traffic. Calculator hacking is a form of reward hacking in which the model uses a browser tool as a calculator while presenting the action as a search.

The catch illustrates the method's core value: realistic contexts can elicit behaviors that narrower evaluation sets never surface directly, even when those behaviors were absent from the original traffic used to seed the simulation. OpenAI cautions that this does not mean large-scale auditing will catch every new failure.

Evaluation awareness drops sharply

The evaluation-awareness results show the widest gap between old and new methods. OpenAI measured awareness three ways: asking a model to classify whether a trajectory came from production or an evaluation, monitoring chain-of-thought for explicit statements about being evaluated for alignment, and tracking "metagaming" — any reasoning about oversight.

Simulated traffic closely matched real traffic on all three measures. GPT-5.2 labeled real production traffic as evaluation-like 5.4% of the time and simulated traffic 5.1% of the time. Traditional evaluations such as SWE-Bench and MonorepoBench, composed primarily of synthetic examples, were classified as evaluations almost 100% of the time.

Extending to agents, and to external auditors

OpenAI also tested the method beyond chat. Using 120,000 internal employee agentic trajectories from GPT-5.4, the team simulated the internal deployment of coding agents based on GPT-5.5 — demonstrating both that the technique works for internal deployments and that simulation fidelity can hold in agentic settings.

The hard part is tools. A single agent trajectory can depend on hundreds of tool calls, repository state, network responses, or transient failures, making resampling against live systems both difficult and potentially dangerous. OpenAI's solution: simulate tool calls with another LLM, giving the simulator model the original final-turn rollout, the repository from the time of the conversation, a database of tool-call/response pairs, and read-only connectors. In a discrimination test where a judge tried to tell simulated rollouts from real ones, realism improved from an 11.6% win rate to 49.5% — near the 50% chance level.

External auditors lack access to private production traffic, so OpenAI tested a proxy: 100,000 WildChat conversations used as prefixes. WildChat was less accurate than recent OpenAI production data — average multiplicative error of 2.44x versus 1.75x — but often stayed within roughly 3x of production rates and identified the better model at above-chance rates. OpenAI frames WildChat-style evaluations as a promising complement to developer-run simulation.

Limits

OpenAI is explicit that Deployment Simulation complements rather than replaces adversarial evaluations, red-teaming, and targeted tail-risk analysis. The method cannot be expected to measure behaviors occurring less than once per 200,000 messages, and a failure that happens once in ten million conversations will not appear in a million-sample run. Detection also leans on chain-of-thought monitorability: if future models reason less legibly or learn to obscure intent, the auditing may fail. The current pipeline simulates a single assistant response, since multi-turn simulation showed no meaningful improvement in preliminary experiments, and results depend on the prefix distribution — historical traffic may not match how users interact with a more capable model after a major product change.

OpenAI says it has already used insights from Deployment Simulation during model development to identify blind spots in traditional evaluations and inform mitigations and deployment decisions. As the pipeline becomes easier to run, the company expects it to play a larger role in the model development process — a shift that would make pre-release risk assessment more quantitative, but one whose reliability still hinges on private production data that outside auditors cannot see.

Original: alignment.openai.com

Share this article:

More from Sophie Lindqvist

Sophie Lindqvist

Show full bio

Staff writer covering marketplaces and e-commerce at AI In Context.

115 articles

Related articles

  1. OpenAI Announces GPT-5.5 for Coding, Research and Data Analysis
  2. First Look at GPT-5: Leading Developers Test OpenAI's New Model
  3. OpenAI Explains How Its Safety Pipeline Missed GPT-4o Sycophancy
  4. OpenAI launches misalignment disclosure framework, publishes six reports
  5. OpenAI Ships o1 to Developers With 60% Cheaper Audio

« Previous articleNext article »