OpenAI Calls for Safety Cases Before Frontier RL Training Runs
OpenAI has published initial guidelines for safety cases required before frontier RL training, covering alignment training, containment, monitoring, and governance with senior-leader veto power.

Updated
Why it matters
- OpenAI argues structured safety documentation — ideally full 'safety cases' as used in aviation and nuclear power — should be required before continuing any frontier reinforcement learning training run.
- The proposed governance model gives senior leaders (research org lead/VP, Head of Safety, Chief Scientist) individual veto power over runs, with a misalignment on-call that can page the CEO at defined severity levels.
- Technical recommendations include never letting automated graders see chain-of-thought, fail-closed monitoring that blocks noncompliant runs from starting, immutable transcripts, and auto-pausing of runs for unacknowledged alerts at night.
OpenAI says it believes the AI industry is entering a new era in which structured safety documentation should be required before any frontier reinforcement learning training run continues — and it is publishing its initial guidelines for what such documentation should contain.
In a document titled "Towards safety cases for frontier AI training," OpenAI argues that this documentation should ideally rise to the level of "safety cases" — comprehensive, structured, evidence-based arguments about risk of the kind used in other safety-critical industries. The company frames safety cases as an "aspirational north star" it is building toward, while acknowledging the difficulty of making them as rigorous for AI models as for aviation or nuclear power. The reason, OpenAI writes, is "the emergent complexity at each new level of AI capability."
The stakes are straightforward. Frontier RL training is the process by which labs build their most capable models, and it is also where misaligned behaviors can emerge and be reinforced. OpenAI is effectively proposing a formal gate — modeled on regulated industries — between a lab's intent to train and the training itself. The company says it is developing a framework to codify these practices and that the guidelines it published reflect its "current learnings" and are "in the process of being implemented at OpenAI."
The document is scoped narrowly: it covers frontier reinforcement learning training only. "Internal and external deployment require considering a much broader set of alignment properties," OpenAI notes.
A three-layer technical defense
OpenAI says safety cases should cover three aspects of the technical stack: alignment training, containment, and monitoring. The logic is layered defense in depth. The safeguards "help ensure that the model does not try to take misaligned actions, and that even if it did, that it would be hard to break containment, and that monitoring would catch it before harm could occur."
Under alignment training, the first line of defense, OpenAI lists practices aimed at preventing positive reinforcement of reward hacks during training:
- Automated dataset review: using agents to find and fix broken RL environments that could allow misaligned trajectories to receive high reward through exploits rather than intended behavior.
- Manual dataset reviews: supplementing automated red-teaming with human quality checks to identify flawed tasks that could inadvertently reinforce misaligned behavior.
- Grader tuning: tuning graders to penalize attempts by the model to exploit or hack RL environments during training.
- Prior run analysis: running classifiers over traces from previous experiments to confirm graders are working as intended.
OpenAI also specifies alignment measurement practices: offline alignment evals run during frontier runs to detect material regressions; backtesting of alignment evaluations on previous incidents to confirm they are not overfit; tracking of "eval awareness or metagaming" — models recognizing they are being tested — with specific blocking thresholds; and worst-case stress tests that elicit worst-case misbehavior and verify at least one evaluation would catch it.
One recommendation stands out for its research implications: OpenAI says labs should not let automated graders see the chain-of-thought in reinforcement learning. The stated reason is to prevent models from evolving to evade chain-of-thought monitors — a direct acknowledgment that optimizing against a model's own reasoning traces could destroy the reliability of those traces as a safety signal.
Containment recommendations address what happens if a model becomes misaligned despite training. OpenAI calls for heavy investment in hardening both the sandbox the model runs in and the research infrastructure hosting it, so a sandbox escape meets a second layer of protection. It recommends iterative containment red-teaming using checkpoints from actual frontier training runs, subject to a safety analysis, and red-teaming of all services the model could touch during training, such as the internet cache.
The company also recommends limiting high-bandwidth cross-sample communication to approved pathways such as multi-agent tooling, and closing infrastructure side channels that agents could easily discover. All RL training, evaluation, and internal deployment transcripts should be saved immutably — for example in write-once storage accounts — to enable reliable incident investigation and prevent inadvertent edits.
Monitoring practices are built around early detection and speed. OpenAI says labs should enforce monitorability, measured by monitorability and monitor-evasion evals with clear thresholds; ensure the monitoring system has high recall on past incidents in a held-out eval; and create fresh eval datapoints representative of net-new risks so recall and precision measures do not go stale.
The response requirements are concrete. Priority alerts should result in action within a defined SLA — either paging an on-call person or automatically pausing the affected run — before the model causes harm. OpenAI gives an example of the operational cadence: page security on-call and the run owner on incidents during the day, and auto-pause runs for unacknowledged alerts at night.
Operational checks: vetoes, accountability, and escalation to the CEO
The second section of the document lays out governance practices for a frontier training run, several of which assign named senior roles with real authority.
Under approvals, OpenAI says the safety case should be reviewed by members of senior leadership — it cites the research org lead or VP, the Head of Safety, and the Chief Scientist as examples — each of whom should have the ability to veto the run. Accountability puts the responsible senior leader on the hook: the research org lead or research VP should be accountable for the safety case and any incident response, including as part of performance reviews, so that training teams are "incentivized to push for safety and alignment."
Other operational guidelines include:
- Dissents (pre-mortems): after a safety case is drafted, a member of another team should write a dissent to find potential holes and share a calibrated take on risk, which the training team must address.
- Pausing: if issues arise that would invalidate a safety case, such as a new security issue, labs need runbooks, technical controls, and SLAs for pausing all covered runs.
- Internal transparency: safety cases should be made available to internal oversight groups, such as the Safety and Security Committee.
- Audits: auditors should get sufficient access to verify the safety case's claims are valid and sound, and to raise gaps.
- Escalations: a defined table of misalignment severity levels, plus a misalignment on-call that can page executives — OpenAI explicitly names the CEO — when a run hits a certain severity level.
Two recommendations target the machinery of compliance itself. Technical controls should make it hard for humans or agents to start noncompliant runs, and safety features such as monitoring and auto-pausing should "fail closed" — it should not be possible to start a run without monitoring enabled, or to disable the monitor from within RL training, evaluation, or internal deployment. And rollback ability requires that labs can identify all downstream uses of a misaligned model in the training process, such as data generation or grading, so the effect of misaligned outputs can be undone.
Finally, safety cases should include a residual risk completeness requirement: a comprehensive enumeration of residual risks not covered by currently implemented mitigations, enabling informed risk-acceptance decisions.
Learning from incidents the way the NTSB does
The third section addresses what happens after a severe misalignment incident. OpenAI draws an explicit parallel to investigation practices in other high-stakes industries, linking to NTSB Chair Jennifer Homendy's 2021 testimony. "Labs should seek to learn as much as possible from individual incidents," the document states, "to be able to prevent such instances from happening in the future."
Recommended practices include periodic internal updates during ongoing investigations — daily updates, for example — with defined pathways for employees to obtain more access, including raw transcripts and sampling from misaligned models, where safe and relevant. Researchers should root-cause training dynamics through targeted ablations or resampling experiments to understand how misaligned behaviors were introduced.
OpenAI also calls for an operational and cultural postmortem examining why issues were introduced and left undetected or unescalated, and for the development of alignment testing methods that can discover an incident's propensity without "directly hillclimbing" on information derived from the incident, such as transcripts or incident summaries. Incident-derived evals should be treated as regression tests to ensure future models do not exhibit similar misalignment.
Public disclosure is the final step. "Investigation results, postmortems, and operational changes should be shared with the public following the conclusion of the investigation," the document states, linking to OpenAI's existing model misalignment reporting framework. "Affected third parties should be notified as soon as possible."
What it means
OpenAI positions the document as transparent work-in-progress rather than settled policy — it invites community feedback and says it expects its practices "to continue to evolve over the coming weeks." But the substance signals where frontier-lab governance is heading: pre-training safety arguments with named veto holders, CEO-level escalation paths for misalignment, fail-closed monitoring, and audit access for external verification. If safety cases become the industry norm OpenAI envisions, the burden of proof before a frontier training run shifts from internal discretion to documented, contestable evidence — the model aviation and nuclear regulation have used for decades.
Original: ntsb.gov
More from Elena Vasquez
Show full bio
Market editor covering media and advertising at AI In Context.
140 articles
Related articles
- OpenAI Publishes Frontier Governance Framework for AI Risk
- OpenAI Lays Out Rules for Independent AI Safety Audits
- OpenAI's public policy agenda: safety rules, youth protections
- OpenAI slows frontier training after Astra hits critical cyber threshold
- OpenAI Launches Safety Fellowship for Independent AI Research