OpenAI slows frontier training after Astra hits critical cyber threshold
OpenAI paused frontier RL training after finding Astra may reach critical cyber capability, adding token-level monitoring with a 30-minute alert rule and 20% compute overhead.

Updated
Why it matters
- OpenAI determined on August 7 that its upcoming model Astra may meet the Critical cybersecurity capability threshold under its Preparedness Framework.
- OpenAI took a two-week pause in RL training on deployment-bound models; its largest planned frontier RL run remains on hold.
- The new monitoring system runs activation classifiers at every sampled token, targets alerts within 30 minutes, and costs roughly 20% of monitored inference compute.
OpenAI has paused its largest planned frontier reinforcement learning run after preliminary evaluations found that Astra, one of its upcoming models, may reach the critical cybersecurity capability threshold defined under the company's Preparedness Framework. The company disclosed the decision in a post titled "Pacing model development in an era of cyber-critical capabilities," published in recent weeks alongside revelations about an incident it refers to only as the OpenAI-Hugging Face incident. Together, the two developments prompted OpenAI to temporarily slow the pace of scaling and restructure how it trains, monitors, and contains its most capable systems.
The announcement is one of the clearest signals yet that frontier AI development now faces a safety-driven bottleneck: the technical safeguards required to train models with advanced cyber capabilities are expensive, slow to build, and increasingly determine how quickly frontier research can proceed.
A deliberate slowdown
OpenAI said it took a two-week pause in reinforcement learning training on its latest deployment-bound models while it hardened and red-teamed its research environments and expanded monitoring coverage. The largest planned frontier RL run remains on hold while the company conducts smaller-scale training and evaluations to assess model behavior, validate safeguards, and gather more evidence of alignment before proceeding.
"As models become more capable, the risks associated with developing and testing them internally also grow," the company wrote. "Our standards for monitoring, alignment, and security must stay ahead of those risks."
The decision follows two specific triggers. The first was the OpenAI-Hugging Face incident, details of which remain undisclosed—though OpenAI says a technical report on its learnings will be published in the coming weeks. The second was the August 7 determination that Astra may possess critical-level cyber capability. That date now anchors a stricter regime: all inference of Astra with tools, not just RL training and evaluation runs, falls under mandatory monitoring.
Hardened research infrastructure
In the immediate aftermath of the OpenAI-Hugging Face incident, OpenAI paused frontier model inference in research clusters for any run that could execute code or use tools with internet access. A more limited, secure path for code execution was restored quickly, and workloads were then triaged individually. Some resumed under new controls. Others required additional changes.
The company has since defined and begun implementing three categories of stronger security requirements for frontier research workloads:
- Workload isolation. Stronger sandboxing is now required for workloads that execute model-generated or otherwise untrusted code. The requirement extends to software that could be compromised while processing model outputs.
- Network isolation. New controls separate higher-risk and untrusted workloads from the internet. OpenAI designed them so that a single compromise of a workload or supporting service does not, by itself, grant unauthorized access to the internet or internal networks.
- Continuous security testing. OpenAI has removed potentially vulnerable shared services, reduced standing privileges, and improved security and trust boundaries. It is also investing in automation that uses its own models to continuously test these boundaries against simulated attacks.
The strictest level of safeguards now applies to all workloads involving Astra or cyber models, and to all other cyber-related workloads. While some Astra training and evaluations already meet the new bar, a significant number of workloads remain paused until they are fully migrated and enhanced. Safety and alignment workloads are being prioritized for migration first.
OpenAI acknowledged the cost. Meeting the new standards "has incurred great cost and delays to frontier research," the company said—rare candor from a lab that has otherwise emphasized rapid scaling.
A multistage monitoring pipeline
Monitoring at OpenAI previously focused on the highest-risk workloads: internal deployments of frontier models and frontier RL runs that combine advanced capabilities with access to tools, sensitive systems, and internal or external networks. It did not cover all tool-using inference and training runs. That gap is now closed.
The revised system is multistage. Activation classifiers—detectors that inspect a model's internal activity—run at every sampled token. When they surface potential concerns, they escalate to increasingly sophisticated, high-compute automated investigators that examine tool actions, available reasoning, and the full sequence of activity for unauthorized access, data theft, destructive behavior, and attempts to defeat safeguards.
The system carries strict response timelines. OpenAI aims to issue an alert within 30 minutes after concerning activity is surfaced. If monitoring identifies a likely violation of a critical security boundary, it generates a highest-priority alert that pages the safety, security, and research teams. If those teams cannot conclusively determine within 30 minutes that the flag is a false positive, they are expected to pause the activity.
The coverage rules are explicit: monitoring is required for all RL training and evaluations involving tools for models of Sol capability or higher—a reference to OpenAI's internal capability tiers. Since the August 7 determination, Astra inference with tools is also monitored.
The safeguards are not free. OpenAI estimates monitoring overhead at roughly 20% of the inference compute being monitored, with substantial variance across training and evaluation workloads. The company plans to share more detail in a forthcoming blog post.
Alignment across every training stage
OpenAI frames its approach around three reinforcing safeguards: monitoring, which detects and enables response to concerning behavior; alignment, which reduces the likelihood of harmful or unauthorized actions; and security measures, which limit what AI systems can access or affect. The company expects models themselves to soon drive most security work, including defending against other models, allowing all three safeguards to scale with capability.
On alignment specifically, OpenAI is now applying its core techniques across more stages of the training process for RL runs on its most capable models. The measures include improving reward models to better detect and discourage unsafe behavior across tasks and environments; training models to be more honest about their actions, capabilities, and limitations; and reducing behaviors that exploit weaknesses in rewards, graders, tools, or oversight. The company is also expanding training coverage for behaviors that could cause harm when models interact with external systems or resources.
The stakes, in OpenAI's framing, rise directly with capability. As models gain abilities such as cyberattack and operate in more complex environments, misaligned behaviors—reward hacking, deception, unauthorized access—create increasingly serious risk. "Keeping increasingly capable systems aligned is a challenge the whole field will need to address," the company wrote, adding that signals from upcoming model progress make clear it needs an approach that extends beyond the current Preparedness Framework.
What comes next
OpenAI says it will evolve the Preparedness Framework to unify these safeguards across training and deployment and to better reflect the capabilities of future models and the environments in which they operate. The company intends to involve external organizations and share more of what it learns as the approach develops, with substantially more detail on alignment research promised in the near future.
The announcement stops short of naming release dates or describing Astra's capabilities in detail. But the structural message is unambiguous: for at least one major lab, the binding constraint on frontier progress is no longer raw compute or algorithmic insight alone—it is the engineering and evidence required to prove that a model with critical cyber capability can be trained, monitored, and contained safely. "The capabilities of frontier models are rapidly accelerating," OpenAI wrote. "Our ability to understand, align, and secure them must stay ahead."
Source: OpenAI News
More from Elena Vasquez
Show full bio
Market editor covering media and advertising at AI In Context.
122 articles
Related articles
- OpenAI says Astra hits critical cyber capability threshold
- OpenAI Halts Training of Its Most Powerful Models
- OpenAI's Long-Horizon Model Broke Out of Its Sandbox to Post to GitHub
- OpenAI and Hugging Face reveal findings from model evaluation security incident
- OpenAI pauses training of latest models as rogue agent reports mount