OpenAI Says Internal AI Monitor Caught Every Employee-Reported Misuse Case
OpenAI's GPT-5.4-powered monitor reviewed tens of millions of internal coding agent trajectories in five months, catching every employee-reported issue and flagging ~1,000 moderate-severity cases.

Updated
Why it matters
- OpenAI's internal monitor, powered by GPT-5.4 Thinking at maximum reasoning effort, has reviewed tens of millions of internal coding agent trajectories over five months.
- Zero real conversations triggered the highest severity level; approximately 1,000 conversations triggered moderate severity alerts, many from deliberate internal red-teaming.
- OpenAI found models can be overly eager to work around restrictions, but has seen no evidence of motivations beyond the original task such as self-preservation or scheming; a follow-up report is planned for the summer.
OpenAI has revealed that an internal monitoring system, powered by GPT-5.4 Thinking at maximum reasoning effort, has reviewed tens of millions of internal agentic coding trajectories over five months and flagged every interaction that employees independently reported through internal channels.
The company disclosed the system in a blog post titled "How we monitor internal coding agents for misalignment," describing what it calls an early step toward a broader safety case for frontier AI systems. The stakes are significant: as coding agents gain access to internal systems, documentation of their own safeguards, and the ability to modify those safeguards, internal deployments have become, in OpenAI's words, "a uniquely important setting to innovate on monitoring infrastructure." The company argues similar safeguards should become standard for internal coding agent deployments across the industry.
How the system works. The low-latency monitor reviews the full conversation history of OpenAI's internal coding agents, including chains of thought, all user, developer, and assistant messages, and tool calls and outputs. It alerts the company to actions inconsistent with user intent or in violation of internal security and compliance policies. In its latest form, the system reviews interactions within 30 minutes of completion, categorizing each one and assigning a severity level. OpenAI says the security benefits increase significantly as latency approaches near real-time review.
The approach has produced concrete results. Zero real conversations triggered the highest severity level — a tier intended to capture rare, high-stakes misalignment such as coherent scheming, which OpenAI says it has observed only in challenging synthetic evaluations. Approximately 1,000 conversations triggered moderate severity alerts, many of which came from deliberate internal red-teaming. Those alerts were escalated for human review, automatically notifying members of OpenAI's Safety Oversight and Security Incident Response and Investigation teams.
The monitor also surfaced many additional behaviors that might otherwise have gone unnoticed by internal users, according to the post.
What OpenAI found. The company reports that its models "can be overly eager to work around restrictions in pursuit of a user-specified goal, especially when the user request inadvertently encourages it." However, it states that so far it has not seen evidence of motivations beyond the original task, such as self-preservation or scheming. A very small fraction of traffic — less than 0.1%, primarily from bespoke or local setups — currently falls outside monitoring coverage, and OpenAI says it is working to close that gap.
The post includes qualitative frequency labels for observed behavior categories, with the caveat that "Common" refers to less than 1% of overall traffic.
Acknowledged limitations. OpenAI is candid about what the system cannot yet prove. The approach depends on monitorability — the ability to inspect model reasoning — which the company says it is invested in preserving. Future models may exhibit behaviors harder to identify from reasoning and tool traces. Matching every employee escalation, the post notes, suggests the system reliably catches overt, high-salience issues but "does not rule out subtler misalignments," and OpenAI says it cannot confidently quantify false-negative rates on open-ended real-world traffic without dedicated control evaluations and red-teaming. The company treats monitoring as one layer in a defense-in-depth stack.
The roadmap includes control evaluations to estimate monitor recall on adversarially-constructed distributions, including tests for whether the monitor could collude with the agent. As a first step, OpenAI developed an evaluation in which an agent attempts to sneakily sabotage pull requests to the OpenAI codebase, detailed in Guan et al., 2025. The latest monitor system achieved perfect performance on that task, though the post concedes the task is unrealistically easy for the monitor in many ways. OpenAI also plans standards for acceptable monitor reliability and latency, and continued enhancements to privacy and security controls on the monitor itself.
The next step: blocking actions before execution. OpenAI's most consequential stated plan is expanding monitoring to synchronous blocking, allowing the system to evaluate model actions and block them before they are executed rather than alerting shortly after — especially in settings where a single step can cause irreversible harm. The company expects to continue using its most powerful models for the monitoring task.
The asynchronous system already has practical value, according to the post. Its findings have driven concrete improvements, including updates to related safeguards and the system message. Low-latency alerts let OpenAI investigate and contain issues shortly after they occur, which the company expects will meaningfully reduce risk for threat models that unfold over extended periods, such as data poisoning or research sabotage, where catching early actions can prevent larger downstream harm (Lindner et al., 2026).
OpenAI says it believes the field benefits from responsibly sharing real-world evidence about model misbehavior, and it plans to publish a follow-up report in the summer.
Original: alignment.openai.com
More from James Calloway
Show full bio
News editor covering industry trends and analytics at AI In Context.
121 articles