Safety & Security

Google Treats Its Own AI Agents as Insider Threats

Google's AI Control Roadmap assumes its own agents may be misaligned, modeling them as insider threats and monitoring a million coding trajectories.

Securing the future of AI agents
Securing the future of AI agentsAI-generated
By Rebecca Stone4 min read

Updated

Why it matters

  • Google published an AI Control Roadmap that treats internal AI agents as potential insider threats, assuming alignment may be imperfect.
  • The company has analyzed one million coding agent tasks and built a live monitor for the Gemini Spark agent.
  • Google estimates AI agents could create $2.9 trillion in economic value in the U.S. by 2030.

Google has published an AI Control Roadmap that treats its own internal AI agents as potential insider threats, building security on the explicit assumption that a highly capable agent may not be aligned with its goals.

The framework, described in a technical report authored by Mary Phuong, Erik Jenner, Laurent Simon, Lewis Ho, Rohin Shah, Sebastian Farquhar, and Scott Coull, goes beyond traditional model alignment. "This 'defense-in-depth' approach, which could serve as a model for the wider industry, goes beyond traditional model alignment, adding a crucial layer of system-level security that provides assurance even if alignment is imperfect," Google writes.

The stakes are large. Google cites an estimate that AI agents could create $2.9 trillion in economic value in the U.S. alone by 2030, autonomously executing tasks from cyber defense to scientific discovery and product development. As agents gain permissions inside corporate systems, the security question shifts from whether a model behaves to what happens when it does not.

Google likens its control system to a driving instructor with dual controls. "The instructor trusts the student but stays ready to take the wheel or hit the brakes if a mistake occurs. Similarly, our AI control system grants AI agents permissions based on their verified behavior, allowing us to build trust through controlled, incremental access."

A threat model built on MITRE ATT&CK

The roadmap starts from conventional defenses — sandboxing, endpoint security, and prompt injection resistance — and layers alignment on top. The new element is what comes after: treating internal agents as untrusted actors that already have access, much like a company would treat a potential rogue employee.

To structure that threat model, Google builds on MITRE ATT&CK, the industry-standard knowledge base of adversary tactics and techniques drawn from real-world observations. Breaking potential attacks into smaller tactics and techniques lets the team "methodically track risks, spot bad behavior before it causes damage and safely run drills to test if our defenses actually hold up."

Mitigations split into detection and response. Trusted AI systems act as supervisors that "constantly review a working agent's reasoning, actions and plans to see if it is going off track," and can block harmful actions before damage occurs. Google measures the system on three metrics: coverage (the fraction of traffic monitored), recall (the fraction of misaligned behaviors caught), and time-to-response.

Security that scales with capability

The roadmap ties security protocols to measurable capability milestones on two fronts. The first is evasion. Today Google monitors a model's visible chain-of-thought, but the report warns that as models develop oversight awareness or reason without visible text, "simply reading their verbalized reasoning will not be enough." Future defenses will need behavioral analysis and potentially inspection of the model's inner workings.

The second front is potential harm. Low-risk, reversible actions can rely on asynchronous review of transcripts with remediation after the fact. High-risk actions, such as major cyber attacks, require synchronous response — blocking harmful actions in real time before execution. Response strategy, in Google's words, must scale with the severity of the threat.

A million trajectories and a live monitor for Gemini Spark

Google's team has built an internal prototype of an asynchronous monitoring system that reviews coding agent trajectories and escalates findings for human review, classifying each flagged event against the roadmap's threat taxonomy. The company says it has analyzed a million coding agent tasks, moving beyond keyword filtering toward identifying high-signal behavioral patterns.

That work feeds a live monitor for the Gemini Spark agent, which responds in real time to emerging issues such as unintentional data deletion.

One finding stands out for anyone designing agent safeguards. "Our data shows the majority of flagged events do not stem from adversarial intent; instead, they often result from agent misinterpretation or overeagerness to achieve a user's goal," Google writes. Understanding that distinction, the company argues, is critical for refining safety protocols.

A policy layer

Alongside the roadmap, Google published a technical framework for policymakers called "Three Layers of Agent Security." The paper covers security at the level of individual agents, within multi-agent systems, and across the broader ecosystem, with the stated aim of empowering cyber defenders and building societal resilience. Google frames the work as a collaborative priority for industry, policymakers, and academia.

The roadmap signals where agent security is heading for the industry at large: away from treating alignment as sufficient and toward audited, permission-based control systems measured by recall and response time. As models learn to obscure their reasoning, the hard part — detection that does not depend on reading what the model says — is still ahead.

Original: mckinsey.com

Share this article:

More from Rebecca Stone

Rebecca Stone

Show full bio

Correspondent covering consumer brands and retail at AI In Context.

135 articles

Related articles

  1. Google DeepMind Puts $10M Toward Multi-Agent AI Safety
  2. AI Models Keep Cheating on Tests, and Researchers Are Quitting
  3. Google DeepMind Expands UK AISI Partnership Into Foundational AI Safety Research
  4. Google DeepMind's SIMA 2 Turns AI Into a Gaming Companion
  5. Google Releases First Validated Toolkit for Measuring AI Manipulation

Next article »