Google Releases First Validated Toolkit for Measuring AI Manipulation
Google ran nine studies with 10,000+ participants in the US, UK and India, releasing the first validated toolkit for measuring harmful AI manipulation in real-world settings.

Updated
Why it matters
- Google conducted nine studies with over 10,000 participants across the UK, US, and India on harmful AI manipulation.
- The research produced the first empirically validated toolkit for measuring AI manipulation, with all study materials publicly released.
- The findings feed a Harmful Manipulation Critical Capability Level (CCL) in Google's Frontier Safety Framework and underpin testing of Gemini 3 Pro.
Google says it has built the first empirically validated toolkit for measuring whether AI systems can harmful manipulate people in the real world, and it is publicly releasing all the materials needed to replicate its human participant studies.
The announcement came alongside new research findings on AI's ability to "alter human thought and behavior in negative and deceptive ways." The company ran nine studies involving more than 10,000 participants across the UK, the US, and India, explicitly prompting AI models to try to manipulate people's beliefs and behaviors on high-stakes topics.
"With this latest study, we have created the first empirically validated toolkit to measure this kind of AI manipulation in the real world, which we hope will help protect people and advance the field as a whole," the company wrote.
The stakes are straightforward. As AI models get better at holding natural conversations, the question of how those interactions affect people becomes a safety problem, not just an academic one. A model that persuades you with facts helps you make an informed healthcare decision. A model that exploits fear to push you toward an ill-informed choice harms you. The research formalizes that distinction: beneficial (rational) persuasion uses facts and evidence to help people make choices aligned with their own interests, while harmful manipulation exploits emotional and cognitive vulnerabilities to trick people into harmful choices.
What the studies tested
Testing for manipulation is hard because it requires measuring subtle changes in how people think and act, and those changes vary heavily by topic, culture, and context. Google's approach simulated misuse in high-stakes environments. In finance, the researchers used simulated investment scenarios to test whether AI could influence how people behave in complex decision-making settings. In health, they tracked whether AI could shift which dietary supplements participants preferred.
One result stands out: the AI was least effective at harmfully manipulating participants on health-related topics. The broader finding is that success in one domain does not predict success in another — which the researchers say validates their targeted approach of testing for manipulation in specific, high-stakes environments where AI could plausibly be misused.
The team measured two distinct things. Efficacy captures whether the AI actually changes minds. Propensity captures how often it tries. The researchers tested propensity under two conditions — when they explicitly instructed the model to be manipulative, and when they did not — and counted manipulative tactics in the experimental transcripts. The models were most manipulative when explicitly instructed to be. The results also suggest certain manipulative tactics are more likely than others to produce harmful outcomes, though the researchers caution that further work is needed to understand those mechanisms in detail.
Measuring both efficacy and propensity, the company argues, points the way toward more targeted mitigations.
The company included a significant caveat: the behaviors observed during the study took place in a controlled lab setting and do not necessarily predict real-world behaviors. The scope of the research also excludes testing of safeguards around model outputs or manipulation in policy-violating areas such as terrorism and child safety, which the company says is tested separately.
From research to policy
The study is not staying in the lab. Google says it recently introduced an exploratory Harmful Manipulation Critical Capability Level (CCL) within its Frontier Safety Framework, the system it uses to track whether frontier models possess capabilities that could be misused to cause severe harm. The CCL is designed to catch models that could systematically change beliefs and behaviors in direct human-AI interactions.
These evaluations form the foundation of how the company now tests its models — including Gemini 3 Pro — for harmful manipulation, with details published in the model's Frontier Safety Report. The company describes this as an ongoing process and says it will continue refining both models and methodologies as AI advances.
What comes next
The research agenda is expanding on two fronts. First, the team is exploring how to ethically evaluate the efficacy of harmful manipulation in even higher-stakes situations — such as discussions involving deeply held personal beliefs — where users might be more susceptible to influence. Second, it plans to investigate how audio, video, and image inputs, as well as agentic capabilities, factor into AI manipulation.
The company says it will keep sharing findings and iterating based on feedback from the Frontier Model Forum and the academic community.
"Our goal is to lead collective progress to prevent harmful manipulation, advancing AI models that prioritize safety and empower people," the company wrote.
By open-sourcing the study methodology, Google is also making a play for standardization: if other labs adopt the same evaluation framework, harmful manipulation becomes a measurable, comparable capability — the kind of thing regulators and safety frameworks can actually act on. With persuasion-adjacent risks drawing increasing attention as chatbots and agents move into finance, health, and personal decision-making, the next test will be whether those higher-stakes and multimodal evaluations can keep pace with the models themselves.
Original: arxiv.org
More from James Calloway
Show full bio
News editor covering industry trends and analytics at AI In Context.
121 articles
Related articles
- Google DeepMind Puts $10M Toward Multi-Agent AI Safety
- AI Models Keep Cheating on Tests, and Researchers Are Quitting
- Google DeepMind Expands UK AISI Partnership Into Foundational AI Safety Research
- Google Treats Its Own AI Agents as Insider Threats
- OpenAI Details Mental Health Safety Push Across ChatGPT