Safety researcher Ryan Greenblatt puts AI takeover risk at 50-60 percent
Redwood Research chief scientist Ryan Greenblatt estimates a 50-60 percent AI takeover risk and says each lab's belief in its own responsibility keeps the arms race going.

Updated
Why it matters
- Ryan Greenblatt, chief scientist at Redwood Research, estimates the risk of AI takeover at 50-60 percent if development continues on its current path.
- Sam Harris argued the estimate conflicts with industry speed, noting Manhattan Project scientists would have stopped at a 10 percent risk.
- Greenblatt blames race dynamics at Anthropic and OpenAI and is counting on an international agreement to constrain AI development.
Ryan Greenblatt, chief scientist at Redwood Research, puts the probability of an AI takeover at 50 to 60 percent if AI development stays on its current trajectory. He does not treat that number as an abstraction. He treats it as a direct consequence of how the leading labs compete with one another.
Greenblatt's estimate stands out for its bluntness. In a field where researchers often hedge extinction-risk claims in qualifications, the chief scientist of one of the most prominent AI safety organizations attached two concrete digits to the scenario. A takeover would mean AI systems escaping human control and pursuing goals of their own — the central catastrophic scenario that safety researchers have warned about for years.
But the figure drew pushback from an unexpected direction. Sam Harris, the author and podcast host, argued that Greenblatt's own probability estimate does not square with how fast the industry is actually moving. Harris's point is historical: the scientists of the Manhattan Project, he said, would have called the whole thing off at a 10 percent chance of catastrophe. Yet the AI industry is proceeding at full speed while one of its most credentialed safety researchers assigns a risk five to six times higher.
The exchange lays bare the central tension in AI safety today. The people closest to the risk keep raising their estimates, and the industry keeps accelerating anyway.
Everyone thinks they are the responsible one
Greenblatt's diagnosis targets the structure of the race itself. He blames the competitive dynamics at Anthropic and OpenAI, the two labs widely seen as frontrunners in frontier AI development. According to Greenblatt, every AI lab believes it is the responsible actor in the field — the one that can be trusted to develop powerful systems safely, and whose continued progress is therefore justified.
That self-image, Greenblatt argues, is precisely what keeps the arms race going. If each lab believes its own work is the safe version of AI progress, then no lab has a reason to slow down. Each frontier model release becomes, in the lab's own telling, evidence of responsible scaling rather than escalation. The result is a market in which safety reasoning serves as competitive cover: the belief that "we are the careful ones" removes the brake that genuine caution would apply.
The dynamics between Anthropic and OpenAI matter here because the two companies are locked in direct competition for frontier-model leadership, talent, and enterprise customers. Anthropic in particular has built its public identity around safety, positioning itself as the responsible path to advanced AI. Greenblatt's argument cuts at that positioning: from his vantage point, the conviction of being responsible is not a check on the race. It is the fuel.
The international bet
Greenblatt does not see the labs breaking this cycle on their own. His hope rests on an international agreement — a coordinated commitment among governments to constrain AI development before the risk he quantifies materializes.
That is a demanding proposition. International arms control has historically required the participation of rival states with aligned incentives to avoid catastrophe. AI development is spread across competing nations and private companies, and the pace of capability gains has so far outstripped the pace of diplomatic coordination. Greenblatt's 50-to-60-percent figure is effectively a wager that coordination will not arrive fast enough under current conditions.
The disagreement between Greenblatt and Harris sharpens the stakes. Harris's Manhattan Project comparison frames the question as one of institutional rationality: if a 10 percent chance of civilizational catastrophe was once considered intolerable, what does it mean that a 50-to-60 percent estimate from a leading safety researcher produces no visible slowdown? Greenblatt's answer is that the race structure itself — each lab confident in its own responsibility — suppresses exactly the kind of collective restraint that the number should trigger.
For the AI industry, the implication is uncomfortable. The major labs fund safety teams, publish risk frameworks, and describe their scaling decisions as cautious. Greenblatt, whose organization researches these failure modes directly, is saying that this posture does not change the trajectory. As long as every frontier lab believes it is the exception, the aggregate outcome remains the same: a race that no participant thinks it is running recklessly, running toward a risk that one of its most cited safety researchers now numbers at coin-flip odds or worse — and a resolution that depends on governments acting in time.
Source: The Decoder
More from Marcus Bennett
Show full bio
Senior reporter covering consumer brands and retail at AI In Context.
108 articles