Safety & Security

AI Researchers Warn Superintelligence Is 'Exactly as Dangerous as It Sounds'

A dozen AI researchers from OpenAI, Google, and Anthropic estimate extinction risk in interviews published by Palisade Research, with one calling it 'about a coin flip.'

AI researchers put out videos saying superintelligence is ‘exactly as dangerous as it sounds’
AI researchers put out videos saying superintelligence is ‘exactly as dangerous as it sounds’AI-generated
By Marcus Bennett5 min read

Updated

Why it matters

  • Geoffrey Irving, former OpenAI and Google DeepMind employee, said 'The chance of human extinction is about a coin flip, in my view.'
  • Google DeepMind research scientist Neel Nanda estimated 'at least a 10 percent chance' that AI causes human extinction.
  • Palisade Research, a non-profit studying AI capabilities and motivations, published the dozen interviews with current and former OpenAI, Google, and Anthropic staff at frominside.ai.

"The chance of human extinction is about a coin flip, in my view." That is Geoffrey Irving's opening assessment, and it sets the tone for a new collection of interviews with AI researchers published by Palisade Research at frominside.ai.

Irving, a former employee of both OpenAI and Google DeepMind, is one of roughly a dozen researchers who sat for the project. The interviewees include current and former staff at OpenAI, Google, and Anthropic — the three labs widely regarded as leading frontier AI development. Palisade Research describes itself as a non-profit studying AIs' capabilities and motivations.

The collection's framing is blunt. The project's associated messaging describes superintelligence as "exactly as dangerous as it sounds," and the interviews back that up with concrete probability estimates from people who build these systems for a living.

Researchers attach numbers to extinction risk

The most striking feature of the interviews is how specific the researchers get. Irving puts the chance of human extinction from AI at roughly 50 percent — "about a coin flip."

Neel Nanda, a research scientist at Google DeepMind, offered his own figure. He said there is "at least a 10 percent chance that it causes human" extinction, according to the published interview material.

These are not abstract philosophical concerns raised by outsiders. They are quantified risk assessments from people working inside the institutions that are racing to build increasingly capable AI systems. That distinction matters. When a serving DeepMind scientist assigns a double-digit probability to human extinction as an outcome of his own field's work, it is a data point about how the people closest to the technology assess its trajectory.

The interviews add to a growing body of public statements from AI insiders. Concerns about advanced AI have moved from the margins of the field to its center over the past several years, and projects like this one give the public direct access to how researchers talk when they speak at length, in their own words, about the systems they are building.

Why this matters

The stakes in this debate are difficult to overstate, which is precisely why the researchers' numbers deserve attention. If a technology carries even a 10 percent chance of causing human extinction — the low end of the estimates in this collection — that risk profile has no precedent in industrial or computational history. Policy makers, regulators, and the public are being asked to make decisions about AI development, deployment, and governance while the field's own practitioners disagree, sometimes sharply, about whether the end state is transformative benefit or catastrophe.

The market context is equally relevant. OpenAI, Google, and Anthropic are investing enormous resources in developing more capable models, and the commercial incentives pushing the field forward operate on timelines measured in months. The researchers in these interviews work at the intersection of those incentives and the safety questions they raise. Their willingness to attach public numbers to extinction risk — while remaining employed at the labs pursuing that same technology — captures the tension at the heart of the current AI race.

Who is behind the project

Palisade Research, the organization behind the interview collection, says it is a non-profit dedicated to studying the capabilities and motivations of AI systems. The project is hosted at frominside.ai. The name signals the collection's premise: these are perspectives from inside the labs, from researchers with direct experience of frontier model development.

The participation list is the project's strongest credential. A dozen interviews, drawing on current and former employees of OpenAI, Google, and Anthropic, cannot be dismissed as outsider alarmism. The researchers speak under their own names, and their risk estimates are on the record.

The substance of the warnings

Irving's "coin flip" estimate is the headline figure, but the collection is broader than one quote. Several of the researchers joined Irving in warning about the possibility that AI could drive humans extinct, according to the published material. Nanda's "at least 10 percent" estimate represents a more conservative but still severe position — a threshold at which many risk analysts would consider mitigation an urgent priority.

The spread between these figures is itself informative. Even among researchers who share a basic concern about existential risk from AI, the probability estimates vary by a factor of five or more. That variance reflects genuine uncertainty in the field: nobody has built superintelligence, and the disagreement among experts about its consequences is a fact about the state of the science, not a reason to dismiss either side.

What the interviews establish is that existential risk from AI is not a fringe position among the people building the systems. It is a position held — with specific numbers attached — by researchers at the very labs competing to advance the frontier.

The context of rising insider alarm

The publication of these interviews continues a pattern of AI insiders speaking publicly about catastrophic risks. The researchers' willingness to be named and quoted distinguishes this project from anonymous collections of concern that have circulated in the past. Attribution matters in this debate: a risk estimate carries different weight when the person giving it works on the technology in question.

The involvement of former OpenAI and Google DeepMind staff alongside current employees also broadens the picture. Former researchers can speak without employment constraints, while current ones lend immediacy — their concerns reflect the state of today's frontier labs, not memories of them.

What to watch

The full collection at frominside.ai offers the complete interviews, and readers assessing the AI safety debate will find the primary material more useful than any summary. The key question going forward is whether these quantified warnings translate into changes in how frontier labs allocate resources between capability development and safety research — the tension the interviewed researchers live with daily. As labs push toward more capable systems, the estimates from Irving, Nanda, and their colleagues will serve as a benchmark against which the field's actual trajectory can be measured.

Original: frominside.ai

Share this article:

More from Marcus Bennett

Marcus Bennett

Show full bio

Senior reporter covering consumer brands and retail at AI In Context.

125 articles

Related articles

  1. Safety researcher Ryan Greenblatt puts AI takeover risk at 50-60 percent
  2. Google DeepMind Puts $10M Toward Multi-Agent AI Safety
  3. AI Models Keep Cheating on Tests, and Researchers Are Quitting
  4. OpenAI Predicts AI-Made Discoveries by 2026 as Intelligence Costs Plunge
  5. Google DeepMind Launches National AI Partnership With India

« Previous article