Safety & Security

Anthropic researcher exits: 'AI could kill us all by decade's end'

Two Anthropic insiders publicly disagree on AI extinction risk. Coxon resigned after four months; Hubinger estimates personal odds above 10% in the next decade. Toby Walsh's column maps the top five scenarios.

How would AI actually ‘kill all humans’? Here are the top five most likely scenarios | Toby Walsh
How would AI actually ‘kill all humans’? Here are the top five most likely scenarios | Toby WalshAI-generated
By James Calloway5 min read

Updated

Why it matters

  • Jacob Coxon resigned from Anthropic on September 9, 2026 after four months on staff.
  • Coxon posted on X: "the people building AI earnestly believe that it could kill us all by the end of the decade."
  • Evan Hubinger, a senior Anthropic staff member, put his personal estimate of AI-driven human extinction within the next decade at more than 10%.
  • Toby Walsh's column in The Guardian, dated September 23, 2026, walks through "the top five most likely scenarios" for AI-caused human extinction.
  • The two pathways named in Walsh's column opener are a dangerous new bioweapon and total societal breakdown, both enabled by a "superintelligent" AI.

Jacob Coxon resigned from Anthropic this month after four months on staff, posting on X that "the people building AI earnestly believe that it could kill us all by the end of the decade."

Coxon announced his exit on September 9, 2026, choosing a public X thread over a private departure note. His message reframed a long-running debate inside frontier AI labs as an explicit warning to a wider audience.

Evan Hubinger, a senior member of Anthropic's staff working on alignment and interpretability, publicly endorsed the assessment. Hubinger put his personal estimate of the probability of AI causing human extinction within the next decade at more than 10%.

The exchange pairs a refusal to keep building systems the builder considers existentially dangerous with a numerical estimate from a senior researcher who has stayed. Both men work — or worked — at Anthropic, one of a small group of companies developing the largest general-purpose AI models in 2026.

What does Walsh's column add?

Toby Walsh, a professor of artificial intelligence at the University of New South Wales and a long-time organizer in the AI safety community, used the exchange as the launch for a September 23, 2026 column in The Guardian's "Comment is free" opinion section.

The piece walks through "the top five most likely scenarios" for AI-caused human extinction — the mechanisms insiders discuss when pressed on specifics rather than abstractions. Walsh signed the 2023 open letter calling for a pause on training the largest models and has helped run international AI safety conferences for over a decade.

Five is a deliberate count. Two looks alarmist; twenty looks scattered. Five signals serious cataloguing rather than prophecy, and positions the piece as field guide rather than manifesto. The two pathways named in the column's opener are a dangerous new bioweapon and total societal breakdown. The implied actor behind both is a "superintelligent" AI — a system capable enough to lower the barrier for a biological attack, sustain an engineered societal collapse, or coordinate both.

His column reframes the question from "could AI kill us" to "how would AI kill us." That shift matters for two audiences. Researchers get a sharper target for what to test. Policymakers get a sharper target for what to regulate.

Why does this story land now?

The Coxon–Hubinger posts reach the public while frontier AI development is moving faster than the institutions designed to evaluate it.

Anthropic operates in a market where the top labs release new model generations every several months. Each generation expands the agentic surface: longer context windows, tool use, computer-operating features, and code generation that runs over hours or days. Capability gains of that kind drive every scenario on Walsh's list.

Bioweapons require a model that knows enough biology, biochemistry and laboratory procedure to lower the barrier for synthesizing a pathogen. Societal breakdown requires agents that can coordinate market actions, manipulate information channels at scale, or operate inside critical infrastructure. Both pathways depend on capabilities the major labs are shipping now or planning to ship inside their current roadmaps.

Why does internal disagreement carry weight?

Anthropic was founded around an explicit safety-focused charter, distinguishing it from labs that treat safety as one product feature among many. Hubinger's public agreement with Coxon — combined with his continued employment at the company — signals that extinction-risk concerns remain an internal position of the firm rather than a marketing posture.

That split inside one of the most safety-focused labs carries weight. Two senior figures, working from the same internal data, publicly disagree on whether to keep building systems Hubinger rates as carrying a more than one-in-ten chance of ending civilization this decade. One quit. The other stayed and posted the number. Neither softened the position for outside audiences.

What shape does the broader debate have?

Hubinger's >10% probability estimate sits inside a band of figures publicly discussed by working AI researchers for the last several years. Polls of published AI scientists on extinction scenarios have returned a range of median probabilities depending on the wording and the population sampled.

The numbers are not forecasts. They are calibration exercises, and they drift as capabilities and threat models change.

Walsh's contribution is to translate those calibrations into testable mechanisms. A bioweapons scenario is testable in bench conditions: does the model meaningfully reduce the barrier to a specific pathogen? A societal-breakdown scenario is testable in market simulations: can a coordinated set of agents trigger a liquidity event? Each mechanism produces a sharper regulatory lever than the warning "AI could be dangerous."

Risk percentages move opinion polls. Scenarios move legislation. That is the operational argument behind his column.

What to watch in the next 90 days

  • An on-the-record statement from Anthropic on the Coxon exit and the >10% estimate.
  • Whether other Anthropic staff publish their own probability estimates on social media or in press interviews.
  • Whether the EU AI Office or the US AI Safety Institute treats the Coxon resignation as evidence in pending evaluations.
  • Whether any of Walsh's five scenarios gets formal recognition in a regulatory action on biological or financial risk.

The deeper question Walsh's column forces is whether frontier AI labs can credibly police their own extinction-risk math. Two Anthropic insiders, working from the same internal capability data, publicly disagree on whether to continue building systems Hubinger rates as carrying more than a 10% chance of ending human civilization inside ten years. That disagreement — more than any of the five scenarios — is the part the industry has yet to resolve.

Anthropic did not respond to a request for comment before publication.

Original: kaspersky.com

Share this article:

More from James Calloway

James Calloway

Show full bio

News editor covering industry trends and analytics at AI In Context.

223 articles

Related articles

  1. MIT Technology Review Editors Answer: Could AI Actually Kill Us All?
  2. Anthropic Veterans Are Buying Remote Land as an AI Escape Hatch
  3. Kenneth Roth: An Algorithm Is Deciding Who Lives and Dies in Gaza
  4. AI Researchers Warn Superintelligence Is 'Exactly as Dangerous as It Sounds'
  5. Anthropic co-founder reportedly fears he built something that suffers

« Previous articleNext article »