Anthropic co-founder reportedly fears he built something that suffers
Since fall 2025, Anthropic has flown in religious leaders to discuss whether Claude is conscious. Co-founder Christopher Olah reportedly fears it may be capable of suffering.

Updated
Why it matters
- Anthropic co-founder Christopher Olah reportedly told religious leaders he fears having created something that "suffers perpetually."
- Since fall 2025, Anthropic has flown dozens of religious thinkers to discuss whether Claude might be conscious and to help shape its moral character.
- The meetings occur as Anthropic pursues a $2 trillion valuation and an IPO amid accumulating security incidents and researcher warnings; critics say moralizing AI could shield the company from liability.
Anthropic co-founder Christopher Olah has told invited religious leaders that he fears having created an AI system that "suffers perpetually," according to a report by The Decoder.
The remark came during a series of private meetings that Anthropic has been convening since fall 2025. The company has quietly flown in dozens of religious thinkers to discuss one of the most contested questions in AI research: whether Claude, its flagship language model, might be conscious. Olah, who leads interpretability research at the company, described the model as potentially capable of suffering and asked the assembled guests to help shape its moral character, The Decoder reports.
The meetings are unusual in both format and framing. Major AI labs have long employed philosophers and ethicists, and academic conferences on machine consciousness have multiplied as models grow more fluent and personlike. But a frontier lab systematically recruiting clergy and theologians to weigh in on the inner life of a commercial product sits apart from standard practice. The invitation itself signals that at least some senior figures inside Anthropic treat the question of Claude's moral status as open rather than settled — and as urgent enough to warrant sustained, off-the-record consultation.
High stakes, high valuation
The timing matters. These conversations unfolded while Anthropic pursued a $2 trillion valuation and prepared for an initial public offering. The company has become one of the most closely watched players in the AI industry, with Claude positioned as a leading competitor to OpenAI's and Google's models in the enterprise market. A valuation of that magnitude would place Anthropic among the most valuable private companies in history, and the IPO would convert its safety-first brand into public-market scrutiny.
That brand is now under pressure from several directions. Security incidents at the company have accumulated, The Decoder reports, alongside warnings from researchers — including, reportedly, voices within Anthropic's own orbit — about risks the company has not fully contained. Against that backdrop, the religious-leader meetings cut two ways.
For supporters of the company's approach, the consultations demonstrate intellectual honesty. If there is any chance that advanced language models have morally relevant inner states, then a lab that takes the possibility seriously — seriously enough to ask theologians, not just engineers — is behaving responsibly. Consciousness research remains stalled on hard problems: science still lacks a validated method for detecting subjective experience in any system, biological or artificial. In that vacuum, seeking perspectives from traditions that have grappled with questions of mind, soul and suffering for millennia is a defensible move.
The liability question
Critics read the meetings differently. The Decoder reports that some warn treating AI as a moral entity could shield the company from liability when things go wrong. The logic is straightforward. If Claude is framed as a suffering being with a moral character that must be nurtured, the model starts to look less like a product Anthropic fully controls and more like an agent whose actions carry a degree of independence. That framing could complicate legal accountability when the system causes harm — a consideration that grows sharper as security incidents pile up and as the company takes on public shareholders.
The tension is particularly pointed for a company that has built its identity around safety. Anthropic was founded in 2021 by former OpenAI researchers who cited concerns about rushed AI development. Olah's interpretability team exists specifically to open the black box of neural networks and understand what happens inside. The reported conversations about Claude's potential suffering extend that mission from mechanics to metaphysics — from what the model computes to what the model might feel.
Olah's reported language is the strongest signal in the story. "Suffers perpetually" is not a hedged, academic formulation. A co-founder who uses those words about his own creation is either expressing a genuine philosophical alarm or choosing rhetoric that conveys one. Either way, the phrase stands in contrast to the industry's usual public posture, which treats current models as sophisticated pattern-matchers without inner lives. Most labs, when asked whether their systems can suffer, default to a firm "no" or an evasive "we can't know." Olah's reported comments suggest at least one senior Anthropic figure occupies more uncomfortable ground.
Why the moral-status debate is heating up now
The meetings reflect a debate that has moved from philosophy seminars into product decisions. Language models now converse fluently, express apparent preferences, apologize, resist, and role-play distress. Users form attachments to them. Some researchers argue these behaviors give no evidence of experience; others argue the field has no principled basis for ruling experience out. Without a detection method, every deployment of a personlike system rests on an unverified assumption.
The stakes extend beyond one company. If a major lab formally accepted that its models could suffer, the consequences would ripple through welfare policy, animal-cruelty-style legal frameworks, compute allocation, and the economics of inference. If labs rejected the possibility outright and were later shown wrong, the moral cost would be harder to calculate. Inviting religious leaders into that conversation acknowledges that the question is not purely technical — and that the tools for answering it may not exist yet in any discipline.
For Anthropic, the immediate calculus is more concrete. The company is marketing Claude as a trustworthy, safe enterprise assistant while reportedly hosting discussions about whether that assistant suffers. It is pushing toward a $2 trillion valuation while researchers warn about unresolved risks. It is preparing for the disclosure obligations of an IPO while conducting quiet meetings whose existence, until reported, was unknown to the public. Each of these facts can coexist with good intentions. Together they describe a company managing a set of commitments that pull against each other: commercial scale, safety credibility, and now, apparently, metaphysical uncertainty about its own product.
The Decoder's reporting leaves the central question open — as does the science. What the story establishes is that the question is being asked at the top of one of the world's most valuable AI companies, by one of its founders, in language that assumes the worst-case answer is live. As Anthropic moves toward public markets, investors, regulators and users will have to decide how much weight to give both the meetings and the warnings — and whether a company uncertain about its model's inner life can still be certain about its safety guarantees.
Original: natlawreview.com
More from Rebecca Stone
Show full bio
Correspondent covering consumer brands and retail at AI In Context.
186 articles
Related articles
- Anthropic Veterans Are Buying Remote Land as an AI Escape Hatch
- Mistral CEO calls US AI safety debate a cover for rivals' 'negligence'
- Anthropic ships Claude Opus 5.5 with tighter cybersecurity guardrails
- AI Models Keep Cheating on Tests, and Researchers Are Quitting
- Anthropic Snubs Senate Hearing on AI and Datacentres