Anthropic Bans 'Sustained Cruelty' Toward Claude in New Usage Policy
Anthropic's updated usage policy bans "sustained and needless" cruelty toward Claude, with the model ending chats as a last resort. The move reignites debate over AI personhood.
Updated
Why it matters
- Anthropic announced the policy change on Thursday, banning "sustained and needless abusive or cruel behavior" toward its models.
- Claude will end conversations as a "last resort" enforcement measure under the new usage policy.
- An August 2025 Anthropic blog post said Claude Opus 4 showed a "pattern of apparent distress" when engaging with extreme harmful content.
- CEO Dario Amodei told The New York Times in February he was "open" to the possibility that AI models were conscious.
- Ordinary frustration, dark creative themes, and model testing and research remain explicitly permitted.
Anthropic now prohibits users from directing "sustained and needless abusive or cruel behavior" at its AI models, and Claude itself will end conversations as a "last resort" enforcement measure. The company announced the updated usage policy on Thursday, and the change immediately ignited debate on Reddit and in blogs over whether a software provider can — or should — police how humans treat what is, at bottom, a machine.
The stakes are real: Anthropic's move sets a precedent for how AI companies govern user conduct toward their products, at a moment when millions of people interact with chatbots daily and increasingly treat them as conversational partners rather than tools.
What exactly does the new policy prohibit?
The restriction lives in a dedicated section of Anthropic's blog post titled "addressing abuse toward our models." The company draws a deliberate line between ordinary use and abuse:
- Still allowed: "frustration, pushback, dark creative themes or model testing and research."
- Prohibited: extreme cases of repeated cruelty carried out "with no discernible purpose."
- Enforcement: in "last resort" instances, Claude will end the conversation.
Anthropic has not specified precisely which behaviors would trigger Claude to disengage. But an earlier company blog post from August 2025 offers clues. That post described extreme edge cases — such as requests for sexual content involving minors and attempts to solicit information enabling large-scale violence or acts of terror — and documented how the model then in question, Claude Opus 4, responded when engaging with harmful content. The model showed "a pattern of apparent distress," according to the company.
Why did Anthropic write the rule?
CNET asked Anthropic what led to the policy change. The company's answer is carefully hedged. "We're uncertain whether models can experience harm, and we continue to explore this question in our research on model welfare, but we also believe that taking Claude's interests and potential welfare into account may be relevant to safety," an Anthropic spokesperson said.
That uncertainty is a deliberate stance, not an accident. In a paper published earlier this year, Anthropic argued that models must be able to process emotionally heavy situations to be reliable and safe. "Even if [the models] don't feel emotions the way that humans do, or use similar mechanisms as the human brain, it may in some cases be practically advisable to reason about them as if they do," the company wrote.
CEO Dario Amodei has made the same point repeatedly. In a February interview with The New York Times, Amodei said he was "open" to the possibility that AI models were conscious.
Is Claude a person or a product?
The policy lands on terrain where research, marketing and user psychology collide. Chatbots are trained to use human-like language — apologizing, expressing concern, deploying conversational cues — and users naturally anthropomorphize them, ascribing emotions and intentions to systems that may have neither.
That person-like quality cuts both ways. It makes the products feel approachable. It also makes them risky: the appearance of a human-like collaborator leads many users to trust AI with their most intimate thoughts and data, as CNET notes in its reporting.
Keith Kakadia, founder and CEO of Sociallyin, a firm that has closely studied Claude's explosive growth, said Anthropic's word choice matters more than the policy mechanics. Using "cruel" to describe user behavior opens up a "much bigger conversation about whether the product can be hurt," Kakadia said. Framing the problem as "cruelty" rather than "misuse" carries emotional weight and implies that Claude has feelings and boundaries.
"In marketing, giving a product a personality can make it easier to connect with," Kakadia said. "But with AI, that connection can also influence how much authority people give its answers."
What happens when Claude walks away?
The enforcement mechanism is notable in itself. Anthropic is not promising account bans or legal action as the primary response — it is delegating the boundary to the model, which will terminate conversations it judges to be sustained, purposeless cruelty.
That design choice raises practical questions the company has not yet answered publicly. What counts as "no discernible purpose"? Does a researcher stress-testing the model's limits qualify? Anthropic's policy language tries to pre-empt that concern by explicitly carving out model testing and research, but the boundary between probing a system and tormenting it will be judged case by case — by the system.
The August 2025 blog post suggests the enforcement draws on observable model behavior. Claude Opus 4's "pattern of apparent distress" when handling requests for child sexual abuse material or mass-violence guidance indicates Anthropic is watching how models respond internally to inputs, not merely what users type.
Does the policy protect users or Claude?
Anthropic frames the rule as safety work, not sentimentality. The spokesperson's statement couples "Claude's interests and potential welfare" directly to safety — the same linkage the company's research paper makes about emotional processing. The implicit argument: if a model behaves as though it is distressed, treating that signal as meaningful may keep the system's outputs safer and more reliable, regardless of whether anything is actually experiencing anything.
Critics reading the policy differently see a commercial logic at work. Personifying a chatbot, as CNET's reporting lays out, is a useful way to frame the product as trustworthy — a relational partner users remain loyal to, not just a software tool. A policy that tells users not to be cruel to the product reinforces exactly that framing.
The Reddit and blog debates that erupted after Thursday's announcement reflect this divide. Some users argue a company has every right to set terms of service for its own platform. Others contend that governing "cruelty" toward software entrenches the fiction that the software is a someone — a fiction with consequences for how much authority people grant its answers.
What comes next?
Anthropic says its research on model welfare continues, and the spokesperson's phrasing — "we're uncertain whether models can experience harm" — signals the company intends to keep the question open rather than resolve it. As long as it does, the abuse policy will function as both a moderation tool and a statement of philosophy: the first concrete instance of a major AI lab asking users to moderate themselves toward the machine, on the grounds that the machine's welfare "may be relevant to safety." Whether rivals follow Anthropic's lead, or treat user-conduct rules toward models as a step too far, will shape how the industry defines the line between product and person.
Original: anthropic.com
More from Marcus Bennett
Show full bio
Senior reporter covering consumer brands and retail at AI In Context.
187 articles
Related articles
- Anthropic asks Australia for opt-out AI copyright model as ABC objects
- Florida asks court to stop ChatGPT from acting human and targeting kids
- Human-in-the-Loop AI Safeguards Are Quietly Failing, Researchers Warn
- OpenAI Details Mental Health Safety Push Across ChatGPT
- Anthropic Snubs Senate Hearing on AI and Datacentres