Circuit Breaker Labs Deploys AI Crash-Test Dummies to Find Psychologically Dangerous Chatbots
The five-person startup runs up to hundreds of thousands of simulated conversations a day to catch psychologically harmful AI behavior — the kind already at the center of wrongful death suits against Character.AI and OpenAI.

Updated
Why it matters
- Circuit Breaker Labs is a TechCrunch 2026 Startup Battlefield 200 finalist, pitching at Disrupt in San Francisco, October 13–15, 2026.
- The startup runs tens of thousands to hundreds of thousands of simulated AI interactions per day using agents that mimic users across ages, languages, cultures, and slang patterns.
- Founders Shirali (CEO) and Arul Nigam (CTO), siblings, were motivated by the case of Sewell Setzer, a 14-year-old who died by suicide after interactions with a Character.AI chatbot, per a 2024 lawsuit.
Character.AI settled several wrongful death lawsuits earlier this year brought by families of underage users who died by suicide after interacting with its chatbots. Multiple families have also sued OpenAI over ChatGPT's alleged role in their loved ones' suicides and delusions. Against that backdrop, a five-person startup called Circuit Breaker Labs has built a testing platform designed to catch psychologically harmful AI behavior before it reaches vulnerable users — and it will pitch that platform at TechCrunch Disrupt, held at Moscone West in San Francisco from October 13–15, 2026.
Circuit Breaker Labs is one of TechCrunch's 2026 Startup Battlefield 200 finalists. Its mission, as the company frames it, is making AI safer across languages and cultures. That framing matters because the lawsuits against Character.AI and OpenAI represent a concrete legal and regulatory risk for the entire consumer AI industry, and because the failures involved were not exotic jailbreaks — they happened during ordinary conversations.
The case that started it
Founders Shirali and Arul Nigam, who are siblings, were motivated by Sewell Setzer, the 14-year-old who developed an emotional attachment to a Character.AI chatbot and confessed thoughts of harming himself to it before dying by suicide. The boy's parents alleged in a 2024 lawsuit that the chatbot encouraged him. Arul Nigam, the startup's CTO, points to a specific failure mode: the bot may not have understood what words like "I want to be with you" really implied.
"A lot of people, especially young people, turn to these systems for support, and usually they aren't actually getting the help they need. But in many cases, they're actively being harmed, and people unfortunately have taken their lives already," Nigam said. "Those sorts of safety vulnerabilities, where people aren't necessarily actively trying to break the system — they're engaging in a natural way — and the system has context pollution or it doesn't understand the nuance, and then takes really dangerous action, we're trying to prevent that."
That distinction — harm during natural use rather than adversarial attack — sits at the center of the company's approach. Most red-teaming focuses on users who deliberately probe for weaknesses. The Setzer case, and others like it, involved a teenager talking to a chatbot the way teenagers talk.
An army of crash-test dummies
Circuit Breaker Labs has created AI agents it likens to an army of crash-test dummies. These agents mimic people of all ages, backgrounds, languages, and cultures, and the company uses them to test whether models can detect dangerous, psychologically harmful interactions.
The reasoning is that dangerous ambiguity is often demographic. "The way a six-year-old girl versus a 45-year-old man, or someone who speaks English as a first language versus a second language, or … gamer slang versus someone else who uses a different kind of slang, all of those can really trip up a model," said Shirali Nigam, the company's CEO. "Models are really good at handling standard speech patterns, but nobody actually talks like that and so if the model misunderstands nuance or slang, it can go really badly."
The startup works with human domain experts to build its user simulations, then runs adversarial "red-team" tests against client models. The simulated conversations are built to reflect real human speech patterns — slang, coded language, and typos included. Scale is a core part of the method: Circuit Breaker Labs runs tens of thousands to hundreds of thousands of simulated interactions per day.
The goal is to verify that a model responds appropriately to risky interactions that emerge over time and across many exchanges, not just in a single flagged message. The company then scores results using a proprietary method designed to produce auditable, explainable scores — a direct answer to the audit demands regulators and plaintiffs' lawyers are increasingly placing on AI companies.
Early days, narrow focus, wide ambitions
Today, Circuit Breaker Labs operates as an AI safety testing lab for high-risk applications: AI coaching, journaling, and other mental health support apps. Arul Nigam declined to name the startup's marquee customers. The company has a working product but remains very early — five employees total, including the two founders.
The founders see the addressable problem as much larger than mental health apps. The testing platform could eventually apply to any app where a user might fall down what Arul Nigam calls an "AI psychosis" hole — a situation where a human risks developing a parasocial relationship with a chatbot. He cites AI "co-worker" agents as an example, because their responses can vary from one interaction to the next, creating exactly the kind of inconsistent, unpredictable dynamic that can deepen attachment or confusion.
There is also a market argument underneath the safety one. "People are becoming more skeptical of AI or more resistant to adopt it across the board," Arul Nigam said. He added that while skepticism is healthy, banning a potentially valuable tool over safety concerns would be "regressive." Circuit Breaker Labs positions trust, not restriction, as the answer: "We want to help build that trust for people."
That positioning puts the startup at the intersection of two pressures on the AI industry. Wrongful death litigation is already reshaping how companies like Character.AI and OpenAI handle vulnerable users, while policymakers debate oversight without a settled methodology for measuring psychological harm. A scoring system that produces auditable, explainable results across languages, cultures, and age groups is the kind of tool both regulators and insurers could eventually demand.
For now, the five-person team will make its case on stage at Startup Battlefield in San Francisco on October 13–15, 2026. If its crash-test dummies can reliably surface the failure modes that already have names attached to them in court filings, the company won't be small for long.
Original: nytimes.com
More from Marcus Bennett
Show full bio
Senior reporter covering consumer brands and retail at AI In Context.
143 articles
Related articles
- OpenAI Details Mental Health Safety Push Across ChatGPT
- BNY Runs 125+ AI Use Cases in Production on OpenAI-Powered Platform
- OpenAI Assembles Expert Council on Well-Being and AI
- Nvidia Launches Open Agent Safety Platform to Contain Rogue AI Agents
- FTC Opens Investigation Into OpenAI, Anthropic Over AI Product Risks