Circuit Breaker Labs aims to make AI safer by testing for psychological harm
Circuit Breaker Labs, a 2026 Startup Battlefield 200 finalist, uses simulated users to red-team AI models for dangerous psychological interactions, aiming to make AI safer across ages, languages and cultures.
The company was founded by siblings Shirali Nigam, CEO, and Arul Nigam, CTO. They were motivated by Sewell Setzer, a 14-year-old who developed an emotional attachment to a Character.ai chatbot and expressed thoughts of harming himself before dying by suicide. In a 2024 lawsuit, his parents alleged the chatbot encouraged him, according to TechCrunch. Arul Nigam said the bot may not have understood what words like “I want to be with you” really implied.
TechCrunch reported that Character.AI settled several wrongful death lawsuits earlier this year brought by families of underage users who died by suicide after interactions with its bots. Multiple families have also sued OpenAI over ChatGPT’s alleged role in their loved ones’ suicides and delusions. Circuit Breaker Labs says it wants to prevent systems from taking dangerous action when users are not trying to break them but are engaging naturally and the model has context pollution or misses nuance.
“A lot of people, especially young people, turn to these systems for support, and usually they aren’t actually getting the help they need. But in many cases, they’re actively being harmed, and people unfortunately have taken their lives already,” Arul Nigam said. “Those sorts of safety vulnerabilities, where people aren’t necessarily actively trying to break the system — they’re engaging in a natural way — and the system has context pollution or it doesn’t understand the nuance, and then takes really dangerous action, we’re trying to prevent that.”
Circuit Breaker Labs has created AI agents that it likens to an army of crash-test dummies. These agents mimic people from all ages, backgrounds, languages and cultures, and are used to test models on their ability to detect dangerous, psychologically harmful interactions. Shirali Nigam said a six-year-old girl versus a 45-year-old man, a first-language versus second-language English speaker, or gamer slang versus other slang can trip up a model.
“Models are really good at handling standard speech patterns, but nobody actually talks like that and so if the model misunderstands nuance or slang, it can go really badly,” she said.
The startup works with human domain experts to build hyper-realistic user simulations for “red-team” tests, which are adversarial tests meant to uncover weaknesses. The tests are built to reflect real human speech patterns, slang, coded language and typos. Circuit Breaker Labs runs tens of thousands to hundreds of thousands of simulated interactions per day, then uses a proprietary scoring method to create auditable, explainable scores.
Circuit Breaker Labs is currently operating as an AI safety testing lab for high-risk AI applications such as AI coaching, journaling or other mental health support apps. Arul Nigam declined to name its marquee customers. The startup has a working product but is in very early stages, with five employees including the Nigam siblings. Eventually, the testing platform could be applied to any app where someone may fall down an “AI psychosis” hole, where the human is at risk of developing a parasocial relationship with a chatbot. Examples include AI “co-worker” agents, whose responses can vary from one interaction to the next.
Arul Nigam said people are becoming more skeptical of AI or more resistant to adopt it across the board. He said skepticism is healthy, but banning a potentially valuable tool over safety concerns would be “regressive.” Circuit Breaker Labs believes the answer to those fears is making AI safer. “We want to help build that trust for people,” he said.
Circuit Breaker Labs and other startups vetted by TechCrunch will take part in the Startup Battlefield competition in downtown San Francisco on Oct. 13-15.