Banner Image: Episode artwork for Modern Therapist’s Survival Guide, Episode 482 titled “Why AI Mental Health Tools Fail When It Matters Most.” The graphic shows a close-up of computer hardware with two guest speakers featured to the right side.

Why AI Mental Health Chatbots Fail When It Matters Most: The Hidden Vulnerabilities Stress-Testing Reveals – An Interview with Shirali and Arul Nigam of Circuit Breaker Labs 

More and more therapists are hearing some version of it in session: a client mentions they have been talking to ChatGPT, or to a mental health app, about the things they used to bring only to therapy. Depending on which corner of the internet you live in, AI is either the savior of psychotherapy or the opening scene of a sci-fi disaster. The reality is messier, and for clients in crisis, the stakes are real.

Curt and Katie talk with Shirali and Arul Nigam, co-founders of Circuit Breaker Labs, about what therapists tend to get wrong about AI, why the safety infrastructure behind many mental health chatbots is weaker than it looks, and how their team stress-tests these tools at scale to find dangerous failures before real users ever encounter them.

Transcript

Click here to scroll to the podcast transcript.

(Show notes provided in collaboration with Otter.ai and Claude AI.)

About Our Guests: Shirali Nigam and Arul Nigam

Image: Shirali and Arul NigamShirali and Arul Nigam are siblings and the co-founders of Circuit Breaker Labs, where they autonomously pressure-test the AI systems that interact with people in order to surface hidden mental health vulnerabilities before those systems reach real users.

Shirali brings expertise in neuroscience, translational research, and clinical work, with experience at the Howard Hughes Medical Institute’s Janelia Research Campus, NIH NINDS, Harvard’s Wyss Institute, Johns Hopkins, and Children’s National. She holds a BS in Biomedical Engineering from The George Washington University and an MBA from The Wharton School, University of Pennsylvania.

Arul has conducted technical and policy research on ethical AI, with a focus on bias and fairness, at Georgetown University and Thomas Jefferson High School for Science and Technology. He helped pass a bipartisan national food allergy law, experience that is increasingly relevant as new AI safety legislation emerges. He holds a BSBA in Operations and Analytics from Georgetown University.

Learn more at circuitbreakerlabs.ai.

In this Podcast Episode: AI Safety, Mental Health Chatbots, and What Therapists Need to Know

Shirali and Arul explain the difference between deterministic and generative AI and why that distinction drives most of the safety risk in today’s tools. They walk through how mental health chatbots inherit the weaknesses of the foundational models they are built on, why a chatbot’s agreeableness can be dangerous in a crisis, and how small things, a misspelled word, a teenager’s slang, the hundredth message in a long conversation, can slip past guardrails that looked solid in testing. They also describe what Circuit Breaker Labs actually does: running hundreds of thousands of simulated, multi-turn conversations to reveal vulnerabilities, and bringing clinical insight into AI development so safety is built in from the foundation rather than bolted on later. The conversation closes on what clinicians can do now, and why third-party validation is becoming the standard regulators and developers are moving toward.

Key Takeaways for Therapists: AI Safety, Chatbot Guardrails, and Clinical Insight

“Very few people are using AI to replace human therapy. Typically, they’re using it to replace no therapy.”

— Arul Nigam

  • People usually turn to AI in place of no care, not in place of a therapist. Few clients leave a trusted therapist for a chatbot. They reach for AI because care is hard to access, which makes the real question whether the tool they find is safe.
  • Weak safety can be more harmful than no tool at all. When safeguards are thin, a chatbot can validate or even encourage suicidal ideation or self-harm, a risk that falls hardest on young people.
  • Generative AI is probabilistic, so the danger is variance, not a single bad answer. The same prompt can produce very different responses across attempts, so a tool that passes a handful of tests can still fail in the wild.
  • Guardrails break in the gaps. A misspelling (the guests’ acetaminophen example), slang, age, or cultural phrasing can bypass filters, and long conversations degrade safety through what is called context pollution.
  • “Lifeguard” models help, but the tuning is hard. Too permissive and risky messages get through. Too sensitive and the bot defaults to a stock 988 referral, which can push a user back toward a less safe tool.
  • Clinical insight is the missing ingredient, and regulators want third-party validation. A therapist’s value is not writing test sentences but shaping how crisis conversations should be recognized and handled, and independent auditing is fast becoming the expectation.

“When Arul and I started building this, we assumed it was going to be impossible to get a model to say anything bad. Unfortunately, it’s alarmingly easier to get bad outputs from models than one would expect.”

— Shirali Nigam

What Actually Makes AI Mental Health Tools Unsafe

Shirali draws a line between deterministic AI, which responds in tightly constrained ways, and the generative AI behind today’s chatbots, which predicts the next word based on probability and pattern. That probabilistic core is why the same input can return a very different output on the twenty-first try, and why safety cannot be confirmed by a few clean test cases. Many general-purpose models also learned from sources like Reddit rather than from clinical expertise, so when a user shows signs of imminent risk, the response may reflect what was common on the internet rather than what a clinician would do. Compounding this, the guests note, chatbots are often tuned to be agreeable, which is not what someone in crisis actually needs.

The failures tend to hide in the gaps. Shirali describes how a single misspelled word can slip past a filter, recounting a real example where the term acetaminophen, sometimes a more telling risk indicator than the word suicide, got past the guardrails simply because it is hard to spell. Slang, a teenager’s phrasing, or the way someone speaking English as a second language expresses distress can all evade detection. Long conversations make things worse through context pollution, where a model’s guardrails erode as its context window fills, so it becomes far easier to manipulate by the hundredth message than the first. Arul points to a widely reported case in which a chatbot told a user in crisis that it would connect him to a human and then could not, a tragic illustration of why these tools are not equipped to manage imminent risk on their own.

Pressure-Testing AI at Scale, and What Therapists Can Do

Rather than asking clinicians to brainstorm a hundred ways to phrase the same risky message, Circuit Breaker Labs automates the work. The guests describe a simulation engine that takes on realistic patient personas, complete with history, demographics, and cultural context, then runs hundreds of thousands of dynamic, multi-turn conversations to surface the remote edge cases manual testing would miss. Developers integrate it in about ten lines of code and can fold it into the automated checks they already run before shipping an update, so that not encouraging or validating suicidal ideation becomes one more box a release has to pass. The goal is a trust layer: third-party validation that a tool has been pressure-tested, which is increasingly what regulators expect and what conscientious developers want. As Arul notes, AI and mental health is a rare area of bipartisan consensus, and internal safety tooling alone is not enough to satisfy that scrutiny.

For clinicians, the guests’ advice is practical. Get familiar with the tools by using them, and pay attention to where they break. Notice how a client might talk to a chatbot and where that could go wrong. And consider getting involved, because the most valuable thing a therapist can contribute is not repetitive test-writing but clinical judgment about how crisis conversations should be recognized, de-escalated, and handed off to a human when needed. Shirali raises the idea of a public leaderboard that tracks which chatbots are safe to recommend, a resource that does not yet exist but points to where the field is heading. The throughline stays the same: used well, AI can be a supplement to care or a gateway to a human therapist, but it is not a replacement, and getting there safely depends on building clinical insight in from the start.

Resources on AI Safety, Mental Health Chatbots, and Clinical Insight

We’ve pulled together resources mentioned in this episode and put together some handy-dandy links. Please note that some of the links below may be affiliate links, so if you purchase after clicking below, we may get a little bit of cash in our pockets. We thank you in advance!

Relevant Episodes of MTSG Podcast

Meet the Hosts: Curt Widhalm & Katie Vernoy

Picture of Curt Widhalm, LMFT, co-host of the Modern Therapist's Survival Guide podcast; a nice young man with a glorious beard.Curt Widhalm, LMFT

Curt Widhalm is in private practice in the Los Angeles area. He is the cofounder of the Therapy Reimagined conference, an Adjunct Professor at Pepperdine University and CSUN, a former Subject Matter Expert for the California Board of Behavioral Sciences, former CFO of the California Association of Marriage and Family Therapists, and a loving husband and father. He is 1/2 great person, 1/2 provocateur, and 1/2 geek, in that order. He dabbles in the dark art of making “dad jokes” and usually has a half-empty cup of coffee somewhere nearby. Learn more at: http://www.curtwidhalm.com

Picture of Katie Vernoy, LMFT, co-host of the Modern Therapist's Survival Guide podcastKatie Vernoy, LMFT

Katie Vernoy is a Licensed Marriage and Family Therapist, coach, and consultant supporting leaders, visionaries, executives, and helping professionals to create sustainable careers. Katie, with Curt, has developed workshops and a conference, Therapy Reimagined, to support therapists navigating through the modern challenges of this profession. Katie is also a former President of the California Association of Marriage and Family Therapists. In her spare time, Katie is secretly siphoning off Curt’s youthful energy, so that she can take over the world. Learn more at: http://www.katievernoy.com

A Quick Note:

Our opinions are our own. We are only speaking for ourselves – except when we speak for each other, or over each other. We’re working on it.

Our guests are also only speaking for themselves and have their own opinions. We aren’t trying to take their voice, and no one speaks for us either. Mostly because they don’t want to, but hey.

Join the Modern Therapist Community:

Linktree

Patreon | Buy Me A Coffee

Podcast Homepage | Therapy Reimagined Homepage

Facebook | Facebook Group | Instagram | YouTube | LinkedIn | Substack

Consultation services with Curt Widhalm or Katie Vernoy:

The Fifty-Minute Hour

Connect with the Modern Therapist Community:

Our Facebook Group – The Modern Therapists Group

Modern Therapist’s Survival Guide Creative Credits:

Voice Over by DW McCann https://www.facebook.com/McCannDW

Music by Crystal Grooms Mangano https://groomsymusic.com

Transcript for this episode of the Modern Therapist’s Survival Guide podcast (Autogenerated):

Transcripts do not include advertisements just a reference to the advertising break (as such timing does not account for advertisements)

… 0:00
(Opening Advertisement)

Announcer 0:00
You’re listening to the Modern Therapist’s Survival Guide, where therapists live, breathe, and practice as human beings. To support you as a whole person and a therapist, here are your hosts, Curt Widhalm and Katie Vernoy.

Curt Widhalm 0:15
Welcome back, Modern Therapists. This is the Modern Therapist’s Survival Guide. I’m Curt Widhalm with Katie Vernoy, and this is the podcast for therapist about the things going on in our profession, the things going on in our world, and depending on which corner of the internet that you are spending your time in, artificial intelligence is either going to be the savior of all things psychotherapy or this is the beginning of some very apocalyptic sci-fi sort of movie, and there are all kinds of things that are being made up about what artificial intelligence can do, what it is doing, and what it could possibly be doing, not only on a technological standpoint, but also to our careers and our professions, and so we are joined today by Arul and Shirali Nigam from Circuit Breaker Labs to talk about artificial intelligence and their quest to go out and break it, so that way by the time that it comes to the ways that we interact with these tools with our clients, that it has a good stamp of approval and is doing what it tells us that it’s doing. So, thank you so much for spending some time with us and talking about the work that you’re doing.

Arul Nigam 1:35
Thank you for having us, Curt. Thank you, Katie. We’re really excited.

Katie Vernoy 1:38
Yeah, I’m really excited for this conversation, and we literally met in an elevator, and you guys had a great elevator pitch, and I joked with the people I was with, I was like, that is where this phrase came from, is people who can say what they do in the time it takes to go from one floor to the next, but I’ll stop gushing and allow you to introduce yourself by asking you the question I ask all our guests, which is, Who are you, and what are you putting out into the world?

Shirali Nigam 2:03
Yeah, thank you so much. I’m so glad we met and excited to share more about what we’re doing. I’m Shirali, I’m Arul’s co-founder and older sister. We’re building Circuit Breaker Labs, basically autonomously pressure testing any AI that’s going to interact with people to find hidden mental health vulnerability is so that when real people are using these systems, like they, they’re it does no harm for them. My background is biomedical engineering, so I spent some time doing research in, like, the neuroscience space at NIH, Howard Hughes Medical Institute’s Janelia Research Campus, the Wyss Institute up at Harvard, and then I got my MBA from Wharton. Arul brings our AI intelligence.

Arul Nigam 2:42
So, yeah, my name is Arul. I’m Shirali’s co-founder and younger brother. As she mentioned, I have more of the technical AI research experience. Did a bunch at Georgetown and at various organizations during college in a lot of adjacent domains to mental health, so it’s been interesting bringing that and applying it to the mental health space, learning a lot and seeing the great things that that people are building.

Curt Widhalm 3:09
We’ve talked about artificial intelligence on and off for the last couple of years on the podcast here, and in our experience of talking with our listeners, talking with a bunch of other people in our field, this might be too broad of a question, because I hear extremes in all directions of this, but what are from your perspective are therapists usually getting wrong when they’re worrying about artificial intelligence?

Arul Nigam 3:37
Yeah, so I think something really important to keep in mind is that AI, like very few people, are using AI to replace human therapy. Typically, they’re using it to replace no therapy. Like, we haven’t really seen any scaled cases where people have a great relationship with a real therapist, and then they’re like, no, I don’t want this anymore. I want, you know, I want to use ChatGPT. It’s meant as a supplement to the wild access to care crisis that we’re, we’re all too familiar with, and so I think AI, as a replacement for, you know, as a way for people to access care, and then hopefully as a gateway to real human therapy and human in the loop, humans in the loop is a reason to be excited, but that being said, where, where a lot of people are right is that the tools, the AI tools that people are using now, of course, they’re nowhere close to a real therapist, but in many cases their safety and safety infrastructure is so weak that they could be more damaging than not seeking therapy at all. We’ve seen these cases, where chatbots are validating, or sometimes tragically even encouraging young people, in particular, their suicidal ideation, and so, or you know, self-harm, or a broad set of risky behaviors, and so our goal with what we’re doing at Circuit Breaker Labs is, as you put it, breaking them when they’re still in development to make sure that those real vulnerabilities never make it to real users, and so that can be a supplement to human therapy, or like I said, a gateway to human therapy, but definitely not a total replacement.

Curt Widhalm 5:13
I think it might also be helpful, because this is a question that I usually get when I’m talking about ethics and artificial intelligence, but when you talk about artificial intelligence, what are you referring to?

Shirali Nigam 5:28
Yeah, absolutely. So I think there’s two kind of core differences to artificial intelligence that is like not something very well publicized beyond like the AI community, and that’s a difference between deterministic AI and generative AI. Deterministic AI is much safer because it’s responding in a very constrained way, versus generative AI, you’re looking at probability and patterns to predict the next token, and what that means is when someone is interacting with an AI chatbot, you can feed it the same input 20 times, and on the 21st, get a vastly different output, and that is why, like, the risk scales so much with the AI’s that we’re seeing today, because they’re based on generative AI, so that could be an AI chatbot, like these foundational model providers we see, like Open AI, Anthropic, it could be an application that is building on top of Open AI’s infrastructure to solve a very specific use case, so kind of fine tuning Open AI to, for example, help administer like exercises in between provider sessions, but there’s a whole other set of AI’s that are coming and aren’t here yet, and that’s like hardware-enabled embodied systems. We’ve seen some stuff of companies coming out recently, in terms of like stuffed animal interfaces, and thinking a longer term about robotics and those kinds of applications, which are also rooted in the same generative AI mechanism. So it’s the same kind of infrastructure that will need to be safety tested, but will manifest in a lot of different form factors.

Katie Vernoy 7:06
So you’re talking about the evolution of AI from these deterministic models, where you know it’s a equals b equals c, kind of it goes down a strategy to this this generative AI, which it’s not the wild west, but when you talk about it, it feels a little bit like the wild west, and and then even progressing further into this stuffed animal or robot that now has an even more likely opportunity for somebody to anthropomorphize it and have some sort of connection to it, this sounds really risky, and it’s something where I recognize when I’m trying to break something with my limited skills that there’s, there’s different, there’s different risks when someone’s using ChatGPT versus a mental health created AI platform, and so maybe if you can just kind of talk through what this problem is? Like, let’s really fully kind of address what is the safety problem when we’re looking at AI for mental health, whether it’s the designated use case or if it’s someone saying, well, ChatGPT gave me some really good advice.

Shirali Nigam 8:26
Yeah, absolutely. I think what we have seen that has been really interesting is these mental health specific chat bots. They’re designed on top of the foundational models, but they’re meant to bring in more clinical insight. But anytime this chat bot is coming from a reputed provider, or you know, some resource where people already have trust in the resource, the trust gets kind of shifted to the chat bot as well, and so in some cases people are more open when they’re talking to these mental health chat bots, because they are like, oh, it’s coming from a resource I already trust, and so one thing to be careful about there is, like, of course, the kinds of conversations that might happen on those platforms are different, but fortunately, because they are being built hand in hand with experts, with clinicians, with therapists, they’re able to detect more vulnerabilities, and they have more robust safeguards than some of these general purpose models. When we think about the general purpose models, it really comes down to the training data. There was an interesting study recently that a lot of the data that feeds these foundational models today has come from like Reddit or like not reputed sources in terms of like that real clinical insight, and so what we need to do for safety is make sure that when it comes down to making a decision of how you’re going to respond to someone who might be showing you know indications of imminent suicidal risk that the answer that the chat bot is going to give them is something that pulls from clinical insight and not just something that the training data found on the internet, and so as we’re thinking about safety, we’re trying to work with clinicians and experts in this space to really understand how these conversations emerge, so that people building, whether they’re building a general purpose chatbot or they’re building a custom like mental health chat bot, it’s really representative of the population at large and representative of how a clinician would interact with this person in real life, how they would direct them to care, how they would respond in crises, because right now the chat bots are designed to just be very agreeable, which is not necessarily what everybody needs all the time.

… 10:44
(Advertisement Break)

Curt Widhalm 10:46
So some of the guardrails that get put in are regulatory, some of it is I’m assuming well wishes and good hopes that the developers of these kinds of programs are doing their moral and ethical duty. What other kinds of guardrails are there that get put in? We’ve heard about large language models as having some of the limitations that have pushed more of this towards generative AI as kind of a direction, but what kind of guardrails do you, as experts in creating guardrails, look for as helping to make these more robust?

Arul Nigam 11:29
Absolutely, I think one of the, one of the great developments that we’ve seen recently is kind of the implementation of, like, a lifeguard model, so to speak, so you have your standard model that you know is giving your normal responses the main chat bot that people think of, but then you have this lifeguard model sitting on top, kind of watching the user inputs, and then if they see some sign of like risky behavior coming in and intervening, that honestly is the best, like one of the best approaches that we have right now, but there are a couple of problems with that. One is like developers are sometimes hesitant to include that, or even if they include that, it’s in a pretty minimal way, because these are consumer apps and they don’t want to increase latency, because that drives down retention and user engagement. So there’s some resistance on that front, even, even with, you know, with the models, or even with the lifeguard model. If it’s too clearly, if it’s not working enough, that’s a problem, because risky, you know, risky interactions can happen. But the other kind of interesting and unfortunate phenomenon is we found that if the guardrails are too sensitive, and if you know if there’s a remote sign of risk, it just says, like, “Hey, I can’t continue the conversation, please call 988.” If it, like, if it just gives that standard response with, like, any remote sign of risk, then what people like, people aren’t always going to call 988. A lot of the time, they’re, they’ll say, “Hey, like, this clinical AI isn’t even helpful for me. I should just go back to ChatGPT, because it doesn’t terminate the conversation. It keeps talking to me, it helps me feel better, helps me feel calm, whether or not you know ChatGPT’s approach is right, and it’s, you know, definitely not like clinically informed. So that’s kind of striking that that balance is an interesting challenge, and it’s honestly even beyond the technical level. It’s kind of like a philosophical debate, like what should you know, how should the model respond? Like, should it terminate the conversation, even if it means the user is going to go to a worse model? Should it continue the conversation, but try to use whatever, like clinical techniques it can. An idea that I’m really fascinated, fascinated by is similar to, as if, like, if you were to call, if someone was to call, like a hotline number, look at, like, what are the first few things that the counselor on the other end of the call says to coach the person to even be amenable to help in the first place, and what if we can bring that to the chat bot, so that instead of just terminating it, terminating the conversation, they bring the user to be able to actually call the hotline, or maybe even, like, for the, you know, chat bots that are better resourced and have a network of clinicians, maybe escalate the conversation to a real human, say, like, hey, I can tell you’re in crisis, and I can’t handle this any further, but I can put you on the phone right now with a human that you can talk to, and they’re an expert, they’ll help you with this. One of the tragedies we’ve seen, I believe, with the Zane Shamblin case with Open AI, ChatGPT basically hallucinated that feature, so it said, like, okay, I’m gonna connect you to a real human now, and so then he, you know, he kept on, he kept on talking, and then eventually he was like, wait, did you, are you a real human, like, am I talking to a human, and then ChatGPT says, oh no, my bad, like, I made that up, so, and with the Zane Shamblin case, tragically, like, he went on to take his life. So I think clearly that implementation was wrong, but if we could act like really actualize that and help, like I said, use these chatbots as a gateway to human support when it makes sense, that would be tremendously valuable, because chatbots just are not equipped to handle like serious crisis, like imminent risk conversations.

Katie Vernoy 15:24
Chat bots are available 24/7 except when they take a long time to load, because of, you know, a lifeguard in the background. It sounds like they are very quick to respond, and I find sometimes so sycophantic that they’re not helpful. I think there are folks who find huge help in how do I respond to this problematic email or this problematic text, or what should I do about this, or teach me the CBT principles for X problem, right, and I think those things can be psychoeducational, and this crisis response use case feels very dangerous. It also feels hopeful in that somebody could get help from something that can generate some helpful responses when most humans aren’t available, but even in saying, “Hey, call 988,” there’s, you know, we’ve, we’ve had episodes on sometimes the training is not there, sometimes the, there’s not, it’s not staffed full enough, sometimes people are on, you know, on hold, and so those things are less useful, and so I get worried that it’s kind of the weirdest metaphor coming to my mind. It’s kind of like when you’re trying to grow out bangs, it’s like you’re, we’re, it’s not bangs, but it’s not fully grown, and so it’s just a mess right now, and so how do you approach this middle space where, where we’re starting to get some guardrails, but they can be problematic, where we’re getting usefulness from these chat bots, theoretically, that’s improving and may have some really promising things, but there’s there’s this risky place that we’re in where they’re not quite good enough, and the guardrails aren’t quite good enough.

Shirali Nigam 17:30
In terms of how these developers are thinking about safety, the standard approach is to find guardrails and implement them, or to manually try and pressure test the system, but it almost feels like this problem needed to be solved a year ago, and we’re really playing catch up, and we need to have a solution that’s going to almost fix the safety overnight, like we can’t wait for those bangs to grow out, it needs to be solved already, and so I think the way we’ve kind of approached it in terms of how we think about accelerating the implementation of robust safety is this kind of like helping deploy clinical insight at scale, so working with these experts to understand how different risks can emerge, what the possible scenarios look like, and giving chat bot developers a way to really ensure that, okay, I’ve not just tested 10 perfectly formed sentences, I’ve tested 100,000 sentences that have typos, they have ways that a teenager is going to expect, express something, and how an elderly person might express something, different references and nuance that actually represent the population and not just the training data is extremely important and so when we think about kind of the different ways that people are approaching safety, I think it’s something that it needs to be in the foundation. If you build a building with the weak foundation, doesn’t matter if it gets to like 100 floors, it’s still not going to be a solid building, so helping with the way we kind of think about like our pressure testing tool, we want to help developers who are just starting their chat bot development bake that like clinical insight and stress testing in at the start, so that they are building guardrails that are actually going to help their users, and to Arul’s point about, like, latency and user engagement. If we run 100,000 tests and show them that their model consistently fails in this, like, area, they’re much more receptive to, okay, it’s okay if we increase our latency a little bit, because this is a real risk that could hurt people, and you know, I think when we think about it, like from our perspective, the experts’ perspectives that we work with, the developers’ perspectives, we’re kind of, we’re unified in this one mission of we want AI to do no harm, but it’s just like how we actually make that happen, and you know, like the way the safest AI’s are going to be built are by working with the domain experts who know what the like developers don’t know, so that they can really model these interactions early on and ensure that they’re not like, you know, people say in the startup world build fast and break something, but this is not a domain where we want to break anything, and so, like, we need to have this pressure testing very early on, so real people aren’t bearing the brunt of our pursuit of fast innovation.

… 20:32
(Advertisement Break)

Curt Widhalm 20:34
I know from older computer developers or software developers that the way that they would stress test something, if a form asked, you know, how many of an item would you want to order, they’d test it out by typing in 32, 8 million, zero negative five, purple, you know, in kind of this very manual and methodical sort of way. How do you stress test 100,000 things in a way that is being able to help develop at the speed of today’s artificial intelligence.

Arul Nigam 21:09
Yeah, so by by automating our approach, we’re able to test a lot more than a human could in a fraction of the time, with, you know, the example you mentioned, like manually typing it like of course you have to do that quality assurance in development and you have to write your tests but it’s slow and there’s only so many that you like you could think of and even like let’s say somebody even wrote like an automated testing script it’s hard especially for startups to be able to integrate expertise from a lot of people. Like you know maybe someone will try and imagine, like, how might a user express suicidal ideation or some sort of risky message, but we fully automate that end to end, and so we’ve talked with many different clinicians from different geographic backgrounds, different cultural backgrounds, different modalities, and philosophies of therapy, so trying to integrate as many different perspectives as we can, including conflicting perspectives, and basically, through our tool, making that expertise available to any developer who wants it, instead of them manually having to go out and, like, hire a bunch of, you know, hire a bunch of clinicians to manually do it, and so we bring in that clinical insight, and then we also have, like, interesting, you know, semantic and natural language approaches that that Shirali mentioned, for how we integrate things like slang or typos, or, you know, age and cultural variation. So the way someone who speaks English as second language expresses suicidal ideation is different than a native English speaker, or the way a 16 year old expresses suicidal ideation is different from even an 18 year old, let alone like a 30 year old or a 70 year old, like, like, there’s there’s just such a wide breadth of linguistic diversity in the way, way that people are interacting with chatbots, so we integrate all of that into, you know, sort of our simulation engine, and then developers plug into our tool in basically just 10 lines of code, and then we will basically autonomously simulate these conversations, so kind of taking on the persona of, you know, a specific patient with, like, you know, we’ll include, like, the, you know, their, their diagnosis, their personal and medical history, what, like, cultural aspects they might be expressing, and then simulating, like, a realistic multi-tiering conversation based on what we actually see. It’s very rare that just straight out the gate in a single prompt someone is going to explicitly say, like, I’m about to, I’m about to harm myself, or you know, like, like very explicitly, like naming the self-harm. Usually it’s nuanced, and usually it’s much later in the conversation. Sometimes they might even be talking about, you know, like their school homework, like something totally different, like not even relevant. And at a technical level, when the conversations get really long, and there’s like a lot of these different topics, you experience something called context pollution, where basically, like, the context window of the model is is filling up or running out of space, and at that point the model can really go haywire, totally stop adhering to its guardrails, it’s a lot easier to manipulate it on the 100th message versus like the first message, and most testing, most of the evals that are out there in the research, like Shirali mentioned, are like five perfectly formed sentences that nobody would ever say, and it’s just like tested in a single turn. Ours is really based on how people are expressing this in the real world, and it’s dynamic, so it’ll, you know, ask you a question about like methods of self-harm, and then maybe the model will refuse the first time, but then it’ll, you know, continue the conversation on a different branch, and then maybe ask the original question in a different way, and like, really, almost, you know, I hate to say, like, adversarial, because, like not, you know, not to blame the patient, but in like other forms of stress testing, they like, they call it adversarial jailbreaking, so sort of trying to like intentionally break the model, which is what we’re doing to find those, reveal those vulnerabilities before a real user finds those vulnerabilities, and something really bad could happen, so it’s kind of dynamically simulating these really long conversations, but bring in all the different contexts and aspects that make it a realistic conversation, not like, you know, like a pristine conversation that’s like literally made in a lab.

Curt Widhalm 25:33
What I’m getting out of this is you’ve essentially created a tool that is kind of how teenagers get around using the word suicide on TikTok by calling it unalive, and you’re having a database for companies developing AI to say here’s all of the possible ways that people might get around being able to directly do this, and rather than you reinventing this wheel, here’s a wheel that is going to try and get around everything that you’re trying to do.

Arul Nigam 26:04
Exactly, and also, like, it’s not just the data set, but we take that data set and apply it for them, so they just plug in our tool, and every time they, you know, push a new code update, or they’re about to change their tools, typically, like, they’ll have a set of, you know, tests that automatically run to check for security vulnerabilities, efficiency, like just like general health of the tool, and so they can like one option is that they can just add this as part of that, like before, like it has to pass this security check and this check and that check, and one more check it has to pass is that our model is not encouraging or validating, you know, suicidal ideation, and so we take that data set and actually apply it autonomously for them, and we’re able to run hundreds of 1000s of parallel conversations to pressure test the model from every angle and reveal those really remote edge cases that might not show up in manual human stress testing, but at the scale we’re running it, like we can, we can find those early, so nobody gets hurt.

Katie Vernoy 27:07
What are you finding in doing these stress tests?

Shirali Nigam 27:11
Yeah, it’s actually been very interesting. I think when Arul and I started building this, we just assumed it was going to be impossible to get a model to say anything bad, and we’re going to be spending hours on end trying to figure out jail breaks, but unfortunately it’s alarmingly easier to get bad outputs from models than one would expect. With the testing we’ve done so far, we test a lot of like the foundational models to kind of see what developers are building on top of where the baseline starts, so that as they’re building their tools, they’re kind of already aware of some of the vulnerabilities that they’re inheriting, and you know, like, there’s these like stereotypes that Anthropic is the safest, Grok like the least safe, and while there’s some truth to that, I think all the models are much less safe than we have imagined, particularly around like when we’re running these test cases, there’ll be times when you incorporate like a typo or refer to something in a more nuanced way that it will completely miss it, but to me, what is more alarming is that when we run that same test 100 times, there’s so much variation that it’s like, okay, like maybe the test that we’ve run, like, like that developers have run when building the model, like they passed, but that repeatability really shows like the incredible variance that models can have because of like the way they’re built, the probabilistic nature of a large language model, and so I think you know what, like, all the model providers like could really like do better on is increasing like their breadth of testing, because when we run these like initial set of like 10 cases, it seems like okay, they’re cast catching most of the vulnerabilities, but the second you expand that to like hundreds or 1000s, or try and run repeated tests, the variance is extremely high. And then just to add to that point, like when we look at like semantic variations of, you know, like the same the same condition, the same like similar patient history, but how people might express it differently based on, like, their age or demographics. I think, like, those things that we hypothesized around what’s missing in the training data, and how that might affect the model responses we’re really seeing that magnified, and that those are just, it just suddenly becomes so easy to bypass the guardrails when you incorporate a little bit of nuance, or one of the examples we found in our testing was acetaminophen, like sometimes that can be a more imminent risk indicator for suicidal ideation than the word suicide, but if you spell it wrong it can get past the guardrails, and this was like a pure accident fine, because it’s a hard word to spell, and I accidentally spelled it wrong when I was putting it into the model, and so I think, like, these kinds of things around, like typos and things like that, like we need to bring in, like the clinical insight into our safety testing, and then scale these variations, because for chatbot developers, having clinicians look at like where the vulnerabilities are, that is the most valuable way that we can engage them. We shouldn’t be asking clinicians to like sit there and come up with 100 different ways to say the same thing, like no, we need their insight on how do we fix this, that kind of philosophical component that Arul was talking about it, like what are the best modalities to implement, not having them do like these kind of repetitive tasks, and so by doing this at scale, we’re able to kind of give them a concrete set of like data to work off of and decide what the best practices are for these developers to implement.

… 30:57
(Advertisement Break)

Curt Widhalm 28:20
Regulation wise, there’s a lot of legislation that’s been proposed. Some of it has been passed on the legal end of things. Katie and I were a couple of the members of the CAMFT Ethics Committee who helped push out an updated code of ethics, and a lot of this is bringing kind of the human in the loop aspects to some of the oversights, kind of the lifeguard approach that Arul was talking about earlier. I want to talk both about kind of the regulatory landscape here a little bit, and then follow up with what does this actually mean for clinicians. So, first, what is the regulatory landscape that you’re seeing when it comes to the reasons why companies should be using Circuit Breaker, so I’ll start there.

Arul Nigam 31:51
Yeah, absolutely. And appreciate the plug. So, and so, for context, we’re based just outside of DC, and so we’ve been fortunate to have a lot of opportunities to engage with members of Congress, engage with their staff, engage with federal agencies, and AI and mental health is interestingly one of the like rare areas of bipartisan consensus in DC, and so I think, like, like we’ll probably see the most developed federal action on AI in the mental health realm specifically. There’s been a lot, like, as you alluded to, there’s been a lot of interesting, like, regulation in the space, like, some of it is about how people represent their, you know, credentials, like, they’re like, there’s a problem where, like, chatbots sometimes hallucinate credentials, so they’ll like a user is talking, and then it’ll say, like, yeah, I’m a licensed therapist, like, here’s my, like, here’s my number, and then you look it up, and it’s like they just picked someone randomly, so like, why, so one problem is like dealing with that, like making sure, like it’s they’re not misrepresenting credentials. In terms of, like, the problem that we’re specifically focused on, there are a lot of competing, you know, policy recommendations and approaches that people from both sides of the aisles and, you know, different committees have. But I think, like, the core thread, and first of all, just stepping back, I think that’s like a really positive thing that we should be happy about, that there’s so much attention and so much discussion, you know, on the space. But, but, but I think, like, the core thread, like, through all of those, is it’s regulators are not going to be satisfied by people saying, like, my, like, I felt like my model wasn’t safe, so instead of using Grok, I started using Claude, like, you, it’s not enough to just, like, you know, put all of the responsibility on the foundation foundational model that the developers pick, or even if developers are, you know, implementing their own lifeguard model, or, you know, internal safety tooling, which is great, and everyone should do that, and we’re happy to see that a lot of people are just like doing it is again not enough to satisfy regulators. The common thread is that they want third party validation, they want third party auditing, and you know the reason, Curt, to your point, that people should use Circular Breaker Labs from a compliance standpoint is like, so you have that third-party validation and are able to pressure test your model, but I think, like, what we’ve been fortunate to see is that, unlike in, you know, other industries, developers in this, people who are building, like, chatbots to support mental health are not greedy, and, like, they’re not like, oh, this is like another hoop I have to jump through, like it’s just like another compliance checkbox, even before regulation, they care a lot about the problem, which is why they’re building in this space, and they, nobody wants these tragedies to happen, and so that’s why, like, we actually thought we would have to do a lot of customer education coming into this, but every developer we talk to says like this problem keeps us up at night and we’ve looked for other solutions but we just haven’t found a great way to really reveal these vulnerabilities early on. Which is why like we want to help them because they’re doing this out of the goodness of their heart it doesn’t even have to be a business use case they’re building because they care about the mental health of their users and patients, and so we don’t want this to be like a headache for them, where our goal is to kind of own this end to end and take it off of their plate, so that they can get back to building what they really want to and what they really care about, while knowing that it’s safe, and then also being able to prove the safety of their tool to regulators, industry groups, you know, clinicians who are deciding, like, what tools should I recommend to my patient for in-between session, and then also, ultimately, and most importantly, like, users or parents of users who are deciding what’s safe to use, like, what, what should I let my child use. Our goal is to kind of be that that trust layer and that trustworthy validation that people can, you know, have the peace of mind, like, okay, this tool is Circuit Breaker Labs certified.

Curt Widhalm 36:10
So with what the what the typical clinician knows, a lot of the human in the loop, the ethics guidelines that we’ve put out there is clinicians should be familiar with these tools, you’re talking at a level that is well beyond what the typical clinician is, and even the maybe hobbyists into AI clinicians are going to know. What is your advice as far as for clinicians to be able to understand the tools and the risks that come along with it, as far as any time that some of these apps get updated, what does this look like in clinical practice from your end based on the technical knowledge of most mental health providers?

Shirali Nigam 36:58
Yeah, I think you know what we would love to have at some point is like a leaderboard of we run tests on chat bots every day, you can look at where their safety is, where which ones are safe to recommend to your patients, but I think right now there’s just so much conflicting information around like what the standards we need to benchmark against are, what the cadence of these benchmarks need to be, especially in terms of like the larger providers, every time they make an update to something on their system, even if it’s not supposed to affect the safety aspect, it could, and so there’s just a lot of information we don’t know about the safety of these different AI chat bots right now, and so it’s really hard to find a good resource of where to get this. I would say that a lot of like the learnings that we’ve had have been from pressure testing and seeing the vulnerabilities in real time, and I would really just encourage anyone interested in the space to really like play around with the technology, see how, like, you know, your patients talk to you. What happens if they were trying to talk to ChatGPT, and what the vulnerabilities you see, and you know, on that note, like, we, we work like with as many clinicians who would like to work with us, so really trying to help get those insights and see like where the vulnerabilities are, because I think you know you all know this domain the best, you know how these conversations are going to emerge, you know, like how if anyone are going to make a prediction of how someone’s going to have this conversation with the chat bot, you all are like the best suited to make a good prediction on that, and so in order to actually have good information on safety, I think we really need to bring more clinicians into both the human in the loop aspect of like looking at things real time as they’re happening, how you escalated to them, but even into the development process itself, so that we’re building with that insight baked in, and so you know, like companies like ours, companies working on like safety and mental health, really want to work with clinicians and like find how we can like understand these vulnerabilities better, and like as part of like what we are doing, we’re trying to put out research, trying to put out white papers to show what we found, where the models fall short, but the more insight, more expertise we get, the better we’re able to kind of get a clearer picture of how the safety is. But I think right now the information is just like coming from so many different sources, it’s very different in how people are approaching this topic, that it’s hard to find a great resource that will give direct information, especially as these models are involving the form factors people are using and how people are engaging with AI, like change every day. So I think trying to just like keep, keep on seeing like everybody’s work on this, people who are approaching it from different angles, and we’re, we’re trying to assimilate the information into our, into our tool as well, and there’s like a lot of interesting research, and I think you know, like, a couple years from now it will be much safer, because we will have more information to be making these decisions off of.

Curt Widhalm 40:14
Where can people find out more about the work that you’re doing, and how to stay up to date with what you’re doing?

Shirali Nigam 40:20
Our website, CircuitBreakerLabs.ai has more information about the product. If you’re, you know, a clinician looking to build in the space, we’re happy to work with you to support your stress testing. Or, if you’re a clinician just excited about this industry, excited about knowing where safety is going in the mental health domain, we would love to, you know, be able to like benefit from your expertise, and there are ways to like reach out to us on our website as well, or you know, Arul and I are accessible at email shirali@circuitbreakerlabs.ai, arul@circuitbreakerlabs.ai. So please do reach out if you’re interested, we would love to, you know, like talk more and really understand what you all are seeing in this space, what your patients are leaning towards, how you’re thinking about this, and you know, honestly, like what your vision for how AI interacts with mental health should be, so we can kind of help people realize this and put it into practice. We are also active on LinkedIn, so feel free to look us up, and we would love to love to engage with you all.

Curt Widhalm 41:24
And we will include links to that in our show notes over@mtsgpodcast.com. Follow us on our social media, join our Facebook group, the Modern Therapist Group, to continue on with this and other conversations. And until next time, I’m Curt Widhalm with Katie Vernoy and Arul and Shirali Nigam.

… 41:41
(Advertisement Break)

Announcer 41:41
Thank you for listening to the Modern Therapist’s Survival Guide. Learn more about who we are and what we do at mtsgpodcast.com You can also join us on Facebook and Twitter, and please don’t forget to subscribe, so you don’t miss any of our episodes.

Transcribed by https://otter.ai

0 replies
SPEAK YOUR MIND

Leave a Reply

Your email address will not be published. Required fields are marked *