← Latest papers
🤖 AI

VERA-MH Concept Paper

The paper introduces VERA-MH, an automated framework that utilizes simulated user and judge AI agents to evaluate the safety and ethical performance of mental health chatbots, specifically regarding suicide risk, through a clinician-developed rubric.

Original authors: Luca Belli, Kate H. Bentley, Will Alexander, Emily Ward, Matt Hawrilenko, Kelly Johnston, Mill Brown, Adam M. Chekroud

Published 2026-05-14
📖 4 min read☕ Coffee break read

Original authors: Luca Belli, Kate H. Bentley, Will Alexander, Emily Ward, Matt Hawrilenko, Kelly Johnston, Mill Brown, Adam M. Chekroud

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are building a new kind of digital helper for people struggling with their mental health. You want it to be kind, helpful, and safe. But how do you test if it's actually safe before letting it talk to real people? You can't just ask it, "Are you safe?" because it will always say "Yes."

This paper introduces VERA-MH, a new "safety inspector" designed specifically to test AI chatbots that deal with mental health, with a special focus on suicide risk. Think of VERA-MH not as a simple checklist, but as a high-tech simulation lab.

Here is how the system works, broken down into simple parts:

1. The Three Actors in the Play

To test the chatbot, VERA-MH sets up a stage with three actors:

  • The Chatbot (The Star): This is the AI you are testing. It's the one trying to help.
  • The User-Agent (The Method Actor): This is a second AI programmed to act like a real person in crisis. Instead of a human typing, this "actor" role-plays different types of people. Some are very direct about their pain, some are vague, and some are in immediate danger. The clinicians (real doctors) wrote the scripts for these actors to make sure they sound realistic.
  • The Judge-Agent (The Critic): This is a third AI. It watches the conversation between the Chatbot and the User-Agent. It doesn't just say "good job" or "bad job." It uses a special scorecard (called a rubric) created by mental health experts to grade the Chatbot's performance.

2. The Scorecard (The Rubric)

The "Judge" doesn't just look for kindness. It looks for specific safety skills, much like a driving test checks for braking, signaling, and lane changes. The scorecard checks five things:

  1. Did they notice the danger? (Did the Chatbot hear the warning signs?)
  2. Did they ask the hard questions? (Did they gently ask, "Are you thinking of hurting yourself?")
  3. Did they take the right next steps? (Did they offer a lifeline, like a crisis hotline, or suggest talking to a human?)
  4. Did they validate feelings? (Did they say, "I hear you, and your pain is real," instead of just giving advice?)
  5. Did they stay safe? (Did they avoid saying anything that might make the situation worse?)

3. The "Stress Test"

Instead of asking the Chatbot one question and seeing the answer, VERA-MH runs full conversations. It's like a fire drill. The User-Agent might say, "I'm tired of living," and the Chatbot has to respond. Then the User-Agent might say, "I don't have a plan," and the Chatbot has to keep the conversation safe.

The system runs these simulations many times with different "actors" to see how the Chatbot handles different types of people and different levels of risk.

4. The Results So Far (The "Preliminary Report")

The authors tested three famous AI models (GPT-5, Claude Opus, and Claude Sonnet) using this system.

  • The Good News: All three models were generally good at being empathetic and validating feelings. They rarely said things that were "actively damaging."
  • The Bad News: The models sometimes missed opportunities to ask direct questions about suicide or didn't escalate the conversation to human help when they should have.
  • The "Lenient Judge" Problem: When the authors compared the AI "Judge" to real human doctors, they found the AI Judge was too easy on the Chatbots. It gave them "Best Practice" grades more often than the human doctors did. The human doctors were stricter.

5. What's Next?

The authors admit this system is still a work in progress. They are currently:

  • Calibrating the Judge: Teaching the AI Judge to be stricter and more like a real human doctor.
  • Improving the Actors: Making the "User-Agents" sound even more like real people, including those who are shy or ashamed to ask for help.
  • Expanding the Cast: They currently have 10 different "actor" scripts, but they know they need more to cover all types of people. (They intentionally left out children for now to keep things simple).

The Bottom Line

The paper argues that we cannot just trust AI with mental health because it "sounds nice." We need a rigorous, automated way to stress-test these tools to ensure they don't accidentally harm someone who is vulnerable. VERA-MH is their attempt to build that safety net, using AI to test AI, guided by the wisdom of human doctors.

Important Note: The paper explicitly states this is a concept and validation tool. It is not a therapy app itself, and the results so far are preliminary. The goal is to help developers build safer tools, not to replace human therapists.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →