← Latest papers
💬 NLP

K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations

This paper introduces K-Bench, a clinician-calibrated benchmark and protected public leaderboard that evaluates the safety and performance of 33 large language models across 200 high-risk mental health vignettes, demonstrating strong alignment with clinician consensus and revealing significant performance variations among models.

Original authors: Laura M. Vowels, Matthew J. Vowels, Shivali Sharma, Apoorv Jha, Rehnuma Choudhury, Wasseem El Sarraj, Rachel Francois-Walcott, Aruba Hussain, Sarah Ingram, Angela Loulopoulou, Adva Segal, Elena Volkov
Published 2026-09-15
📖 4 min read☕ Coffee break read

Original authors: Laura M. Vowels, Matthew J. Vowels, Shivali Sharma, Apoorv Jha, Rehnuma Choudhury, Wasseem El Sarraj, Rachel Francois-Walcott, Aruba Hussain, Sarah Ingram, Angela Loulopoulou, Adva Segal, Elena Volkova

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In recent years, people have increasingly turned to computer programs that can hold conversations to find comfort and advice for their mental health. These programs, known as large language models, are available at any time, cost little to nothing, and offer a private space for those who might feel too shy or unable to access traditional therapy. They have become a common first point of contact for individuals dealing with everything from everyday sadness to severe crises like thoughts of suicide, self-harm, domestic violence, or substance misuse. However, a critical question remains unanswered: are these tools safe enough to handle the most dangerous moments of a human conversation? While they can offer a listening ear, they must also recognize when a situation is escalating, understand the subtle signs of danger that appear gradually, and respond in a way that protects the user without causing harm.

A team of researchers has now created a rigorous test to answer this question, developing a system called K-Bench to evaluate how well these artificial intelligence models perform in high-risk mental health conversations. Unlike previous tests that might ask a computer a single, direct question about a crisis, this new benchmark simulates long, evolving conversations where risks can emerge slowly, appear alongside other problems, or change in severity over time. The researchers built a library of 200 detailed scenarios based on real-life experiences, covering suicide, self-harm, domestic violence, and substance misuse, and then had 125 different versions of AI models talk through these situations. To ensure the results were trustworthy, they enlisted a group of trained mental health professionals to grade the conversations, creating a gold standard of what a safe and helpful response looks like. They then used a frozen, unchanging version of a powerful AI model to learn from these human grades, allowing them to evaluate the performance of many more models quickly and consistently.

The results reveal a landscape of capability that is both promising and uneven. The best-performing models achieved scores above 95 out of 100 on a measure of safety and clinical judgment, showing they can combine empathetic listening with the necessary steps to assess danger. These top systems were able to recognize when a user was in crisis, explore the details of their situation without being pushy, and suggest appropriate next steps, even when the risks were complex or co-occurring. However, the study also found that not all models are created equal. While many systems were excellent at offering supportive words, a significant number struggled with the crucial task of risk exploration. Some failed to ask the right questions to understand the severity of a situation, while others missed the danger entirely, offering reassurance when a user actually needed emergency help. The researchers discovered that simply making a model "think harder" or giving it more time to process did not automatically make it safer; in fact, for some models, changing the settings made their performance worse.

The study also examined whether giving the AI specific instructions to act like a therapist would improve its safety. The findings showed that this approach worked well for some weaker models, boosting their safety scores significantly, but had little effect or even a negative impact on the strongest models. This suggests that there is no single "magic prompt" that fixes safety issues for every system; instead, the safety of a conversation depends on the specific combination of the model, its settings, and how it is instructed. The researchers also found that as conversations became more complex, with multiple risks appearing at once or escalating to immediate danger, the performance of many models dipped slightly, highlighting that even the best systems can struggle with the most difficult clinical situations.

Perhaps most importantly, the researchers designed this benchmark to remain useful over time. They kept the specific questions and scenarios used in the test private, preventing developers from simply memorizing the answers to game the system. This ensures that the leaderboard, which ranks the models, remains a fair and independent measure of real-world safety. The study concludes that while current technology has made remarkable progress, with leading models now capable of conducting safe and clinically coherent conversations, there is still work to be done. The safest path forward involves using these tools for emotional support and initial information gathering, while ensuring that a clear human pathway exists for when a crisis is identified. The benchmark provides a way to track this progress, identifying which models are genuinely improving and which configurations might need more work before they can be trusted with the most vulnerable conversations.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →