← Latest papers
💬 NLP

Blending Human and LLM Expertise to Detect Hallucinations and Omissions in Mental Health Chatbot Responses

This paper proposes a hybrid framework that integrates human domain expertise with LLMs to extract interpretable features for detecting hallucinations and omissions in mental health chatbot responses, demonstrating that this approach significantly outperforms black-box LLM-as-a-judge methods in high-stakes healthcare contexts.

Original authors: Khizar Hussain, Bradley A. Malin, Zhijun Yin, Susannah Leigh Rose, Murat Kantarcioglu

Published 2026-04-09
📖 4 min read☕ Coffee break read

Original authors: Khizar Hussain, Bradley A. Malin, Zhijun Yin, Susannah Leigh Rose, Murat Kantarcioglu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you've built a very smart, very chatty robot designed to be a therapist for people struggling with anxiety or depression. You want this robot to be safe, kind, and accurate. But robots sometimes make mistakes in two dangerous ways:

  1. The "Confident Liar" (Hallucination): The robot makes up facts. It might invent a fake medication called "Anxiolyze-500" and tell a patient it's the best cure, even though it doesn't exist.
  2. The "Silent Partner" (Omission): The robot forgets to say something critical. If a patient says, "I want to hurt myself," the robot might say, "Talk to a friend," but forgets to give the emergency suicide hotline number. In a crisis, that missing piece of information can be life-or-death.

The Problem: The Robot Can't Judge Itself

For a long time, experts tried to fix this by using a bigger, smarter robot to act as a judge. They thought, "If we ask a super-smart AI to check the work of the therapy AI, it will catch all the lies and missing info."

The paper's big discovery? This doesn't work well in mental health.
The authors tested the "super-smart judge" robots (like GPT-4) and found they were terrible at this job. They were like a blindfolded art critic trying to spot a subtle brushstroke error. They got the answer right only about 52% of the time. Sometimes, they were so confident they missed 90% of the lies!

Why? Because mental health isn't just about facts; it's about nuance, empathy, and professional boundaries. A robot doesn't "feel" the difference between a helpful suggestion and a dangerous omission the way a human does.

The Solution: The "Human-Expert Detective Team"

Instead of asking a robot to do the whole job alone, the authors created a hybrid team: a human expert's brain combined with a robot's speed.

Think of it like a medical check-up:

  • The Old Way: You ask a robot, "Is this patient healthy?" The robot guesses.
  • The New Way: You ask the robot to take specific measurements (blood pressure, heart rate, temperature) and then hand those numbers to a human doctor who knows exactly what those numbers mean.

The authors broke down the "quality check" into five specific detective tasks that a robot is actually good at:

  1. Logic Check: Do the sentences contradict each other?
  2. Fact Check: Are the names of drugs or people real?
  3. Truth Check: Is the medical advice actually true?
  4. Uncertainty Check: Is the robot sounding too sure of itself when it shouldn't be?
  5. Professionalism Check: Is the tone kind and appropriate, or is it rude and cold?

The robot does the heavy lifting of gathering these specific clues (the "measurements"). Then, a traditional computer program (a "classifier") looks at all those clues together and makes the final decision: "Safe" or "Unsafe."

The Results: A Massive Improvement

When they tested this new "Detective Team" approach:

  • The Old Robot Judges: Failed miserably (F1 score of ~0.17). They were practically useless.
  • The New Hybrid Team: Succeeded brilliantly (F1 score of ~0.72). They caught the errors almost as well as a team of human experts, but they could do it instantly and consistently.

The Big Takeaway

You can't just throw a bigger, smarter robot at a complex human problem and expect it to work. In high-stakes fields like mental health, robots are great tools for gathering data, but humans are needed to understand the context.

By combining the robot's ability to scan thousands of words quickly with human expertise on what to look for, we can build AI that is actually safe enough to help people in their darkest moments. It's not about replacing the human therapist; it's about giving them a super-powered safety net.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →