← Latest papers
💬 NLP

LLAMADRS: Evaluating Open-Source LLMs on Real Clinical Interviews--To Reason or Not to Reason?

The paper introduces LlaMADRS, a benchmark for evaluating open-source LLMs on real psychiatric interviews, demonstrating that structured "Item-then-Sum" strategies and well-designed prompts often outperform direct scoring or reasoning-augmented models in clinical assessment tasks.

Original authors: Gaoussou Youssouf Kebe, Jeffrey M. Girard, Einat Liebenthal, Justin Baker, Fernando De la Torre, Louis-Philippe Morency

Published 2026-04-23
📖 5 min read🧠 Deep dive

Original authors: Gaoussou Youssouf Kebe, Jeffrey M. Girard, Einat Liebenthal, Justin Baker, Fernando De la Torre, Louis-Philippe Morency

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a group of very smart, but very different, robots how to act like a psychiatrist. Specifically, you want them to listen to a conversation between a patient and a doctor and then fill out a specific checklist called the MADRS (a 10-item scale used to measure how depressed someone is).

This paper, titled LLAMADRS, is the report card on how well 25 different open-source AI models did this job. The researchers wanted to answer one big question: Do these robots need to "think out loud" (reason) to do a good job, or is just giving them a clear checklist enough?

Here is the breakdown using some everyday analogies:

1. The Setup: The "Robot Psychiatrist" Exam

The researchers used real recordings of actual therapy sessions (541 of them!). They asked the AI models to listen to these conversations and score the patient on 10 different symptoms (like "sadness," "lack of sleep," or "suicidal thoughts").

  • The Goal: Get the score as close as possible to what a human doctor would give.
  • The Contenders: 25 different AI models, ranging from tiny (0.6 billion parameters) to massive (400 billion parameters). Some are "Standard" models (just answer the question), and some are "Reasoning" models (they are told to show their work, like a student writing out math steps).

2. The Big Discovery: The "Lego vs. The Whole House" Strategy

The researchers tested two ways to get the final score:

  • Strategy A (Direct Total Score): Ask the AI, "Listen to the whole conversation and give me one final number."
    • The Analogy: This is like asking a student to look at a messy room and instantly guess the total number of toys without counting them one by one. It's hard, and they often get it wrong.
  • Strategy B (Item-then-Sum): Ask the AI to score each of the 10 symptoms individually, then add them up at the end.
    • The Analogy: This is like telling the student: "Count the red blocks, then the blue blocks, then the green blocks. Now, add those three numbers together."

The Result: Strategy B (Item-then-Sum) was a huge winner. It reduced errors by 30% to 80% for almost every model. Even the "Reasoning" models, which tried to do the math in their own heads while answering, couldn't beat the simple method of scoring each item separately and adding them up later.

Lesson: Don't ask the AI to do the whole job in one giant leap. Break it down into small, manageable steps.

3. The "Thinking Out Loud" Debate

The paper looked at whether "Reasoning" models (those that think out loud) were better than "Standard" models.

  • The Finding: It depends entirely on how you ask the question (the "Prompt").
    • Scenario 1: The Vague Test. If you just say, "Here is a conversation, tell me how depressed this person is," the Reasoning models win. They need to "think out loud" to figure out what to do.
    • Scenario 2: The Detailed Test. If you give the AI a clear rulebook, definitions of symptoms, and examples of how to score, the Standard models perform just as well (or sometimes even better) than the Reasoning ones.

The Analogy: Imagine a student taking a test.

  • If the teacher says, "Solve this mystery," the student who writes out their detective notes (Reasoning) does better.
  • If the teacher says, "Here is the formula, here is an example, now solve this," the student who just knows the formula (Standard) does just as well, but faster and without wasting energy.

Key Takeaway: "Reasoning" isn't magic. It's mostly a crutch for when the instructions are unclear. If you give the AI good instructions, it doesn't need to overthink.

4. Size Matters (But Not Everything)

  • Bigger is Better: Generally, the larger the model, the better it scored. A 70-billion-parameter model was much more reliable than a 1-billion-parameter one.
  • The "Tiny" Problem: The smallest models (under 8 billion parameters) struggled significantly. They often failed to produce a valid answer or gave wildly inaccurate scores. They simply didn't have enough "brain power" to handle the complexity of a real therapy session.
  • The "Thinking" Trap: Interestingly, for the Reasoning models, if they produced too much text (too many "thinking" steps), they actually got worse at the task. It's like a student who talks so much during the test that they run out of time to write the answer.

5. The "Blind Spot"

The AI models were great at scoring things like "Reduced Appetite" (Did they eat less?) because that's easy to hear in a conversation.
However, they struggled with "Apparent Sadness" (Does the person look sad?).

  • Why? The AI only had the text of the conversation. It couldn't see the patient's face, posture, or tears.
  • The Metaphor: It's like trying to judge if someone is crying by listening to a phone call where they are trying to hold back their voice. You might hear a sniffle, but you miss the tears. The paper suggests that for these specific symptoms, we need AI that can "see" and "hear" (multimodal), not just read text.

Summary: What Should We Do?

If you want to use AI to help doctors assess depression:

  1. Don't ask for a magic number. Ask the AI to score each symptom one by one, then add them up.
  2. Give clear instructions. Don't rely on the AI to "figure it out." Give it definitions and examples. You don't need the expensive "Reasoning" models if your instructions are good.
  3. Use big models. Small models aren't reliable enough for medical work yet.
  4. Remember the blind spot. AI can't see facial expressions yet, so it might miss some signs of sadness.

The Bottom Line: We don't need robots that "think" like humans to do this job; we need robots that follow clear, step-by-step instructions very well.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →