← Latest papers
💬 NLP

Before the Labels: How Dataset Construction Shapes Suicidality Detection in Clinical Text

This paper argues that clinical NLP models for suicidality detection often misinterpret dataset labels as ground truth, overlooking how construction choices in datasets like ScAN encode specific, limited operationalizations of suicidality that obscure clinical ambiguity and heterogeneity.

Original authors: Priyanshi Garg, Ishita Rao, Jieqiong Ding, Amandalynne Paullada

Published 2026-06-19
📖 5 min read🧠 Deep dive

Original authors: Priyanshi Garg, Ishita Rao, Jieqiong Ding, Amandalynne Paullada

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot to understand human sadness and thoughts of self-harm. To do this, you give the robot a massive library of doctors' notes. The paper argues that before we trust the robot's answers, we need to look closely at how the library was built, because the way the books are organized changes what the robot actually learns.

Here is the story of the paper, broken down into simple concepts and analogies.

1. The Core Problem: The "Filtered" View

The paper starts with a big idea: Electronic Health Records (EHRs) are not a direct window into a patient's mind. They are more like a photograph taken through a specific lens.

  • The Lens: The "lens" is the doctor. Doctors write notes based on what patients say, what they observe, hospital rules, and their own comfort levels.
  • The Distortion: If a patient is hesitant to talk about suicide, or if a doctor asks a question in a way that makes the patient say "no," the note will say "no suicide risk." The robot sees "no risk," but the reality might be much more complicated.
  • The Paper's Claim: When researchers turn these notes into data labels (like "Suicidal" or "Not Suicidal"), they aren't capturing the raw truth of the patient's feelings. They are capturing the doctor's documented judgment of those feelings.

2. The Case Study: The "ScAN" Dataset

The authors looked at a specific dataset called ScAN, which is built from real hospital records (MIMIC-III). They treated this dataset like a crime scene investigation to see how the "evidence" was collected. They found three main ways the data was "shaped" or "filtered":

A. The "Single Voice" Filter (Documentation-Mediated)

  • The Metaphor: Imagine a news report where only the police officer's voice is heard, and the witness's direct quotes are never recorded.
  • The Reality: The dataset only includes notes written by doctors. It excludes patient journals, intake forms filled out by patients, or direct conversations.
  • The Result: The robot learns what doctors think patients are feeling, not necessarily what patients are actually feeling. If a doctor misses a cue, the robot misses it too.

B. The "Snapshot" Filter (Episodic)

  • The Metaphor: Imagine trying to understand a person's entire life story by looking at a single photo taken on their birthday.
  • The Reality: The dataset treats suicide risk as a "snapshot" of a single hospital stay. It doesn't connect the dots between visits that happened years ago or visits that will happen in the future.
  • The Result: Suicide is often a long, recurring struggle. By cutting the data into isolated "hospital episodes," the dataset flattens a long movie into a single still image, losing the context of the patient's history.

C. The "Yes/No" Filter (Intent-Resolved)

  • The Metaphor: Imagine a teacher grading a test where "I don't know" and "The answer is definitely wrong" are both marked with the same red "F."
  • The Reality: In the real world, doctors often write notes that are ambiguous (e.g., "Patient seems unsure," or "Intent is unclear"). In this dataset, these "unsure" cases were often lumped together with "definitely not suicidal" cases for the sake of training the computer model.
  • The Result: The robot learns to treat genuine confusion as a clear "no." This erases the nuance of uncertainty, which is actually a very important part of clinical judgment.

3. The Linguistic Detective Work

To prove their point, the authors acted like linguists, looking at the actual words used in the notes. They found that identical labels hid very different stories.

  • The "Present" Lie: They found that notes labeled as "Current Suicidal Ideation" (meaning the patient is thinking about it right now) often contained words like "previously," "history of," or "in the past."
    • Analogy: It's like a weather report saying "It is raining" when the note actually says, "It rained last week." The label says "Rain," but the text says "Past Rain."
  • The "Unsure" Collapse: They found that notes marked "Unsure" had a lot of words indicating doubt (like "possibly," "maybe," "unclear"), while notes marked "Negative" (No risk) were very direct. But because the dataset merged them, the robot couldn't tell the difference between "I'm not sure" and "Definitely not."

4. Why This Matters (The "So What?")

The authors aren't saying the doctors did a bad job or that the dataset is "wrong." They are saying the dataset is a specific tool built for a specific purpose, and we need to know its limitations.

  • The Risk: If we treat these labels as absolute "Ground Truth" (the ultimate reality), we might build AI systems that are confident but wrong.
    • Example: A system might flag a patient as "high risk" just because they mentioned a past suicide attempt, even if they are safe now, because the label didn't distinguish between "past" and "present."
  • The Solution: The paper suggests we need to be transparent. We should tell users (and doctors) exactly how the data was built:
    • "This data only looks at single hospital visits."
    • "This data merges 'unsure' with 'no risk'."
    • "This data only sees what the doctor wrote, not what the patient said."

Summary

Think of the dataset as a map. The paper argues that this map is drawn by doctors, using a specific set of rules that cut out long-term history and merge confusing areas into clear zones.

The authors are saying: "Don't just look at the map and assume it's the whole territory. Look at the legend to understand what the mapmaker decided to include, exclude, and simplify." Only then can we trust the robot that reads the map.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →