AnchorSIPS: A Synthetic Dataset and Evaluation Resource for Evidence-Supported Psychosis-Risk Symptom Measurement
This paper introduces AnchorSIPS, a synthetic dataset of 10,000 structured psychosis-risk interviews with transcript-grounded measurement targets designed to overcome data-access bottlenecks and evaluate AI models' ability to perform evidence-supported symptom assessment and diagnosis.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Detective's Dilemma: Why AI Needs a Clue, Not Just a Guess
Imagine you are a detective trying to solve a mystery, but instead of a crime scene, you are looking at a person's mind. In the world of mental health, there is a specific kind of mystery called "psychosis-risk." This is when someone starts having thoughts or seeing things that feel very real to them but don't match the shared reality of everyone else—like hearing whispers when no one is there or believing in secret messages from the TV. Doctors use special interviews to figure out if these experiences are just early warning signs or something more serious.
The problem is that these interviews are like top-secret files. Because they contain deeply personal and private stories, doctors and researchers can't just share them with computers to learn from. It's like trying to teach a detective to solve crimes without ever letting them see a real case file. To get around this, scientists have started making "synthetic" data—fake but realistic stories that look and feel like real interviews but don't belong to any real person. However, most of these fake stories are just summaries or simple chats. They don't teach the computer how to think like a detective; they just ask it to guess the final answer. This paper introduces a new tool called AnchorSIPS, which is designed to teach AI not just the answer, but exactly where in the story the answer came from.
The New Detective Training Manual: AnchorSIPS
Meet AnchorSIPS. Think of it as a massive, 10,000-page training manual for AI detectives, but instead of real crimes, it uses perfectly crafted, fake interviews about people who might be developing psychosis. The goal isn't just to see if the AI can guess the final diagnosis (like "Yes, they have a risk" or "No, they don't"). The real goal is to see if the AI can point its finger at the exact sentence in the conversation that proves its guess.
In the real world, when a doctor interviews someone, they don't just jump to a conclusion. They ask a series of 24 specific questions. If the person says "Yes" to a question, the doctor asks follow-up questions: "How often does this happen?" "Does it bother you?" "Does it stop you from going to school?" Finally, the doctor makes a chain of decisions based on those answers to reach a final conclusion. This paper created a dataset where every single one of those steps is mapped out. It's like having a script where the "hidden truth" of the patient's life is written down first, and then the conversation is built around that truth, ensuring the clues are always there, even if the patient is being vague or shy.
The researchers built this dataset using a clever "Plan-then-Realize" method. Imagine a playwright who first writes a strict outline of the plot (the "Plan"), deciding exactly what the character will admit and what they will hide. Then, a talented actor (an AI) is hired to improvise the dialogue based on that outline. The actor can choose how to say the lines—maybe they stutter, maybe they sound scared, maybe they try to hide the truth—but they cannot change the plot. This ensures that the "clues" (the medical labels) are always correct and anchored to the specific words the patient said, preventing the AI from making up facts or getting confused.
The Big Surprise: AI is Good at Guessing, Bad at Proving
The researchers tested seven different super-smart AI models on this new dataset to see how well they could act like a doctor. They asked the AIs to do two things: guess the final diagnosis and, more importantly, cite the exact parts of the conversation that supported their guess.
Here is the twist: The AIs were surprisingly good at the easy stuff. They could guess the final diagnosis with high accuracy, often getting it right. It was like they could look at a messy room and guess, "Someone was here," without needing to see the footprints. But when the researchers asked them to point to the footprints (the specific sentences in the transcript), the AIs stumbled.
The results showed a big gap. While the models could make the "coarse" decisions (like "Yes, this is a risk"), they failed miserably at the "grounded" work. They struggled to extract the specific details, like how often a symptom happened or how much it caused distress. Even worse, when they tried to cite the evidence, they often got it wrong. One model might get the final answer right but cite the wrong sentence as proof, or it might make up a reason that wasn't actually in the text.
The paper suggests that this is a major problem. If an AI can guess the right answer but can't show its work, we can't trust it in a real hospital. It's like a student who gets the right answer on a math test but writes down the wrong formula; they might get lucky once, but they don't actually understand the math. The study found that current AI models are "overconfident" in their final guesses but "under-qualified" in their ability to find and use the evidence to back them up.
Why This Matters
This isn't just about making better chatbots; it's about safety. In mental health, especially with something as serious as psychosis, you can't rely on a machine that just "feels" like it's right. You need a machine that can say, "I think this person is at risk, and here are the three sentences they said that prove it."
The authors of this paper are clear: AnchorSIPS is a research tool, not a medical device. It is a synthetic dataset, meaning the patients and stories are made up by computers, not real people. It is designed to stress-test AI systems, to see where they break, and to teach them how to be more honest about what they know and what they don't. The study suggests that while AI is getting better at talking, it still has a long way to go before it can be trusted to listen, analyze, and prove its conclusions in the high-stakes world of mental health. Until AI can learn to anchor its answers to the evidence, just like a good detective, it remains a guesser, not a healer.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.