← Latest papers
💬 NLP

SynSym: A Synthetic Data Generation Framework for Psychiatric Symptom Identification

The paper introduces SynSym, a framework that leverages large language models to generate diverse, clinically grounded synthetic data for psychiatric symptom identification, demonstrating that models trained on this synthetic data achieve performance comparable to those trained on real-world annotations.

Original authors: Migyeong Kang, Jihyun Kim, Hyolim Jeon, Sunwoo Hwang, Jihyun An, Yonghoon Kim, Haewoon Kwak, Jisun An, Jinyoung Han

Published 2026-03-24
📖 5 min read🧠 Deep dive

Original authors: Migyeong Kang, Jihyun Kim, Hyolim Jeon, Sunwoo Hwang, Jihyun An, Yonghoon Kim, Haewoon Kwak, Jisun An, Jinyoung Han

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot to understand human sadness. You want the robot to spot specific signs of depression—like "can't sleep," "feel worthless," or "want to give up"—in the millions of posts people write on social media every day.

The problem? Teaching the robot is incredibly hard.

Real human posts about mental health are messy. Sometimes people are very direct ("I am depressed"). Sometimes they are poetic ("My heart is a sinking ship"). Sometimes they are sarcastic. To train a robot to understand all these nuances, you need a massive library of examples, each carefully labeled by a team of expensive, highly trained psychiatrists. But there aren't enough psychiatrists in the world to label millions of posts, and the ones who do often disagree with each other. It's like trying to build a library with only a few books, written in different languages, by authors who can't agree on the plot.

Enter SynSym.

Think of SynSym not as a robot, but as a super-powered, tireless ghostwriter who has read every medical textbook and every social media post ever written. Its job is to write a massive library of fake (synthetic) stories that look and feel exactly like real human posts, but are generated by a computer.

Here is how SynSym works, broken down into four simple steps:

1. The "Brainstorming" Phase (Concept Expansion)

Imagine you ask a human writer to write about "sadness." They might just write "I feel sad." That's boring and not very helpful for training a robot.
SynSym asks the AI to break "sadness" down into tiny, specific pieces first. Instead of just "sadness," it thinks of: feeling empty, crying for no reason, losing interest in hobbies, feeling heavy.
Analogy: It's like a chef who doesn't just say "make soup." They first list every possible ingredient (carrots, onions, broth, spices) so they can cook thousands of unique variations of soup, ensuring the robot learns to recognize every way soup can taste.

2. The "Two Voices" Phase (Dual Styles)

Real people talk differently depending on where they are. A doctor in a hospital writes differently than a teenager on Twitter.
SynSym forces the AI to write in two distinct voices for every symptom:

  • The Doctor Voice: Formal, clinical, using medical terms (e.g., "I am experiencing insomnia").
  • The Friend Voice: Casual, emotional, using slang or metaphors (e.g., "I haven't slept in three days and my brain won't shut up").
    Analogy: It's like training a translator who needs to understand both a formal legal contract and a funny text message. If you only teach the robot the legal language, it will fail when a teenager sends a text.

3. The "Real Life" Phase (Multi-Symptom Mixing)

In the real world, people rarely have just one problem. If someone is depressed, they might also be anxious and tired.
SynSym doesn't just write about one symptom at a time. It uses medical knowledge to mix symptoms together realistically. It knows that "suicidal thoughts" often go with "feeling worthless," so it writes posts that combine them naturally.
Analogy: Imagine training a detective. If you only show them a crime scene with a single broken window, they might miss the bigger picture. SynSym creates complex crime scenes with multiple clues, teaching the robot to spot how different problems overlap in real life.

4. The "Quality Control" Phase (The Editor)

Before the AI's work is published, it gets a second look. The system checks: "Does this sentence actually mean what we think it means?" If the AI wrote something weird or confusing, it gets thrown in the trash.
Analogy: It's like a strict editor at a newspaper who rejects any story that doesn't make sense, ensuring the final library is full of high-quality, accurate stories.

The Results: Does it work?

The researchers tested this "ghostwriter" by training a robot using only the fake stories SynSym wrote.

  • The Surprise: The robot trained on fake data performed just as well as robots trained on real, human-labeled data.
  • The Superpower: When they gave the robot a little bit of real human data to fine-tune its skills, it became even better than any robot trained on real data alone.

Why is this a big deal?

Think of it like this:

  • Old Way: You try to teach a student by giving them 100 real textbooks, but the pages are torn, the handwriting is messy, and the teachers disagree on the answers.
  • SynSym Way: You give the student 10,000 perfect, clear, and diverse practice books written by a genius AI. The student learns faster, understands more, and is ready for the real world.

In short: SynSym solves the "not enough data" problem in mental health research. It creates a safe, endless supply of high-quality training examples, allowing AI to become a better tool for spotting mental health struggles early, potentially saving lives.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →