← Latest papers
🧬 biology

Frequency-Stratified Benchmarking Reveals Strong Long-Tail Limitations of Large Language Models in Diagnosis of Monogenic Diseases

This paper introduces MendelianBench-5K, a frequency-stratified benchmark of 5,000 cases, to demonstrate that large language models exhibit significant performance biases in diagnosing monogenic diseases, achieving high accuracy for well-studied genes while struggling with rare, recently characterized ones.

Original authors: Jihao Cai, Jianle Yang, Maimaitiyibubaji Abudukadier, Aoxiang Luo, Lina Zhao, Guozhuang Li, Kexin Xu, Zhihong Wu, Terry Jianguo Zhang, Sen Zhao, Zefu Chen, Nan Wu

Published 2026-06-29
📖 4 min read☕ Coffee break read

Original authors: Jihao Cai, Jianle Yang, Maimaitiyibubaji Abudukadier, Aoxiang Luo, Lina Zhao, Guozhuang Li, Kexin Xu, Zhihong Wu, Terry Jianguo Zhang, Sen Zhao, Zefu Chen, Nan Wu

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine the world of rare genetic diseases as a massive library. In this library, a few famous books (common diseases) are stacked on the front desk, read by everyone, and discussed constantly. But the vast majority of books are tucked away in the dusty, hard-to-reach back corners. These are the "long-tail" rare diseases—newly discovered, rarely reported, and known by very few people.

This paper is like a report card for a new kind of super-smart librarian: the Large Language Model (LLM). These are AI systems trained on huge amounts of text to answer questions. The researchers wanted to see if these AI librarians could help doctors diagnose patients by looking at their symptoms and guessing the right genetic "book" (gene) responsible.

Here is what they found, using simple analogies:

1. The Test: A Balanced Library Tour

The researchers built a special test called MendelianBench-5K. Instead of just testing the AI on the famous books at the front desk, they created a fair test with 5,000 patient cases. They made sure to pick an equal number of cases from the "famous" section (common genes) and the "dusty corner" section (rare, new genes).

2. The Result: The AI is a "Famous Book" Fanatic

When the AI tried to diagnose these patients, it did okay on the famous books but struggled terribly with the rare ones.

  • The Famous Books: When a disease was well-known and talked about a lot in medical literature, the AI was reasonably good at guessing the right gene.
  • The Dusty Corners: When the disease was rare or recently discovered, the AI's performance dropped off a cliff.

The Analogy: Imagine a student taking a history test. If the test asks about World War II (a famous topic), the student gets an A. But if the test asks about a tiny, obscure village uprising from last week that only one newspaper mentioned, the student fails miserably. The AI isn't "thinking" like a doctor; it's just remembering what it has read the most.

3. The "Star" Bias: Guessing the Popular Names

The researchers noticed a funny pattern in the AI's mistakes. When the AI didn't know the answer, it didn't guess randomly. Instead, it kept guessing the same few "famous" genes over and over again, even when they were wrong.

  • The Metaphor: It's like a detective who, when they can't solve a specific crime, just accuses the most famous criminal in town (like Sherlock Holmes' Moriarty) every single time, regardless of the evidence. The AI kept pointing to the same well-known genes (like MECP2 or FMR1) because those names appeared most often in its training data.

4. The "New Book" Problem

The AI also struggled with cases published very recently.

  • The Analogy: If a new book was published yesterday, the AI librarian hasn't had time to read it yet. The study found that the AI was much better at diagnosing diseases described in old reports (from 2004 or earlier) compared to reports from the last few years. The AI's knowledge is stuck in the past.

5. Too Much Information?

The researchers also tested how much information the AI needed.

  • The Sweet Spot: Giving the AI a little bit of symptom information helped it guess better.
  • The Overload: However, giving the AI too many symptoms (a long, messy list) didn't help. In fact, it sometimes made the AI worse. It's like trying to solve a puzzle while someone is shouting 50 different clues at you at once; the AI gets confused and stops performing well.

The Bottom Line

The paper concludes that current AI cannot be trusted to diagnose rare genetic diseases on its own.

  • It is biased toward what is popular and well-documented.
  • It fails when the disease is rare, new, or under-studied.
  • It tends to "hallucinate" by guessing the most famous genes when it's unsure.

The authors say that to truly test if these AI tools are safe for doctors to use, we need to stop testing them only on common diseases. We must test them on the "long tail" of rare diseases, just like they did in this study. Until then, these AI tools are like a student who memorized the most popular textbooks but hasn't learned how to handle the unknown.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →