← Latest papers
💬 NLP

TreeProbe : A Tibetan Medicine Benchmark for Cultural Bias in LLMs

This paper introduces TreeProbe, the first benchmark designed to evaluate cultural bias in large language models by testing their adherence to the structured theoretical framework of Tibetan medicine, revealing that current models often exhibit systematic epistemic drift toward dominant biomedical or Traditional Chinese Medicine paradigms.

Original authors: Jin Zhang, Linyu Li, Weili Jiang, Yuqing Cai, Yutong Liu, Guanquecairang, Yongbin Yu, Jingye Cai, Nyima Tashi, Gadeng Luosang

Published 2026-08-04
📖 3 min read☕ Coffee break read

Original authors: Jin Zhang, Linyu Li, Weili Jiang, Yuqing Cai, Yutong Liu, Guanquecairang, Yongbin Yu, Jingye Cai, Nyima Tashi, Gadeng Luosang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you're walking into a giant, high-tech library where the books are written by super-smart robots. These robots, called Large Language Models (LLMs), have read almost everything on the internet. They are amazing at answering questions, writing stories, and even pretending to be doctors. But here's the catch: most of the books they read were written in English and based on Western science. So, when you ask them about a medical system from a different culture—like Tibetan medicine, which has its own unique way of thinking about the body and illness—the robots get confused. Instead of using the Tibetan rules, they accidentally switch to the Western rules or even Chinese medicine rules, like a chef who only knows how to make pizza suddenly trying to make sushi but using pizza dough. This paper is about building a special test to see exactly how and why these robots make that mistake, and how much they "drift" away from the truth.

The researchers behind this study, led by Jin Zhang and a team from universities in China and Tibet, created a new tool called TreeProbe. Think of Tibetan medicine as a giant, ancient tree with deep roots. The paper uses the "Tree of Medicine," a classic framework from Tibetan texts, to build a test with 4,719 questions covering 467 different diseases. They didn't just ask the robots simple trivia; they asked them to explain why a disease happens, how to diagnose it using the Tibetan "three humors" (a concept similar to balancing energies), and how to treat it with specific diets or herbs.

The team found that even the smartest robots today are still struggling. When tested, the best models only got about 60% of the multiple-choice questions right. But the real story isn't just that they got questions wrong; it's how they got them wrong. The robots didn't just guess randomly; they consistently drifted toward other medical systems. Some models, like the one from Anthropic (Claude), tended to switch to Western biomedical thinking, while others, like some Chinese models, drifted toward Traditional Chinese Medicine (TCM). It's as if the robots have a "default setting" that overrides the specific Tibetan knowledge they are supposed to use.

The researchers also discovered that the robots are good at sounding the part. They can write fluent Tibetan sentences that look perfect grammatically, but the medical advice inside is often wrong or mixed up with other systems. It's like a robot that speaks perfect French but gives you a recipe for a German dish when you asked for a French one. The study suggests that this happens because the robots were trained on data where Western and Chinese medical ideas are much more common than Tibetan ones. When the robot gets stuck, it falls back on what it knows best, accidentally erasing the unique logic of Tibetan medicine.

In short, TreeProbe acts like a diagnostic tool for the robots themselves. It shows us that to make AI truly fair and useful for everyone, we can't just feed it more data; we have to teach it to respect and understand different ways of knowing the world. The paper doesn't claim to have fixed the problem yet, but it provides the first clear map of where the robots are getting lost, helping developers build better, more culturally aware AI in the future.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →