← Latest papers
💬 NLP

PersianMedQA: Evaluating Large Language Models on a Persian-English Bilingual Medical Question Answering Benchmark

This paper introduces PersianMedQA, a large-scale bilingual dataset of 20,785 expert-validated Persian medical exam questions, to evaluate 41 state-of-the-art LLMs and reveal that while closed-weight general models outperform specialized Persian models, cultural and clinical nuances in the original language remain critical for accurate medical reasoning.

Original authors: Mohammad Javad Ranjbar Kalahroodi, Amirhossein Sheikholselami, Sepehr Karimi, Sepideh Ranjbar Kalahroodi, Heshaam Faili, Azadeh Shakery

Published 2026-05-27
📖 4 min read☕ Coffee break read

Original authors: Mohammad Javad Ranjbar Kalahroodi, Amirhossein Sheikholselami, Sepehr Karimi, Sepideh Ranjbar Kalahroodi, Heshaam Faili, Azadeh Shakery

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a group of students how to be doctors, but you have to test them in two different languages: Persian (their native tongue) and English (the language of most medical textbooks).

This paper, called PersianMedQA, is like a massive, 14-year-old exam archive from Iran. It contains over 20,000 multiple-choice questions covering 23 different medical specialties, from heart surgery to pediatrics. The authors used this "exam bank" to see how well different AI "students" (Large Language Models) can answer medical questions in both languages.

Here is the breakdown of what they found, using some simple analogies:

1. The "Star Students" vs. The "Struggling Locals"

The researchers tested 41 different AI models.

  • The "Superstars" (Closed Models): The top performers were like the elite, private-school students who have access to the best resources. Specifically, GPT-4.1 (a closed, paid model) got about 83% of the Persian questions right and 80% of the English ones. It was the clear winner.
  • The "Local Heroes" (Persian Models): The authors were surprised to find that models specifically built to speak Persian (like Dorna) performed terribly. They scored around 35%, which is barely better than guessing. It's as if a student who grew up speaking the language forgot how to read the medical textbooks in that same language. They struggled to follow instructions and couldn't reason through the problems.
  • The "Specialized Students" (Medical Models): Models trained specifically on medical data also didn't do as well as the general "Superstars." Being a "medical AI" didn't automatically make them better at Persian medical questions.

2. The "Translation Trap"

You might think, "If an AI is smart in English, just translate the Persian questions to English, and it will work."

  • The Reality: The paper found that while AI generally does better in English, 3% to 10% of the questions can ONLY be answered correctly in Persian.
  • The Analogy: Imagine a recipe that says, "Add a pinch of Saffron." If you translate it to English, it just says "Saffron." But in the original Persian context, the recipe might imply a specific type of Saffron grown in a specific Iranian village that is cheaper and more common there. If you translate the question, you lose that local context.
  • The Result: In some cases, the AI got the answer right in Persian because it understood the local rules (like which vaccines are given at what age in Iran), but got it wrong in English because the translation made it sound like a Western rule.

3. Bigger Isn't Always Better

The researchers checked if making the AI "brain" bigger (more parameters) helped.

  • The Finding: For the general "Superstar" models, bigger was better. But for the specialized medical or Persian models, just making them bigger didn't help. It's like giving a student a bigger backpack; if they don't have the right books inside, a bigger bag doesn't make them smarter. They need the right data, not just more of it.

4. The "Cheating" Test

To see if the AIs were actually thinking or just guessing based on patterns, the researchers tried a trick: they showed the AI only the answer choices without the question.

  • The Result: In the "Medical Ethics" section, the AI got surprisingly high scores just by looking at the answers. It learned that answers containing words like "patient autonomy" or "consent" were usually the right ones. It was like a student who didn't study the lesson but knew that if an answer sounds very polite and formal, it's probably the correct one. This suggests the tests might be overestimating how well the AI actually understands ethics.

5. The "Teamwork" Experiment

Since no single model was perfect, the researchers tried having the top 3 or 5 models vote on the answer together (like a jury).

  • The Result: This "team" approach helped the open-source models improve their scores slightly, but it still couldn't catch up to the single "Superstar" model (GPT-4.1).

The Bottom Line

The paper concludes that you cannot simply translate medical exams from English to other languages and expect AI to work well.

  • Medical knowledge is deeply tied to local culture, laws, and habits (like vaccination schedules).
  • Current AI models, even those trained on Persian, are not yet reliable enough to be trusted with high-stakes medical questions in that language.
  • We need to build better, culturally-aware AI specifically for low-resource languages, rather than just translating English models.

Important Note from the Authors: The paper explicitly states that these models are not ready to be used as real doctors. They are currently just being tested on exam questions, and they should not be deployed in real hospitals without strict human supervision.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →