← Latest papers
💬 NLP

Prompting language influences diagnostic reasoning and accuracy of large language models

This study demonstrates that prompting language significantly impacts the diagnostic reasoning and accuracy of most large language models in clinical settings, with English generally outperforming French across four of five evaluated models, highlighting a critical barrier to equitable global deployment.

Original authors: Adrien Bazoge, Josselin Corvellec, Sofiane Djillali Sid-Ahmed, Pierre-Antoine Gourraud

Published 2026-05-20
📖 4 min read☕ Coffee break read

Original authors: Adrien Bazoge, Josselin Corvellec, Sofiane Djillali Sid-Ahmed, Pierre-Antoine Gourraud

Original paper dedicated to the public domain under CC0 1.0 (http://creativecommons.org/publicdomain/zero/1.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a team of five brilliant medical students. They are incredibly smart, have read millions of medical books, and can solve complex puzzles. However, there's a catch: they all went to school in different countries, and their "native language" for thinking is English.

This paper is like a report card testing how well these students perform when asked to solve medical cases in English versus French. The researchers wanted to see if the language used to ask the question changes the quality of the answer.

Here is the breakdown of what they found, using simple analogies:

1. The Setup: The "Medical Puzzle" Test

The researchers created 180 medical puzzles (called vignettes). These weren't just multiple-choice questions; they were short stories about patients with symptoms, similar to what a doctor sees in a real clinic.

  • The Test: They gave these puzzles to five different AI models (the "students"): o3, DeepSeek-R1, GPT-4-Turbo, Llama-3.1, and BioMistral.
  • The Twist: They asked the same puzzles twice to each student: once in English and once in French.
  • The Graders: Two real doctors acted as the teachers. They didn't just check if the final answer was right; they graded the entire thought process (the reasoning) on a scale of 0 to 18. They looked at things like:
    • Did the student notice all the clues?
    • Was the logic sound?
    • Did they consider other possibilities?
    • Was the final diagnosis correct?

2. The Results: The "English Advantage"

For four out of the five students, the results were clear: They performed significantly better when speaking English.

  • The Gap: Even though these models can speak French fluently, their "brain" worked better in English. It's like a musician who can play a song in French, but their fingers move faster and their rhythm is more perfect when playing in English.
  • The Score: On average, the English answers were better by a noticeable margin (between 0.37 and 0.91 points on the 18-point scale).
  • Where the Gap Happened: The difference wasn't just about the final answer. The students made more logical mistakes, missed more clues, and had weaker reasoning when forced to think in French.
    • Analogy: Imagine a detective solving a mystery. In English, they follow the trail of clues perfectly. In French, they might get distracted, miss a clue, or jump to a conclusion that doesn't quite fit the evidence, even if they eventually guess the right name.

3. The Exception: The "Super-Student"

One model, o3, was the only one that showed no difference between English and French.

  • Why? The paper suggests this model has a special "superpower" called advanced reasoning capabilities. It seems that as AI gets better at thinking through problems step-by-step (like doing complex math in its head), the language barrier starts to disappear. It's like a student who has mastered the logic of the puzzle so well that it doesn't matter which language the instructions are written in.

4. Why This Matters (According to the Paper)

The paper argues that language isn't just a wrapper for the AI; it's part of how the AI thinks.

  • The "Training Diet" Problem: Most of these AIs were trained on huge piles of English data (like a student who only read English textbooks). Even if they can translate French, their "muscle memory" for medical reasoning is built on English patterns.
  • The Risk: If a doctor in a French-speaking hospital uses these tools, the AI might give a slightly less accurate diagnosis or a messier explanation than it would for an English-speaking doctor.
  • The Conclusion: The paper concludes that for AI to be truly fair and useful everywhere, we can't just rely on English. We need to ensure these models are just as sharp in French (and other languages) as they are in English.

Summary

Think of these AI models as multilingual chefs. They can cook a great meal (diagnosis) if you give them the recipe in English. If you give them the same recipe in French, four of the five chefs will still cook a good meal, but it might be slightly less seasoned or the presentation might be a bit off. Only one chef (o3) cooks the exact same perfect meal regardless of the language of the recipe.

The paper warns us that until we fix this "language bias," relying on these tools in non-English settings might mean getting a slightly lower-quality result.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →