← Latest papers
💰 quantitative finance

How much does context affect the accuracy of AI health advice?

This study demonstrates that the accuracy of large language models in verifying health claims is highly dependent on linguistic, topical, and source contexts, revealing significant performance gaps in non-English languages and complex real-world scenarios that challenge the generalizability of English-language benchmarks.

Original authors: Prashant Garg, Thiemo Fetzer

Published 2026-02-25
📖 4 min read☕ Coffee break read

Original authors: Prashant Garg, Thiemo Fetzer

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a team of seven super-smart robots (AI chatbots) that everyone is starting to trust for medical advice. You ask them, "Is this health tip true?" and they answer "Yes" or "No."

This paper is like a report card for those robots, but with a twist: the teachers are checking if the robots get the answers right not just in English, but in 21 different languages, and not just for simple facts, but for messy, real-world news stories.

Here is the breakdown of what they found, using some everyday analogies:

1. The "Textbook" Test vs. The "Street Smarts" Test

The researchers gave the robots two very different types of tests:

  • Test A: The Official Rulebook (The "Textbook")
    They used a list of health claims that are legally approved by the UK and EU governments (like "Vitamin C supports the immune system").

    • The Result: The robots were like star students when speaking English or languages close to English (like Spanish or French). They got almost 97% of these right.
    • The Catch: As the language got further away from English (like Persian, Hindi, or Korean), the robots started to stumble. It's like a student who memorized the textbook in English but gets confused when the teacher switches to a language with a totally different grammar structure.
  • Test B: The Real-World News Feed (The "Street Smarts")
    They then tested the robots on 9,000 real health claims found in news articles, government reports, and social media (covering topics like COVID-19, abortion, and politics).

    • The Result: This is where the robots really struggled. Their accuracy dropped significantly. They were like confused tourists trying to navigate a busy, noisy city.
    • The Pattern:
      • They were best at spotting truths in simple, direct news snippets (like "The CDC says...") or about COVID-19 (because everyone talked about that so much, the robots "studied" it heavily).
      • They were worst at understanding complex scientific abstracts (like those from the journal Nature) or general health advice. These are like trying to read a dense, legal contract while wearing foggy glasses.

2. Why Do They Fail? (The "Why" Behind the Mistakes)

The paper suggests three main reasons why the robots aren't perfect:

  • The Library Bias (Training Data): The robots were trained mostly on the "open web." This is like a library that has millions of copies of The New York Times and Fox News (short, punchy sentences) but very few copies of Nature (complex, nuanced science). The robots are great at reading the newspapers but bad at reading the science journals.
  • The "Hedge" Problem: Real science often uses words like "might," "could," or "suggests." The robots prefer clear, bold statements. When a sentence says, "This drug might help," the robot gets confused and might guess wrong.
  • The "Specialist" Gap: The robots were tuned heavily to handle the COVID-19 crisis (like a student who only studied for the math final). But when asked about politics or general health, they didn't have that same "specialized training," so they made more mistakes.

3. The Big Warning

The main takeaway is a warning label for anyone thinking, "Hey, this AI works great in English, so it must be safe for everyone."

It's not.
Relying on an English-speaking AI to give health advice in a different language or on a complex topic is like hiring a tour guide who only knows the city center to show you around the dangerous, winding backstreets. They might get you lost, or worse, give you dangerous directions.

4. What Should We Do?

The authors suggest we stop assuming the robots are ready for the whole world. Instead:

  • Test locally: Before letting an AI give health advice in a specific country or language, we need to test it specifically for that culture and language.
  • Human in the loop: For serious health issues, a human expert should double-check the robot's work, especially when the topic is complex or the language is rare.
  • Build better libraries: We need to train these robots on more diverse, high-quality scientific data in many languages, not just English.

In short: AI health advice is a powerful tool, but right now, it's a "one-size-fits-all" suit that fits English speakers well but is too tight or too loose for everyone else. We need to tailor the suit before we let the robot walk into the hospital.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →