← Latest papers
💻 computer science

Trust, Safety, and Accuracy: Assessing LLMs for Routine Maternity Advice

This study evaluates the performance of large language models like Perplexity AI and ChatGPT-4o against maternal health experts in rural India, finding that while Perplexity offers superior semantic accuracy, ChatGPT-4o provides clearer, more understandable medical advice, highlighting the potential of AI as a scalable tool for maternal health education when balancing accuracy and clarity.

Original authors: V Sai Divya, A Bhanusree, Rimjhim, K Venkata Krishna Rao

Published 2026-03-19
📖 4 min read☕ Coffee break read

Original authors: V Sai Divya, A Bhanusree, Rimjhim, K Venkata Krishna Rao

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a pregnant woman living in a small village in India. You have a question about your health, but you feel shy asking a doctor because of cultural taboos, or maybe the nearest clinic is too far away. So, you turn to your phone and ask an AI chatbot. But here's the big worry: Is the AI giving you safe, accurate advice, or is it just making things up in a language you can't understand?

This paper is like a report card for three famous AI chatbots (ChatGPT, Perplexity, and Gemini) to see how well they act as "digital midwives" for Indian mothers.

Here is the breakdown of their performance using simple analogies:

1. The Test: The "Myth-Busting" Exam

The researchers didn't just ask the AIs random questions. They gave them 17 specific questions that real Indian women often ask, including some common myths and misconceptions (like "Can I eat this specific fruit?" or "Is this old wives' tale true?").

They treated the AIs like students taking a test. To grade them, they had a panel of real, expert doctors answer the same questions first. The doctors' answers were the "Gold Standard" (the perfect answer key).

2. The Grading Criteria: Three Ways to Judge

The researchers used three different "rulers" to measure how good the AI answers were:

  • Ruler A: The "Grandma Test" (Readability)

    • The Goal: Can a woman with a basic school education understand this?
    • The Analogy: Imagine the AI is a teacher. If the teacher uses big, complicated words like a university professor, the student gets lost. If the teacher speaks simply, like a friendly neighbor, the student learns.
    • The Result: ChatGPT was the best teacher. It spoke in clear, simple sentences (like a 7th or 8th grader reading level). Perplexity was a bit too academic, using long, complex sentences that might confuse someone in a hurry.
  • Ruler B: The "Meaning Match" (Semantic Similarity)

    • The Goal: Did the AI actually understand the concept, or did it just guess?
    • The Analogy: Imagine you and a friend are describing a movie. If you both say, "It was a sad story about a dog," you have high "meaning match." If you say, "It was a comedy about a cat," the match is low.
    • The Result: Perplexity was the best at matching the core meaning of what the doctors said. It got the "gist" of the medical advice very well.
  • Ruler C: The "Keyword Check" (Noun Overlap)

    • The Goal: Did the AI mention the important medical terms (like "blood pressure," "vitamins," "delivery") that the doctors mentioned?
    • The Analogy: Think of a grocery list. If the doctor says, "Buy milk, eggs, and bread," and the AI says, "Buy milk, eggs, and bread," they have a perfect match. If the AI says, "Buy milk, eggs, and... a bicycle," the match is low.
    • The Result: ChatGPT was the best at including the right medical keywords in the right context.

3. The Final Scoreboard

  • 🏆 ChatGPT (The Clear Communicator): It won the overall race for clarity. Its answers were the easiest to read and it used the right medical terms. It felt like a knowledgeable friend explaining things simply.
  • 🥈 Perplexity (The Accurate Researcher): It was the closest to the doctors in terms of deep meaning, but its answers were sometimes too complicated to read easily.
  • 🥉 Gemini: It did a decent job but didn't stand out as much as the other two in this specific test.

4. Why Does This Matter?

The paper concludes that AI is a powerful tool, but it needs to be the right kind of tool.

  • The Good News: AI can help break the silence. Women who are afraid to ask doctors due to shame or distance can get quick, private, and generally accurate answers.
  • The Warning: If the AI speaks too "smartly" (like Perplexity did sometimes), a mother might not understand the advice, which could be dangerous.

In a nutshell: This study tells us that ChatGPT is currently the best "digital health assistant" for everyday pregnancy questions in India because it speaks a language that regular people can understand, while still getting the medical facts mostly right. However, we still need to keep checking these AIs to make sure they don't spread myths and that they keep getting better at understanding local cultures.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →