How Good LLMs Are at Answering Bangla Medical Visual Questions? Dataset and Benchmarking
This paper introduces BanglaMedVQA, the first clinically validated dataset and benchmark for Bangla Medical Visual Question Answering, revealing that current state-of-the-art foundation models, including Gemini and GPT-4.1 mini, perform significantly poorly on specialized diagnostic tasks due to challenges inherent in low-resource languages and complex medical reasoning.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a group of very smart, well-read robots (called AI models) that are great at looking at pictures and answering questions about them. You've taught them to speak English fluently, and they can even describe a picture of a broken bone or a tumor quite well.
But what happens if you ask them the same questions in Bangla, a language spoken by nearly 300 million people? That is exactly what this paper investigates.
Here is the story of their findings, broken down into simple parts:
1. The Missing Puzzle Piece
For a long time, scientists have been testing these AI robots on medical pictures in English. They have big, high-quality test sets for English. But for Bangla? There was almost nothing. There was one old dataset, but it was like a puzzle with missing pieces and blurry pictures—nobody knew if the answers were actually correct or if a doctor had even checked them.
The authors of this paper decided to build a brand new, high-quality test set called BanglaMedVQA.
- The Ingredients: They took thousands of medical images (like X-rays and CT scans) from two famous, trusted English databases.
- The Recipe: They used advanced AI to translate the medical questions and answers into Bangla.
- The Quality Control: This is the most important part. They didn't just trust the computer. They hired two real, certified doctors to check every single question and answer. If the doctors said, "No, that's wrong," they fixed it. The result was a dataset with a 97% accuracy rate, verified by humans.
2. The Big Test: How Smart Are the Robots?
The authors put the top AI robots (both the expensive, closed-source ones like Google's Gemini and OpenAI's GPT, and the free, open-source ones) through this new Bangla test.
The Results were surprising and a bit scary:
- The "General" Questions: When asked simple things like, "What kind of picture is this?" (Is it an X-ray or an MRI?) or "What body part is this?" (Is it a lung or a brain?), the robots did okay. They got about 40% of the answers right.
- The "Specialist" Questions: When the questions got harder—like, "What specific disease is this?" or "Exactly where in the lung is the tumor?"—the robots crashed.
- Even the best robots (Gemini and GPT) performed worse than random guessing on these hard questions. It was as if they were just throwing darts at a board and hoping to hit the bullseye.
- The free, open-source robots did even worse, often getting less than 10% right.
3. The Language Barrier
The researchers also tested the robots in English using the same questions.
- The Gap: The robots were significantly better at English than Bangla. It's like a student who is an A+ in English class but fails the same test when it's translated into a language they barely understand.
- The Takeaway: The robots haven't been trained enough on medical knowledge in Bangla. They are "low-resource" learners in this language.
4. Can We Help Them Think Better?
The researchers tried a trick called "Chain-of-Thought." Instead of asking the robot for the answer immediately, they told it: "Let's think step-by-step before you answer."
- Did it work? Yes, but only a little bit. It helped the robots reason a bit more, especially the free ones, but it didn't fix the core problem. They still struggled to pinpoint exactly where a disease was or what specific condition it was.
- The Bottleneck: The paper suggests the main problem isn't just the language; it's that the robots don't truly "see" the medical details in the picture well enough to explain them in Bangla.
5. The Bottom Line
This paper is a wake-up call.
- Current Status: We have built a high-quality, doctor-verified test set for Bangla medical AI.
- The Reality: Even the most advanced AI models today are not ready to be trusted with complex medical diagnoses in Bangla. They are good at describing the picture generally, but they fail at the critical job of diagnosing the patient.
- The Future: We need to build better models and train them specifically on medical data in Bangla before we can ever use them in a real hospital.
In short: The robots are like medical students who have memorized a textbook in English but are currently failing their final exam when asked to diagnose a patient in Bangla. We have created the exam paper (the dataset) to prove this, and the scores show we have a long way to go.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.