← Latest papers
💻 computer science

IyawoBench v2.0: Extended Diagnostic Evaluation of Large Language Model Clinical Triage in Nigerian Primary Care

This paper introduces IyawoBench v2.0, a rigorous diagnostic framework and benchmark using Nigerian primary care data to reveal that conventional safety metrics mask critical failure modes in large language models, necessitating scenario-specific model selection for clinical triage in low-resource settings.

Original authors: Anthonio Oladimeji Gabriel, Dimeji Olawuyi

Published 2026-08-03
📖 5 min read🧠 Deep dive

Original authors: Anthonio Oladimeji Gabriel, Dimeji Olawuyi

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are the captain of a massive ship, but instead of steering through an ocean, you are navigating a stormy sea of medical emergencies. In many parts of the world, there aren't enough trained doctors to check every single patient who walks through the door. So, scientists have started building "digital first responders"—super-smart computer programs called Large Language Models. Think of these models as incredibly well-read librarians who have read every medical textbook ever written. They can look at a patient's symptoms and guess what's wrong, or more importantly, decide how urgent the problem is.

But here is the tricky part: how do you know if a digital librarian is actually safe to let loose in a real hospital? In the past, we checked their safety with simple "pass or fail" tests. It was like asking, "Did the librarian tell the patient to go to the hospital if they were bleeding?" If the answer was "yes," we gave them a gold star and said, "Safe!" But this paper argues that a gold star isn't enough. Just because a model doesn't send a bleeding patient home doesn't mean it's perfect. It might be sending them to the wrong kind of hospital, or it might be sending everyone to the hospital, even people with just a cold, which would clog up the emergency rooms. We need a way to see how the model makes mistakes, not just if it makes them.

This is exactly what the researchers at Iyawo Health in Nigeria set out to do. They created a new, super-detailed test called IyawoBench v2.0. Instead of just asking "Is it safe?", they built a mathematical microscope to look at three specific ways these AI models can trip up. They tested three of the smartest AI models available today (Claude Sonnet 4.6, Llama 3.3 70B, and Llama 3.1 8B) using 200 fake but realistic patient stories based on real data from 19 Nigerian health centers.

The results were a bit of a shocker. The researchers found that every single model they tested had at least one major flaw, even the ones that looked perfect on the old, simple tests.

Here is what they discovered about the three "flaws" they invented to catch these errors:

  1. The "Over-Protective" Flaw (Conservative Escalation Bias): Imagine a bodyguard who is so scared of danger that they tackle everyone who walks through the door, even the mailman. One of the models, Claude Sonnet 4.6, acted like this. It was great at spotting real emergencies, but it also sent too many people with minor problems to the hospital. This is bad because it overwhelms the hospital, making it harder for the truly sick people to get help.
  2. The "Under-Protective" Flaw (Systematic Downgrade Bias): This is the most dangerous one. Imagine a guard who sees a fire but tells the people to "wait and see" instead of calling the fire department. The Llama 3.1 8B model did this. On the old tests, it looked 100% safe because it never sent a critical patient home. But the new test revealed a scary secret: it downgraded 77 out of 100 critical patients to a "come back tomorrow" status. In real life, that delay could be fatal for someone with a severe infection.
  3. The "Confused Middle" Flaw (Middle-Tier Instability): Imagine a traffic cop who can't decide if a car should go fast or slow, so they wave it in both directions randomly. The Llama 3.3 70B model was great at spotting the extremes (super sick or totally fine) but got totally confused about the "medium" cases. It couldn't decide if these patients needed to go to the hospital today or tomorrow, flipping back and forth.

The paper also showed that there is no single "best" AI model for every situation. It depends on what the hospital needs most.

  • If the hospital is desperate to catch every single emergency and doesn't mind a crowded waiting room, the "over-protective" model (or even a simple rule that says "send everyone to the hospital") might be the best choice.
  • If the hospital is already full and can't handle more people, the "under-protective" model (which keeps people at home unless they are very sick) might actually be the most efficient, despite its risks.
  • If you want a balance, the "confused middle" model might be the winner.

The big takeaway is that we can't just pick one AI and say, "This is the winner." The "best" model changes depending on the specific problems the hospital is facing. The authors suggest that instead of just looking for a high score, we need to use their new "cost" math to figure out which mistakes a specific hospital can afford to make and which ones it absolutely cannot. They proved that the old way of testing AI was hiding these dangerous patterns, and we need this new, more detailed map to navigate the future of medical AI safely.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →