IyàwóBench: A Benchmark for Evaluating Large Language Model Clinical Triage Accuracy on Undifferentiated Febrile Illness in Nigerian Primary Health Settings
The paper introduces IyàwóBench, the first benchmark for evaluating large language models on clinical triage for undifferentiated febrile illness in Nigerian primary care, revealing that while modern models demonstrate high safety by avoiding dangerous downgrades, their diagnostic accuracy varies significantly and is substantially improved by systems specifically engineered with WHO guidelines.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a busy clinic in Nigeria where a community health worker is the first line of defense for thousands of people. Every day, they see patients with fevers. The problem? A fever is like a generic "check engine" light on a car. It could be something simple like a cold, or it could be a life-threatening emergency like meningitis or severe sepsis. The health worker has to decide instantly: "Do I treat this person right here?" or "Do I send them to a hospital immediately?"
Getting this wrong is dangerous. If you send a critical patient home, they might die. If you send a mild patient to the hospital, you clog up the system and waste resources.
This paper introduces a new "test drive" called IyàwóBench to see if Artificial Intelligence (AI) can help these health workers make the right call.
The "Test Drive" (The Benchmark)
The researchers created a digital simulation, like a video game for doctors. They built 200 fake patient stories (vignettes) based on real data from 1,200 actual patients in Nigeria. These stories cover eight different types of febrile illnesses, ranging from simple malaria to deadly bacterial infections.
They asked six different AI models (the "drivers") to read these stories and pick one of three actions:
- REFER NOW: Go to the hospital immediately (Emergency).
- REFER TODAY: Go to the hospital later today (Urgent).
- TREAT HERE: Stay here and get medicine (Safe).
The Results: Safety vs. Accuracy
The paper tested the AI on two main things: Safety and Accuracy.
1. The Safety Score (The "Don't Kill Anyone" Test)
This was the most important result. The researchers checked: Did any AI ever tell a critical, life-threatening patient to "Stay here and treat yourself"?
- The Result: Zero. Every single AI model got a 100% safety score.
- The Analogy: Imagine a group of new drivers taking a test. Even if they are bad at parallel parking, none of them drove into a crowd of pedestrians. They all recognized the "stop" signs and the red lights of a medical emergency. They were all too scared to make the most dangerous mistake.
2. The Accuracy Score (The "Right Call" Test)
This measured how often the AI picked the exact right level of care (e.g., saying "Refer Now" when the patient actually needed it).
- The Result: The AI models were all over the map.
- The Star Performer: One model (Claude Sonnet) got about 67.5% right. It was the best, but still missed about 1 out of every 3 cases.
- The Middle Pack: Other models like Llama 4 Scout got around 59.5%.
- The Strugglers: Some models, like Llama 3.1 8B, only got 39% right.
- The "Silent" Failures: Two models (Qwen and GPT OSS) got almost 0% accuracy.
- The Analogy: Think of this like a multiple-choice test. The best AI got a "C" grade. The worst "parseable" AI got an "F." The two that got "0%" didn't fail because they didn't know the answers; they failed because they refused to write the answer in the specific box the teacher asked for. They wrote long essays instead of checking the box.
Key Takeaways from the Paper
- AI is Cautious: The AI models are very good at spotting when a situation is scary and dangerous. They would rather send a patient to the hospital unnecessarily than risk sending a critical patient home.
- Size Doesn't Always Mean Smarts: The researchers found that having a "bigger" brain (more computer parameters) didn't guarantee a better score. A smaller, specialized model with specific instructions about Nigerian health rules performed better than a massive, general-purpose model.
- The "Format" Problem: Two powerful models failed the test not because they were bad at medicine, but because they couldn't follow the strict rules of the test format. They gave great medical advice, but they didn't say it in the exact code the computer needed to read.
- The Gap: There is a big difference between the best AI (67.5%) and the worst (39%). This means we can't just pick any AI and hope it works; we need to choose carefully and maybe "train" it specifically for Nigerian clinics.
What the Paper Does Not Say
It is important to stick to what the paper actually found:
- This paper does not say AI is ready to replace human doctors in Nigeria.
- It does not claim these models are perfect (67.5% accuracy means they are wrong nearly a third of the time).
- It does not test if the AI can prescribe the right medicine or dosage, only if it can decide where the patient should go.
- It does not test if the AI works on a real phone with a slow internet connection, only in a cloud simulation.
The Bottom Line
The paper introduces a new ruler (IyàwóBench) to measure AI in Nigerian clinics. It found that while AI is very careful and won't make the most deadly mistakes, it is still inconsistent at figuring out the exact right level of care. To be useful, future AI tools will need to be trained specifically on local data and forced to follow strict formatting rules.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.