EpiCurveBench: Evaluating VLMs on Epidemic Curve Digitization
This paper introduces EpiCurveBench, a benchmark of 1,000 real-world epidemic curve images, and EpiCurveSimilarity (ECS), a novel evaluation metric that accounts for temporal structure and local shifts, to demonstrate that current vision-language models struggle with epidemic curve digitization and that ECS provides a more discriminative and epidemiologically relevant assessment than existing key-value metrics.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a library of old, dusty public health reports. Inside these reports are colorful charts (called "epicurves") that show how many people got sick, were hospitalized, or died during various outbreaks over the last 175 years. The problem? These charts are just pictures. To use this data to predict future outbreaks, computers need the numbers written out in a list, not hidden inside an image.
This paper introduces a new way to test if Artificial Intelligence (AI) can read these charts and turn them into lists of numbers accurately. Here is the breakdown in simple terms:
1. The Problem: The AI is "Blind" to Time
Previously, researchers tested AI on chart-reading using simple, clean, computer-generated charts. The AI was doing great on those tests (scoring over 89%). But those tests were like a driving test on an empty, straight highway. Real-world epidemic charts are like driving through a chaotic, rainy city with blurry signs, overlapping lines, and tiny text.
The old way of grading the AI was like a teacher checking a student's homework by looking at the answers in random order.
- The Flaw: If the AI got the numbers right but wrote them down one day late (a "temporal shift"), the old grading system gave it a zero. It treated a "right answer, wrong day" the same as a "completely wrong answer." This is unfair because, in real life, knowing the trend is right is often more important than being off by a single day.
2. The Solution: A New Test and a New Grader
The authors created two things to fix this:
A. The Test: EpiCurveBench
They gathered 1,000 real-world epidemic charts from public health agencies around the world. These aren't clean, perfect images. They are messy:
- Some are scanned from old, blurry paper.
- Some have lines that cross over each other.
- Some have text that is rotated or tiny.
- Some have hundreds of data points (like a dense forest of dots).
- They cover diseases from 1849 to 2025.
B. The Grader: EpiCurveSimilarity (ECS)
They invented a new scoring system called ECS. Think of it like a rubber ruler instead of a rigid metal one.
- If the AI gets the shape of the curve right but is slightly shifted in time, the rubber ruler stretches to match it and gives partial credit.
- If the AI misses a chunk of the data (like cutting off the end of the chart), the ruler shrinks and penalizes the score heavily.
- This metric is designed to care about the story the data tells (the trend), not just if every single number is in the exact right box.
3. The Results: The AI Still Has a Long Way to Go
The authors tested six different AI models (three famous "closed" ones, one open-source one, and two specialized chart-readers) on this new, hard test.
- The Score: Even the best AI model only scored 52.3% on this new test. This means that even the smartest AI currently available is still making significant mistakes when trying to digitize real-world public health data.
- The Specialized vs. General: The models built specifically for charts (like "OneChart") actually did worse than the general-purpose AI models. They got confused by the messy, real-world styles.
- The "Reasoning" Trap: Telling the AI to "think harder" or use a calculator tool didn't always help. Sometimes, it made the AI crop the image wrong or get confused, lowering its score.
4. Why This New Grader Matters
The paper proves that the old grading system (RMS) was hiding the truth.
- The Old Grader: Put all the AI models in a tiny, crowded room where they all looked about the same (scores within a 5-point range). It couldn't tell who was actually better.
- The New Grader (ECS): Spread the models out into a large field (a 25-point range), clearly showing which ones were better at understanding the shape and timing of the data.
Most importantly, the authors checked if a higher ECS score actually meant better results for real-world math. They found that yes, it did.
- If an AI had a high ECS score, the total number of cases it calculated was closer to the truth.
- It was better at predicting when the peak of the outbreak happened.
- It was better at predicting how fast the disease was spreading.
The Bottom Line
We have a treasure trove of historical disease data trapped in pictures. We have AI that is getting better at reading them, but it's not ready for prime time yet. The authors built a harder test and a fairer grading system that shows us exactly where the AI is failing (like missing data points or getting numbers slightly wrong) so engineers can fix it. Until the AI can score higher on this new test, we can't fully trust it to unlock decades of vital health data.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.