← Latest papers
💻 computer science

MI-CXR: A Benchmark for Longitudinal Reasoning over Multi-Interval Chest X-rays

The paper introduces MI-CXR, a new benchmark designed to evaluate the longitudinal reasoning capabilities of vision-language models over multi-visit chest X-ray sequences, revealing that current state-of-the-art models struggle with temporal consistency and global trajectory summarization despite achieving only modest accuracy above random guessing.

Original authors: Sunghwan Steve Cho, Yunseok Han, Jaeyoung Do

Published 2026-05-18
📖 5 min read🧠 Deep dive

Original authors: Sunghwan Steve Cho, Yunseok Han, Jaeyoung Do

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a doctor trying to understand a patient's health story. Usually, you don't just look at one X-ray taken today. You look at a series of X-rays taken over weeks or months to see how a disease started, how it changed, and whether treatment worked. This is called longitudinal reasoning.

The paper introduces a new test called MI-CXR to see if modern AI "doctors" (specifically Vision-Language Models) can actually do this kind of storytelling.

Here is the breakdown of what the paper found, using simple analogies:

1. The Problem: AI is Good at Snapshots, Bad at Movies

Current AI models are great at looking at a single photo and saying, "That looks like pneumonia." They are also okay at comparing two photos and saying, "This one looks worse than that one."

But real life isn't a snapshot or a simple "before and after." It's a movie with many scenes. The paper argues that most AI tests only check the snapshots or the two-photo comparisons. They miss the hard part: connecting the dots across a whole timeline.

2. The New Test: MI-CXR

The researchers built a new benchmark (a standardized test) called MI-CXR.

  • The Setup: Instead of showing the AI one or two images, they show it five X-rays taken from the same patient over time (like five frames from a movie).
  • The Goal: The AI has to answer questions about the entire story, not just one frame.

The test is divided into three types of "challenges," which the authors call Task Families:

  • Temporal Event Localization (The "When" Game):
    • Analogy: Imagine a crime scene investigation. You have five security camera clips. The question is: "Exactly when did the thief enter the room?"
    • The Task: The AI must pinpoint the specific time interval (e.g., between Visit 2 and Visit 3) when a disease first appeared or disappeared. It can't just say "it happened sometime"; it must pick the one correct interval.
  • Interval-wise Change Reasoning (The "What Changed" Game):
    • Analogy: You are a referee watching a soccer game. You need to say, "Between the 10th and 20th minute, the team's defense got weaker."
    • The Task: The AI looks at two specific visits and describes exactly what changed between them. The tricky part is that the AI has to figure out which pair of visits the question is talking about first, then describe the change.
  • Global Trajectory Summarization (The "Story" Game):
    • Analogy: A sports commentator summarizing the whole season. "The team started strong, had a slump in the middle, and finished strong."
    • The Task: The AI must look at all five visits and write a summary of the patient's entire health journey. Did the disease get better overall? Did it come back?

3. The Results: The AI Got Lost

The researchers tested 14 of the smartest AI models available (including famous ones like GPT-5, Claude, and specialized medical AIs).

  • The Score: The average score was 29.3%.
  • The Reality Check: Since these are multiple-choice questions with 5 options, a random guess would get you 20%. The AI models were only slightly better than a monkey throwing darts at a board.

Why did they fail?
The researchers dug deeper to find out why the AI struggled. They found that the AI wasn't failing because it couldn't "see" the disease.

  • Local vs. Global: If you asked the AI to look at just two images and describe the change, it did much better. It could see the details.
  • The Bottleneck: The problem was connecting the dots. The AI could describe what happened between Visit 1 and 2, and then between Visit 2 and 3, but it couldn't put those pieces together to form a consistent story for the whole timeline.
    • It would get confused about when something happened.
    • It would get confused if a disease appeared twice.
    • It would fail to keep the story consistent (e.g., saying a disease disappeared, then later saying it was still there, without realizing the contradiction).

4. The Conclusion: A "Diagnostic Tool" for AI

The paper concludes that current AI models have a major blind spot. They are like students who can memorize individual facts but fail a history exam because they can't understand the sequence of events.

The authors created MI-CXR not to say "AI is useless," but to provide a diagnostic tool. Just as a doctor uses a specific test to find a broken bone, this benchmark helps researchers see exactly where the AI's reasoning breaks down (is it the vision? the memory? or the logic?).

Key Takeaway:
AI is currently very good at looking at a single picture, but it is still very bad at watching a movie and understanding the plot. Until AI can reliably connect the dots across time, it cannot be trusted to make complex medical decisions that require tracking a patient's history over weeks or months.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →