← Latest papers
💻 computer science

LoMeVQA: A Comprehensive Benchmark for Longitudinal Medical VQA

This paper introduces LoMeVQA, a comprehensive benchmark comprising 206K longitudinal medical VQA pairs designed to evaluate temporal reasoning in multimodal models, alongside the proposed MedLong-8B model that achieves state-of-the-art performance on these tasks.

Original authors: Zhilin Wu, Zhangkai Ni, Chengmei Yang, Longzhen Yang, Yihang Liu, Ying Wen, Lianghua He

Published 2026-07-31
📖 3 min read☕ Coffee break read

Original authors: Zhilin Wu, Zhangkai Ni, Chengmei Yang, Longzhen Yang, Yihang Liu, Ying Wen, Lianghua He

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery, but instead of looking at a single crime scene photo, you have a whole album of snapshots taken over months or years. In the world of medicine, this "album" is called longitudinal data. It's the record of a patient's health journey, showing how their body changes from one doctor's visit to the next. For decades, artificial intelligence (AI) has been getting really good at looking at just one picture and saying, "That looks like a broken bone" or "That looks like a tumor." These AI systems, known as Multimodal Large Language Models (MLLMs), are like super-smart students who can read text and look at images at the same time. They are amazing at answering questions about a single moment in time. But real life isn't a single moment; it's a movie. Doctors don't just look at today's X-ray; they compare it to last month's to see if a disease is getting better, staying the same, or getting worse. The big question scientists have been asking is: Can these super-smart AI students actually watch the whole movie and understand the plot, or do they just get confused when the story changes?

This is exactly what the paper LoMeVQA tackles. The authors, a team of researchers from Tongji University and East China Normal University, realized that while AI is great at single snapshots, it's terrible at watching the "movie" of a patient's health. To prove this, they built a massive new playground called LoMeVQA (Longitudinal Medical Visual Question Answering). Think of it as a giant, 206,000-question quiz designed specifically to test if AI can spot changes over time. They didn't just ask simple questions; they created five different types of challenges, ranging from "Did the disease get better or worse?" (Progress Classification) to "Point exactly where the change happened on the image" (Differential Region Grounding).

To build this quiz, the researchers didn't just guess; they built a clever robot pipeline. They took real patient records from a huge database, organized them by date like a photo album, and used a medical "encyclopedia" (a knowledge graph) to pull out the important facts. Then, they asked a large language model to write questions and answers based on how those facts changed over time. After a strict quality check—where they even had human doctors review the best 2,500 questions to make sure they were accurate—they ended up with a gold-standard dataset.

When they put the top AI models to the test, the results were a bit of a shocker. Even the smartest medical AI models, which usually ace single-image tests, stumbled badly on this new quiz. They got about 55% right on simple "better or worse" questions and barely 25% right on the "point to the change" tasks. It turns out, these AIs are like students who memorized the textbook but can't follow a story with a twist. They often missed subtle changes or got the timeline mixed up.

To fix this, the authors created their own AI student, named MedLong-8B. They took a powerful existing model and "tutored" it specifically on their new 206,000-question quiz. The result? MedLong-8B became the star of the show, achieving the best scores ever recorded on this benchmark. The paper shows that while current AI struggles to understand the passage of time in medical images, it is possible to train models to get much better at it. The authors suggest that by giving AI the right kind of practice with longitudinal data, we can finally build systems that don't just see a picture, but truly understand a patient's story.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →