Helping Figures Tell their Story! Paper-Grounded Video Generation Explaining Complex Scientific Figures
This paper introduces MINARD, a novel pipeline and the FigTalk benchmark designed to generate narrated, region-grounded walkthrough videos from scientific figures and their source papers, effectively addressing the gap in current systems for explaining complex visual pipelines through step-by-step, paper-faithful storytelling.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are at a science conference. The speaker flashes a complex diagram on the screen—a tangle of boxes, arrows, and labels representing a new computer system. They say, "The figure speaks for itself," and move to the next slide before you've even figured out where to look.
This paper argues that AI should be able to do what a good human speaker does: take that static, confusing picture and turn it into a step-by-step video tour that explains what is happening, why it matters, and where to look at every moment.
Here is the breakdown of their solution, MINARD, and their new testing ground, FigTalk, using simple analogies.
The Problem: The "Static Map" vs. The "Tour Guide"
Current AI systems treat scientific figures like a static photograph. They might say, "This is a box with an arrow," but they miss the story.
- The Missing Piece: They don't know why the arrow points there, or that the orange box is the most important part because of a specific sentence in the research paper.
- The Hallucination Risk: If you ask a standard video AI to explain a diagram, it might invent parts of the diagram that don't exist (like adding a red arrow where there is none) because it's trying to "guess" the story rather than reading the source material.
The Solution: MINARD (The Three-Person Team)
The authors built a system called MINARD (named after a famous 19th-century mapmaker who visualized complex data beautifully). Instead of one giant AI trying to do everything at once, MINARD acts like a small production team with three specific jobs:
The Narrator (The Scriptwriter):
- Job: Reads the entire research paper and looks at the figure.
- Analogy: Imagine a tour guide who has read the entire history book of the museum before stepping in front of the exhibit. They don't just say, "Here is a painting." They say, "This painting shows the battle of 1812, and notice how the general is pointing here because..."
- Key Feature: They use a "Critic" team to fact-check the script, ensuring the AI doesn't make up facts or skip important steps.
The Perception Team (The Spotter):
- Job: Looks only at the figure (ignoring the script for now) to find every single valid part: every text label, every box, and every arrow.
- Analogy: This is like a stage manager who has a master list of every prop on stage. They create a "whitelist" of valid things to point at. This prevents the AI from pointing at empty space or inventing a box that isn't there.
The Grounding Team (The Director):
- Job: Takes the script from the Narrator and the list of props from the Spotter, then decides exactly when to highlight which part.
- Analogy: This is the director saying, "Okay, when the guide says 'look at the engine,' the spotlight must hit only the engine, not the whole car." They ensure the highlight matches the words perfectly.
The Result: A Faithful Walkthrough
When MINARD creates a video:
- It doesn't redraw the picture (which often leads to errors).
- It keeps the original image and simply overlays highlights (like a laser pointer) that move in sync with the voiceover.
- It explains the "why" by pulling facts from the paper, not just describing the "what" from the image.
The Test: FigTalk (The Exam)
To prove their system works, the authors created FigTalk, the first "exam" for this specific task.
- The Setup: They collected 114 real scientific figures and the videos of humans explaining them in conferences.
- The Challenge: They asked AI systems to recreate these explanations.
- The Metric: They didn't just check if the AI sounded smart; they checked if the AI pointed at the right thing at the right time (Grounding) and if the story was factually correct based on the paper (Faithfulness).
The Findings
- The "Hard" Figures Win: On simple diagrams, other AIs sometimes do okay. But on complex diagrams with many steps (the "Hard" tier), MINARD shines. It stays accurate while other systems get confused, point at the wrong things, or invent fake details.
- Human Preference: When people watched the videos, they preferred MINARD's explanations 74% of the time over other methods. They found it easier to follow and more trustworthy.
- The "Paper" Matters: The biggest improvement came from making the AI read the paper. Without the paper, the AI could describe the picture but couldn't explain the science behind it.
What It Is NOT (Limitations)
The authors are clear about what this system doesn't do yet:
- It is designed for architecture diagrams (flowcharts, system maps). It is not currently built to explain bar charts, graphs, or statistical plots (which require a different kind of explanation).
- It is slower than some other methods because it takes time to "think" and fact-check, but this slowness is what makes it accurate.
In short: MINARD is a system that turns a confusing scientific diagram into a clear, fact-checked video tour by having a team of AIs work together: one to write the script from the paper, one to find the parts, and one to direct the spotlight.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.