InterChart: Benchmarking Visual Reasoning Across Decomposed and Distributed Chart Information
The paper introduces InterChart, a diagnostic benchmark designed to evaluate vision-language models' ability to reason across multiple related charts, revealing that current state-of-the-art models struggle significantly with cross-chart integration and complex multi-step reasoning despite improved performance on decomposed visual tasks.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to solve a mystery, but instead of looking at one clue, you have to look at three different maps, a weather report, and a stock ticker all at once to figure out the answer. That is the challenge INTERCHART sets for artificial intelligence.
Here is a simple breakdown of what the paper does, using everyday analogies.
The Big Problem: AI is Good at One Thing, Bad at Many
Current AI models (called Vision-Language Models) are like students who are great at reading a single textbook page and answering a question about it. If you show them one chart (like a bar graph of sales), they can usually tell you the highest number.
But in the real world, data is rarely just one chart. Scientists, bankers, and journalists often use multiple charts together to tell a story. One chart might show temperature, another shows ice cream sales, and a third shows rainfall. To understand the full picture, you have to connect the dots between them.
The paper argues that current AI models are terrible at this "connecting the dots" task. They get confused when they have to look at two or three different charts and figure out how they relate to each other.
The Solution: A New "Gym" for AI (INTERCHART)
The researchers built a new testing ground called INTERCHART. Think of this as a gym with three different levels of difficulty, designed to test how well AI can handle complex visual puzzles.
Level 1: The "Decomposed" Warm-up (DECAF)
- The Analogy: Imagine a messy room full of toys. In this level, the researchers take a messy room and organize it into neat, separate boxes.
- What it tests: They take complex charts and break them down into simple, single-variable pieces. They ask the AI simple questions like, "What is the number on this specific line?"
- The Result: AI does pretty well here. When the information is clean and separated, the AI can find the facts.
Level 2: The "Synthetic" Middle Ground (SPECTRA)
- The Analogy: Now, imagine two different maps of the same city, but one is drawn in blue and the other in red. They show related things (like traffic and weather), but they look different.
- What it tests: The AI has to look at two charts that are related (e.g., "How does rain affect traffic?") but are drawn in different styles. It has to realize that "Rain" on the first chart matches "Precipitation" on the second.
- The Result: The AI starts to stumble. It struggles to realize that the two different-looking charts are talking about the same thing.
Level 3: The "Real-World" Boss Fight (STORM)
- The Analogy: This is the hardest level. Imagine looking at a newspaper article with three different charts from different years, drawn by different people, with different colors and labels. You have to figure out a trend that spans all of them.
- What it tests: This uses real charts from the internet (like economic reports). The charts are messy, the labels might not match perfectly, and the AI has to do "time travel" reasoning (e.g., "In 2015, Country A was higher than Country B, but by 2020, they swapped. What happened in between?").
- The Result: The AI performance crashes. Even the smartest models get confused. They can't handle the messiness and the need to connect ideas across different visual styles.
Key Discoveries from the "Gym"
- Simplification Helps: The paper found that if you take a messy chart and turn it into a simple text table (like a spreadsheet) before asking the AI, the AI gets smarter. It's like giving the detective a written list of clues instead of a messy pile of photos. The AI loves tables; it hates messy images.
- More Steps = More Confusion: The more steps the AI has to take to connect the dots (e.g., "Look at Chart A, then Chart B, then compare them, then guess the future"), the more likely it is to fail.
- Real vs. Fake: AI is okay at reasoning on "fake" charts made by computers (which are neat and perfect), but it fails miserably on "real" charts from the news (which are messy and imperfect). This means AI hasn't truly learned to reason; it just memorized patterns from clean data.
- The "Judge" Problem: The researchers also realized that simply checking if the AI's answer matches the correct answer word-for-word isn't fair. If the answer is "25%" and the AI says "about a quarter," a computer might say "Wrong," but a human would say "Right." So, they used other AIs as "judges" to check if the meaning was correct, not just the spelling.
The Bottom Line
The paper concludes that while AI is getting better at looking at pictures, it is still very bad at synthesizing information from multiple, messy sources. It's like a student who can read a single sentence perfectly but fails when asked to write an essay connecting three different books.
The INTERCHART benchmark is a tool to show researchers exactly where and why these AI models are failing, so they can build better ones that can handle the messy, multi-chart reality of the real world.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.