← Latest papers
💻 computer science

ChartAct: A Benchmark for Dynamic Chart Understanding

This paper introduces ChartAct, a new interactive benchmark comprising 1,440 high-quality samples across 7 chart types and two environments to evaluate dynamic chart understanding, revealing that current multimodal models and GUI agents still face significant limitations in handling interactive chart operations.

Original authors: Muye Huang, Wu Lin, Lingling Zhang, Hang Yan, Zhiyuan Wang, Yumeng Fu, Zesheng Yang, Jun Liu

Published 2026-05-27
📖 4 min read☕ Coffee break read

Original authors: Muye Huang, Wu Lin, Lingling Zhang, Hang Yan, Zhiyuan Wang, Yumeng Fu, Zesheng Yang, Jun Liu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: Charts That Move and Talk Back

Imagine you are looking at a static painting of a map. You can see the mountains and rivers, but you can't zoom in to see the tiny villages, and you can't click a river to see how fast the water is flowing. This is how most current AI models "see" charts today. They look at a frozen picture and try to guess the answer.

But in the real world, charts are more like interactive video games.

  • You might need to hover your mouse over a dot to see a secret number.
  • You might need to click a legend to hide a line and see another one clearly.
  • You might need to drag a slider to see how data changed over time.

The authors of this paper realized that AI is great at reading the "frozen painting" but terrible at playing the "interactive game." To fix this, they built a new test called ChartAct.

What is ChartAct? (The Video Game Test)

Think of ChartAct as a video game level designed specifically to test if an AI can actually play with a chart, not just look at it.

  1. The Setup: The researchers went out and collected 673 real charts from actual websites (like financial dashboards or weather sites). These aren't fake drawings; they are the real, interactive things people use every day.
  2. The Mission: They created 1,440 questions for these charts. Some questions are easy (you can see the answer immediately). But the tricky ones require the AI to take action.
    • Example Question: "What was the flow value at 4:00 on September 14, 2009?"
    • The Catch: That specific number is hidden. The AI has to zoom in to find the right date, then hover over the specific point to reveal the number. If it just guesses based on the blurry initial view, it fails.
  3. The Two Levels: To make it even harder, they tested the AI in two different "rooms":
    • The Clean Room (Dynamic Chart): The chart is alone on the screen. It's easy to find.
    • The Busy Office (Dashboard Chart): The same chart is buried inside a cluttered webpage full of other buttons, titles, and graphs. The AI has to find the right chart first, then interact with it. This is like finding a needle in a haystack while the haystack is moving.

How Did the AI Do? (The Scoreboard)

The researchers put 11 advanced AI models (including big names like Claude, GPT, and open-source models) through this test. Here is what happened:

  • The "Smart" Students: The best model, Claude-Opus-4.7, got about 84.5% of the answers right. It was good at figuring out, "Oh, I need to zoom in first," and then doing it.
  • The Struggling Students: Most other models scored below 60%. Many of them failed because they tried to guess the answer from the blurry initial picture without taking any action.
  • The "Clutter" Problem: When the charts were moved into the "Busy Office" (the Dashboard environment), every single model got worse. Some dropped by 30% or more. This shows that when there are too many distractions, AI gets confused about which chart to touch and which button to click.

Why Do They Fail? (The Three Mistakes)

The paper found that AI models usually trip up in three specific ways:

  1. The "Lazy Reader": The model sees the question, looks at the initial picture, and guesses an answer without ever clicking or hovering. It's like trying to read a menu in a dark restaurant without turning on the light.
  2. The "Clumsy Mouse": The model knows it needs to hover, but it hovers over the wrong dot. It's like trying to pick a specific apple from a tree but grabbing the one next to it instead.
  3. The "Lost Tourist": In the busy dashboard, the model gets confused by the surrounding text and clicks on the wrong chart entirely. It's like trying to find a specific shop in a mall but walking into the wrong store because the signs looked similar.

The Conclusion

The paper concludes that while AI is getting very good at understanding static images, it is still struggling with dynamic, interactive data.

To truly understand the world's data, AI needs to learn how to interact—to click, zoom, and explore—just like a human does. ChartAct is the new "driving test" to see if AI can learn those skills. Right now, most AI drivers are still learning how to turn the steering wheel without crashing into the dashboard.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →