← Latest papers
💬 NLP

ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows

The paper introduces ScienceBoard, a realistic multi-domain environment and a benchmark of 169 rigorously validated scientific tasks designed to evaluate multimodal autonomous agents, revealing that current state-of-the-art models achieve only a 15% success rate in complex scientific workflows and highlighting the need for improved design principles to support automated scientific discovery.

Original authors: Qiushi Sun, Zhoumianze Liu, Chang Ma, Zichen Ding, Fangzhi Xu, Zhangyue Yin, Haiteng Zhao, Zhenyu Wu, Kanzhi Cheng, Zhaoyang Liu, Jianing Wang, Qintong Li, Xiangru Tang, Tianbao Xie, Xiachong Feng, Xi
Published 2026-04-21
📖 4 min read☕ Coffee break read

Original authors: Qiushi Sun, Zhoumianze Liu, Chang Ma, Zichen Ding, Fangzhi Xu, Zhangyue Yin, Haiteng Zhao, Zhenyu Wu, Kanzhi Cheng, Zhaoyang Liu, Jianing Wang, Qintong Li, Xiangru Tang, Tianbao Xie, Xiachong Feng, Xiang Li, Ben Kao, Wenhai Wang, Biqing Qi, Lingpeng Kong, Zhiyong Wu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a brilliant, super-smart robot assistant. You tell it, "Go to the library, find a book about black holes, open it to page 42, and read the diagram to me."

In the world of computers, we've been teaching these robots to do simple things for a long time. But now, scientists want to see if these robots can handle real, messy, complex scientific work. Can they act like a real researcher, using specialized software, clicking through menus, typing commands, and actually doing the science?

This paper, SCIENCEBOARD, is the answer to that question. It's like a giant, high-tech training gym built specifically to test these "AI Co-Scientists."

Here is the breakdown in simple terms:

1. The Gym: SCIENCEBOARD

Think of SCIENCEBOARD as a virtual laboratory. Inside this lab, there are six different "workstations," each representing a different field of science:

  • Biochemistry: Like a digital microscope for looking at proteins (using software called ChimeraX).
  • Astronomy: A 3D planetarium where you can zoom in on stars and planets (using Celestia).
  • Geography: A map maker for analyzing terrain and water flow (using GrassGIS).
  • Math & Logic: Tools for solving complex equations or proving mathematical theorems.
  • Writing: A system for writing scientific papers.

The catch? The AI agents can't just "think" about the answer. They have to physically interact with the computer, just like a human would. They have to move the mouse, click buttons, type in the terminal, and wait for the software to load.

2. The Test: 169 Real-World Challenges

The researchers created 169 specific tasks for the robots to solve. These aren't riddles; they are real jobs a scientist might do.

  • Example: "Find this specific protein structure, run a simulation to see how it folds, and tell me the result."
  • Example: "Zoom out from the Sun in our solar system simulation and show me the orbit of Mars."

The tasks range from "Easy" (click a button) to "Hard" (a multi-step process requiring logic, math, and navigating complex menus).

3. The Players: The AI Agents

The researchers put the world's smartest AI models (like GPT-5, Gemini, and various open-source models) into this gym to see how they performed. They tested them in different ways:

  • The "Eyes Only" approach: The AI just looks at screenshots of the screen.
  • The "Text Only" approach: The AI reads a list of what's on the screen (like a blind person reading a menu).
  • The "Hybrid" approach: The AI sees the picture and reads the menu text.

4. The Results: A Reality Check

Here is the big news: The robots are still very clumsy.

Even the smartest, most expensive AI models only succeeded about 20% of the time on average.

  • The Good News: They are getting better at math and biochemistry.
  • The Bad News: They struggle terribly with geography and astronomy. Why? Because those fields rely heavily on visuals (maps, star charts). The AI gets confused by the clutter on the screen. It can't tell the difference between a star and a speck of dust, or a mountain range and a river.

The Analogy: Imagine asking a brilliant mathematician to drive a car through a busy city. They know the rules of the road (the logic), but they can't see the traffic lights or the pedestrians (the visual grounding). They keep crashing because they can't translate their smart thoughts into safe, physical actions.

5. Why Does This Matter?

The paper concludes that while AI is amazing at chatting and solving riddles, it is not yet ready to be a Scientific Co-Pilot.

To fix this, the researchers suggest we need to build "Team AI" instead of "Solo AI."

  • The Planner: One AI that is good at logic and planning the steps.
  • The Doer: A different AI that is specialized in seeing the screen and clicking the right buttons.

The Bottom Line

SCIENCEBOARD is a wake-up call. It shows us that giving AI a mouse and a keyboard isn't enough. To truly automate science, we need to teach these robots how to see the world through a computer screen and how to handle the messy, unpredictable nature of real scientific software.

We are currently in the "training camp" phase. The robots are learning, but they are far from being the Nobel Prize-winning scientists of the future.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →