← Latest papers
💻 computer science

DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces

This paper introduces DataSpace, a comprehensive benchmark and KDD Cup 2026 evaluation framework designed to assess data agents' ability to generate verifiable tabular results from heterogeneous, multi-modal workspaces, revealing significant performance gaps in multimodal evidence integration and the substantial impact of agent harness choices on reliability.

Original authors: Boyan Li, Zhuowen Liang, Yupeng Xie, Xiaotian Lin, Tianqi Luo, Xinyu Liu, Yizhang Zhu, Zhangyang Peng, Yuan Li, Zhengxuan Zhang, Jiayi Zhang, Nan Tang, Guoliang Li, Yuyu Luo

Published 2026-08-05
📖 3 min read☕ Coffee break read

Original authors: Boyan Li, Zhuowen Liang, Yupeng Xie, Xiaotian Lin, Tianqi Luo, Xinyu Liu, Yizhang Zhu, Zhangyang Peng, Yuan Li, Zhengxuan Zhang, Jiayi Zhang, Nan Tang, Guoliang Li, Yuyu Luo

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery, but your clues are scattered everywhere. Some are in a neat spreadsheet on your computer, others are buried in a 50-page PDF report, some are hidden inside a database, and the most important clue might be a video of a security camera. This is the daily reality for "data agents"—smart computer programs designed to answer questions by hunting through an organization's messy digital files. For a long time, scientists tested these agents on clean, single-type puzzles, like finding a name in a phone book or asking a database a simple question. But in the real world, data is a chaotic mix of languages, formats, and media. The big question researchers are asking is: Can these digital detectives actually piece together a complete story from a jumbled pile of evidence, or do they get lost when the clues don't look the same?

This paper introduces a new, super-challenging test called DataSpace to see exactly how good these data agents are at solving these messy, real-world mysteries. The researchers built a massive playground containing 410 different tasks and 7,439 digital artifacts (files) totaling 15.01 GB of data. They didn't just throw random files together; they created a "heterogeneous workspace" where a single question might require you to read a video, scan a PDF, query a database, and cross-reference a JSON file, all while dealing with clues written in both English and Chinese. The goal isn't just to find a single fact; the agent must produce a complete, verifiable table of answers, like a final report card.

The results of this test are a bit of a wake-up call. Even the smartest, most advanced AI models tested (including giants like Grok 4.5 and GPT-5.6) only managed to solve about 66.34% of the tasks correctly. That means they got roughly one out of every three puzzles wrong. The researchers found that the agents struggle most when they have to combine information from different types of media (like joining a video clue with a spreadsheet) or when they need to perform complex "joins" (connecting data from two different sources). Interestingly, the choice of the "harness"—the software framework that tells the AI how to use its tools—mattered almost as much as the AI brain itself, causing a 15.36-point difference in scores.

Ultimately, the paper shows that while data agents are getting better, they are far from perfect. They often get the right numbers but fail to format the final table correctly, or they miss a crucial clue hidden in a video. The authors conclude that we still have a long way to go before these agents can be fully trusted to handle complex, multi-format data analysis on their own, and they've provided this benchmark as a map to help future engineers fix these specific weaknesses.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →