← Latest papers
🤖 AI

Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation

The paper introduces Messier, a unified high-resolution corpus of nearly one million standardized records spanning 30 benchmarks and diverse domains, which enables comprehensive cross-benchmark agent evaluation, reveals uneven progress across task types, and provides a foundational infrastructure for auditing and scaling AI agent capabilities.

Original authors: Stefan Krsteski, Charlotte Meyer, Guillaume Allegre, Tony O'Halloran, Alexandre Sallinen

Published 2026-07-29
📖 5 min read🧠 Deep dive

Original authors: Stefan Krsteski, Charlotte Meyer, Guillaume Allegre, Tony O'Halloran, Alexandre Sallinen

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to figure out which video game character is the absolute best at everything. You have a character who is amazing at jumping, another who is a wizard at solving math puzzles, and a third who can navigate a maze perfectly. But here's the catch: every game measures "skill" differently. One game gives you a gold star for jumping over a single pit, while another requires you to solve a complex algebra problem to open a door, and a third demands you collect every single coin in a level without touching a single enemy. If you just look at the final score—a single number like "95%"—you might think they are all equally good. But that number hides the messy details. Did the character fail because they couldn't jump, or because they got stuck on a door? Did they get 95% because they did everything perfectly, or because they failed one tiny thing and the game counted the whole run as a zero?

This is the exact problem facing scientists who study AI agents. These are not just chatbots; they are digital workers that can use tools, write code, browse the web, and make decisions in real-time environments. Right now, there are hundreds of different tests (benchmarks) for these agents, but they all speak different languages. Some tests are strict: if you miss one step in a 10-step recipe, you get a zero. Others are more forgiving. Because of this, it's incredibly hard to compare an agent that is great at coding with one that is great at legal research. It's like trying to compare a marathon runner to a chess grandmaster using only a single "athleticism" score. To understand how AI is actually improving, we need a way to look under the hood of these scores and see exactly where the agents succeed and where they stumble.

Enter MESSIER, a massive new project that acts like a giant, universal translator for AI testing. The researchers behind MESSIER didn't just build a new test; they built a giant library that collects the results from 30 different existing tests, covering 11,891 specific tasks and 74,205 different ways of checking if an agent succeeded. They gathered data on 714 different AI agents and even ran new, fresh tests on six difficult areas where data was missing, like legal reviews and quantum circuit design. In total, they organized nearly one million trial records into one neat, standardized format. Think of it as taking a chaotic pile of receipts from 30 different stores, each written in a different currency and with different math rules, and turning them all into a single, clear spreadsheet that lets you see exactly what was bought, how much it cost, and which store had the best deals.

The most exciting thing MESSIER found is that the way we count points changes the story. The paper shows that many tests use a "strict" rule: if an agent fails even one small part of a task, the whole thing is marked as a failure. The researchers showed that if they relaxed this rule and looked at how many parts the agent did get right, the picture of progress looks very different. For example, on a legal benchmark, the strict "all-or-nothing" score said agents were succeeding only 0.2% of the time. But when they looked at the details, they found that on average, agents were actually satisfying 28.8% of the criteria. The strict rule was hiding a lot of partial progress. This suggests that while AI is getting better, the way we measure it might be too harsh, making it look like agents are failing more often than they really are.

The study also mapped out where AI is truly "winning" and where it is still struggling. They found that AI is almost perfect at "function calling" (using tools), with a success rate of 0.97. It is also getting very good at programming, improving rapidly over the last two years. However, AI is still having a tough time with "enterprise workflows"—complex, multi-step office jobs like managing a business process. In these areas, the success rate is only 0.57, meaning agents still fail more often than they succeed. The researchers also proved that you can use this massive library to create a new "skill scale" for AI that matches up very closely with other major rankings (with a correlation of 0.81), without needing to spend millions of dollars running new tests.

Ultimately, MESSIER is a tool for clarity. It suggests that by looking at the tiny details of how an agent fails, rather than just the final score, we can get a much truer picture of what AI can actually do. It doesn't say AI has solved everything, but it does suggest that our current measuring sticks might be too blunt. By standardizing these results, the authors hope to stop the waste of money on repeating the same tests and help researchers understand exactly which jobs AI is ready for and which ones still need a human hand. The data is now open for anyone to use, turning a confusing mess of numbers into a clear map of the AI landscape.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →