APEX-Agents
The paper introduces APEX-Agents, an open-source benchmark and evaluation infrastructure designed to assess the capabilities of AI agents in executing complex, long-horizon, cross-application tasks typical of investment banking, management consulting, and legal work, revealing that current top models achieve a maximum Pass@1 score of 24.0%.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you're hiring a new assistant to run a complex business project. You don't just want them to answer a simple question like "What's the weather?" You want to see if they can:
- Find a specific file in a messy digital filing cabinet.
- Open a spreadsheet, do some math, and update a chart.
- Draft an email to a client based on that data.
- Do all of this without getting confused, deleting the wrong files, or giving up halfway through.
That is exactly what this paper, APEX–Agents, is about. It's a giant "final exam" for AI agents (smart computer programs that can do tasks for you) to see if they are ready for real-world office work.
Here is the breakdown of the paper using some simple analogies:
1. The Test: "The Digital Office Simulator"
Most previous AI tests were like asking a student to solve a math problem on a whiteboard. It's clean, controlled, and doesn't look like real life.
APEX–Agents is different. The researchers built 33 "Digital Worlds."
- The Setup: Imagine a video game where you play a consultant, a lawyer, or a banker. You have a computer with a calendar, email, spreadsheets, PDFs, and a code editor.
- The Mission: You are given a realistic, messy project (e.g., "Analyze this company's growth potential and send a report to the CEO").
- The Catch: The AI has to navigate this digital office just like a human would. It can't just "know" the answer; it has to find the data, open the right file, edit it, and send it.
2. The Teachers: Real Humans, Not Robots
To make sure the test was fair and realistic, the researchers didn't just ask AI to write the questions. They hired 256 real experts (investment bankers, management consultants, and corporate lawyers) to build the test.
- These experts spent 5–10 days working on fake projects to create the "Digital Worlds."
- They then wrote the specific tasks the AI had to solve.
- They also created a "grading rubric" (a checklist) to decide if the AI did a good job.
3. The Students: The AI Models
Eight different AI models took this test. Think of them as students with different study habits:
- The Big Names: Models like Gemini 3 Flash, GPT-5.2, and Claude Opus 4.5.
- The Open Source: Smaller, free models like GPT-OSS and Kimi K2.
4. The Results: "The Good, The Bad, and The Inconsistent"
The results were surprising but honest.
- The Score: Even the best AI only got about 24% of the tasks perfect on the first try.
- Analogy: If you hired the smartest AI assistant today to do a week's worth of complex analyst work, they would mess up about 3 out of every 4 assignments if you only gave them one chance.
- The "Try Again" Factor: If you let the AI try the same task 8 times, the best models could get about 40% right.
- Analogy: They aren't totally useless, but they are very inconsistent. Sometimes they get it right; other times they get stuck in a loop or delete the wrong file.
- The Open Source Struggle: The free, open-source models scored under 5%. They were like students who hadn't studied at all compared to the top-tier private models.
5. How They Failed (The "Rogue" Behavior)
The paper looked closely at how the AI failed, which is just as important as the score:
- The "Doom Loop": Some AIs got stuck repeating the same action over and over (like a hamster on a wheel) until they ran out of time.
- The "Accidental Arsonist": Some AIs deleted files they weren't supposed to touch.
- The "Over-Thinker": The most successful AI (Gemini 3 Flash) used 5 times more computer power (tokens) than its competitors to get the job done. It was effective, but very inefficient.
6. The Big Takeaway
The paper concludes that AI agents are not ready to replace human professionals yet.
- Current State: They are like a very smart intern who has read every book in the library but has never actually worked in an office. They know the theory but struggle with the messy reality of clicking buttons, managing files, and planning long-term projects.
- The Future: The researchers are releasing all their data and tools for free (open source) so other scientists can help build better agents. They believe that with more training, these agents could eventually become the "on-demand team of experts" that changes how we work.
In short: This paper built a realistic "office simulator" to test AI. The results show that while AI is getting smarter, it still gets lost easily, deletes important files, and needs a lot of tries to get a simple professional job done correctly. We are on the right track, but we aren't there yet.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.