OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding
This paper introduces OmegaUse-OfficeVal, a new benchmark comprising 100 long-horizon office-suite tasks with economic grounding (human labor time and task price) to evaluate LLM agents, revealing that while current models are significantly cheaper and faster than humans, they still fall short of achieving human-level quality in task completion.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a world where your computer doesn't just wait for you to click buttons, but actually does the work for you. This is the dream of "AI agents"—smart programs that can read your instructions, open your files, and finish a job, much like a digital intern. For a while, these agents have been great at short, simple tricks, like writing a single email or summarizing a short article. But the real test of their power is whether they can handle a whole day's worth of office work: taking a messy report, reformatting it, adding charts, and saving it perfectly, all without getting confused or breaking the file. The big question researchers are asking is: Can these digital workers actually do the job, and is it worth the money compared to hiring a real human?
Enter OmegaUse-OfficeVal, a new "report card" designed by researchers at Baidu to test exactly this. They didn't just make up fake, easy tasks; they built a benchmark using 100 real-world office challenges, like fixing a broken spreadsheet or turning a rough draft into a polished presentation. What makes this test special is that they attached a "price tag" to every single task. They measured exactly how long a human took to do it and estimated how much it would cost to pay someone to do it. Then, they let several of the world's smartest AI models try to solve these same puzzles. The results? The AI agents are incredibly fast and cheap—often finishing in minutes for pennies—but they still make too many mistakes to be trusted with the final product. While they are getting better, they haven't quite reached the level of a reliable human junior worker yet.
The Setup: A Digital Obstacle Course
To understand what the researchers did, think of the office as a giant, messy workshop. In the past, testing AI was like asking it to stack a few blocks. If it dropped one, it failed. But real office work is more like baking a complex cake from scratch: you have to gather ingredients (files), follow a recipe (instructions), and if you burn the frosting, the whole cake is ruined.
The researchers created a dataset called OmegaUse-OfficeVal containing 100 of these "cake-baking" tasks. These weren't simple "write a poem" requests. They were long, multi-step jobs involving Word documents, PowerPoint slides, and Excel spreadsheets. Some tasks took humans an average of 2.32 hours to complete. That's a long time for a computer to keep its focus!
What makes this benchmark unique is its "economic grounding." The researchers didn't just ask, "Did the AI finish the task?" They asked, "How much value did it create?" To do this, they tracked two things for every task:
- Human Labor Time: How long it took a real person to do it.
- Task Price Proxy: A rough estimate of what that task would cost if you hired a freelancer to do it (ranging from about $3.65 to $29.22 per task).
This allowed them to compare the AI not just on speed, but on whether it was actually saving money or just making cheap mistakes.
The Test: AI vs. The Human Intern
The researchers pitted several top-tier AI models against a "human baseline." This baseline wasn't a CEO or a genius; it was a skilled junior worker or intern. The humans were given the same instructions and files as the AI and asked to do the work.
To grade the results, the researchers didn't just have a human look at the final file and say, "Looks good!" That's too slow and subjective. Instead, they wrote code-based verifiers. Imagine a robot inspector that automatically checks the final file. It opens the document, checks if the cover page is there, counts if the charts are the right size, and makes sure no text was accidentally deleted. If the file is broken or missing a key part, the robot gives it a zero. This made the testing fair, fast, and repeatable.
The Results: Fast and Cheap, But Not Perfect
When the dust settled, the results told a clear story.
The Human Baseline:
The human workers scored an average of 27.79 out of a possible perfect score. They were reliable. They rarely broke the files, and they got most of the details right. However, they took time and money. On average, a human task cost about $6.86 and took 2.32 hours.
The AI Agents:
The AI models were a different story. They were lightning fast and incredibly cheap.
- Speed: The fastest AI (DeepSeek-V4-Pro) finished a task in just 0.184 hours (about 11 minutes).
- Cost: The cheapest AI (Qwen3.7-Plus) cost only $0.2152 per task. That's less than a cup of coffee!
But here is the catch: They weren't very good at the work.
The best AI model (GLM-5.2) only scored 17.91. That is significantly lower than the human score of 27.79. While the AI was faster and cheaper, it made more mistakes. It often forgot to add a table of contents, messed up the formatting, or created a file that couldn't be opened.
The researchers found that the AI struggled the most with the hardest, longest tasks. When a task required a human to work for a long time (over 2 hours), the AI's performance dropped even further. It seems that while AI is great at quick, simple jobs, it gets lost when the "long-horizon" planning gets complicated.
The Verdict: Not Ready for Prime Time (Yet)
The study suggests that we are in a "valley" of AI office work. The technology is undeniably useful because it is so cheap and fast. If you need a rough draft or a quick summary, an AI is a fantastic tool. But if you need a polished, professional document ready to send to a client, the AI isn't quite there yet.
The authors explicitly ruled out the idea that AI has already solved office work. They showed that while the AI is "substantially cheaper and faster," it has "not yet approached human-level deliverable quality." The gap between a human worker and an AI agent is still wide when it comes to the final quality of the product.
However, the paper is optimistic about the future. By creating this rigorous test with real economic data, the researchers have built a roadmap. They know exactly where the AI is failing (long tasks, complex formatting) and can now work to fix those specific problems. For now, the best strategy seems to be a partnership: let the AI do the heavy lifting and the quick drafts, but keep a human in the loop to check the final product and fix the mistakes. The dream of "vibe working"—where AI just handles the whole job—is still a work in progress, but the journey has officially begun.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.