← Latest papers
💬 NLP

Office Comprehension Benchmark

This paper introduces Office Comprehension Bench (OCB), the first public benchmark designed to evaluate LLMs on Word, Excel, and PowerPoint comprehension through two tracks—File Fidelity Q&A and Domain Q&A—revealing that even top-tier models struggle with complex document reasoning, achieving only 59.3% accuracy on domain tasks.

Original authors: Firoz Shaik, Mateus Picanço Lima Gomes, Tanvir Aumi, Jingci Wang, Milos Milunovic, Filip Basara, Ivana Jovanovic, Vishwas Suryanarayanan, Neha Nandan Kenkare, Weiyao Xie, Zhipeng Han, Zheng Zhang, Wal
Published 2026-07-03
📖 5 min read🧠 Deep dive

Original authors: Firoz Shaik, Mateus Picanço Lima Gomes, Tanvir Aumi, Jingci Wang, Milos Milunovic, Filip Basara, Ivana Jovanovic, Vishwas Suryanarayanan, Neha Nandan Kenkare, Weiyao Xie, Zhipeng Han, Zheng Zhang, Waleed Shahid, Jay Rathi, Russell Scherer, Thong Q. Nguyen, Michael Bentley, Tamara Stankovic, Rasika Chakravarthy, Vishal Chowdhary

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a giant, super-smart robot assistant that can read documents. You want to know: Is it actually good at reading real office files, or is it just guessing?

Microsoft researchers built a giant test called Office Comprehension Bench (OCB) to find out. Think of this test as a "driver's license exam" for AI, but instead of driving a car, the AI has to drive through Word documents, Excel spreadsheets, and PowerPoint slides.

Here is how the test works, broken down into simple parts:

1. The Two Types of Tests

The exam has two different "tracks," like two different levels of a video game.

  • Track 1: The "Eagle Eye" Test (File Fidelity)

    • The Goal: Can the AI see exactly what is on the page?
    • The Analogy: Imagine you hand the AI a printed menu and ask, "What is the price of the soup?" or "Is the soup listed in bold text?"
    • What it checks: Can the AI find specific numbers in a table? Can it spot a chart? Can it read the tiny speaker notes at the bottom of a slide? It's about spotting and reading facts without making things up.
    • The Result: The AI is actually really good at this! It acts like a tireless librarian who never misses a detail. In fact, the AI scored higher than average humans on these tasks because humans get tired and miss small things, while the AI scans everything perfectly.
  • Track 2: The "Expert Detective" Test (Domain Q&A)

    • The Goal: Can the AI understand complex business problems and solve them?
    • The Analogy: Imagine you hand the AI a stack of 50-page financial reports, a spreadsheet with sales data, and a slide deck about a new product. Then you ask: "If we buy this company, how much money will we make in three years, and what are the risks?"
    • What it checks: This isn't just about finding a number. The AI has to read everything, connect the dots between different files, do math, and build a logical argument. It's like asking a detective to solve a mystery using clues scattered across three different notebooks.
    • The Result: This is where the AI struggles. Even the smartest AI models only got about 59% to 63% correct. They are like a smart intern who knows the facts but sometimes gets the big picture wrong or makes calculation errors.

2. How They Graded the Exam

The researchers didn't just ask, "Was the answer right?" They broke every answer down into tiny, atomic pieces (like checking off a checklist).

  • The "Three-Judge" Panel: Instead of one person grading the test, they used three different AI models to act as judges. If two out of three judges agreed the answer was correct, it counted as a pass. This is like having a panel of three teachers grade a student's essay to make sure the grade is fair.
  • The "Atomic" Claims: If a question had 50 parts to the answer, the AI had to get all 50 parts right. If it got 49 right but missed one tiny detail, it lost points. This ensures the AI isn't just "vaguely" correct; it has to be precise.

3. The Big Findings

The paper reveals a split personality in today's AI:

  • The Super-Reader: When it comes to just looking at a document and finding facts (like "What is the date on this letter?"), the AI is a super-human. It is faster and more accurate than a human who reads the document once.
  • The Sub-Expert: When it comes to thinking deeply about complex business problems (like "Analyze this supply chain risk"), the AI is still a sub-expert. It is smart, but it's not quite at the level of a seasoned professional yet.

4. Does "Thinking Harder" Help?

The researchers tested if making the AI "think longer" or "think deeper" (using more computer power) would fix the problems.

  • The Result: Surprisingly, not really. Making the AI think harder within the same "family" of models didn't improve the score much. It was like asking a student to stare at a math problem for an hour instead of 10 minutes; they still made the same mistakes.
  • The Real Fix: The only way to get a better score was to switch to a "bigger, more expensive" model (a higher tier). It's like upgrading from a standard calculator to a super-computer; the bigger machine just does the job better, but it costs more and takes longer.

Summary

The paper introduces a new way to test if AI can truly understand our office files. It shows that while AI is amazing at finding information in documents, it is still learning how to reason through complex, real-world business problems. The test is now public, so other researchers can use it to see if their AI is getting better at being a true office assistant.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →