← Latest papers
🤖 AI

Measuring Agents in Production

Original authors: Melissa Z. Pan, Negar Arabzadeh, Riccardo Cogo, Yuxuan Zhu, Alexander Xiong, Lakshya A Agrawal, Huanzhi Mao, Emma Shen, Sid Pallerla, Liana Patel, Shu Liu, Tianneng Shi, Xiaoyuan Liu, Jared Quincy Dav
Published 2026-06-08
📖 5 min read🧠 Deep dive

Original authors: Melissa Z. Pan, Negar Arabzadeh, Riccardo Cogo, Yuxuan Zhu, Alexander Xiong, Lakshya A Agrawal, Huanzhi Mao, Emma Shen, Sid Pallerla, Liana Patel, Shu Liu, Tianneng Shi, Xiaoyuan Liu, Jared Quincy Davis, Emmanuele Lacavalla, Alessandro Basile, Shuyi Yang, Paul Castro, Daniel Kang, Koushik Sen, Dawn Song, Joseph E. Gonzalez, Ion Stoica, Matei Zaharia, Marquita Ellis

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you've just built a fleet of incredibly smart, robotic assistants (AI Agents) that can do complex tasks like writing code, managing finances, or diagnosing medical issues. The tech world is buzzing with excitement, imagining these robots running the world autonomously.

But here's the twist: The robots actually running in the real world right now aren't the super-autonomous, "do-it-all" robots from the sci-fi movies. They are more like cautious, highly supervised interns.

This paper, titled "Measuring Agents in Production" (MAP), is the first major report card on what these AI assistants are actually doing in real companies. The researchers didn't just guess; they interviewed 20 teams building these systems and surveyed 86 real-world deployments.

Here is what they found, explained simply:

1. Why are companies building these robots?

The Goal: To save time and get more work done.
The Analogy: Imagine a factory. The boss isn't trying to build a robot that can invent a new car engine from scratch. They want a robot that can assemble the parts 10 times faster than a human.

  • 80% of the teams said they built these agents to boost productivity.
  • They aren't trying to replace humans entirely yet; they are mostly helping human employees (like doctors, engineers, or customer service reps) do their jobs faster.
  • Surprise: They aren't building them to "mitigate risk" or "reduce expertise" as much as you might think. The main driver is just getting things done faster.

2. How are they actually built? (The "Secret Sauce")

If you look at research labs, they try to build the most complex, autonomous robots possible. But in the real world, companies are doing the exact opposite. They are choosing simplicity and control.

  • The "Off-the-Shelf" Rule: 70% of these agents don't have their own custom brains. They just use powerful, pre-made AI models (like the ones you might have heard of) and give them very specific instructions. They aren't retraining the brain; they are just writing better manuals for it.
  • The "Short Shift" Rule: In research, agents might try to solve a problem by taking 100 steps on their own. In production, 68% of agents are stopped after 10 steps and asked to wait for a human to check their work.
    • Analogy: Think of it like a GPS. A research robot might try to drive the whole way to the destination without a driver. A production robot drives for 10 minutes, then stops and asks, "Is this the right turn?" before continuing.
  • The "Human in the Loop" Rule: 74% of the time, a human is the final judge. The robot suggests an answer, but a human has to say "Yes, that's correct" before it gets sent out.

3. How do they know if the robot is doing a good job?

This is one of the biggest headaches.

  • No Standard Tests: There is no "SAT test" for AI agents yet. 75% of teams don't use formal benchmarks because they are too hard to make or don't exist for their specific job.
  • The "Human Judge" Method: Instead of a computer grading the robot, humans are the graders. Experts look at the robot's work and say, "Good job" or "Try again."
  • The "A/B Test" Method: Sometimes they just run the robot alongside a human or an old software system to see which one finishes the task faster.

4. What is the biggest problem?

Reliability.

  • The Problem: The biggest fear isn't that the robot is "dumb"; it's that it might make a mistake and do something weird or dangerous.
  • The Solution: Companies aren't trying to fix this by making the robot "smarter" (which is hard). They are fixing it by putting up guardrails.
    • They limit how many steps the robot can take.
    • They force the robot to ask for permission before doing anything risky.
    • They run the robot in a "sandbox" (a safe, fake environment) before letting it touch real data.

5. How fast do they need to be?

  • The Analogy: You don't need a Formula 1 car to mow your lawn.
  • The Reality: Most of these agents are surprisingly slow. 66% of them can take minutes (or even hours) to finish a task.
  • Why? Because they are usually doing background work (like processing insurance claims or organizing files). As long as they finish the job faster than a human would (even if it takes 5 minutes), they are considered a success. They don't need to be instant like a voice assistant.

The Big Takeaway

The paper concludes that successful AI in the real world isn't about building the most autonomous, complex AI. It's about building simple, controllable tools that work well with humans.

Companies are deliberately choosing to be "boring" and safe. They trade the "cool factor" of a robot that can do anything on its own for the reliability of a robot that does a small part of the job perfectly and asks a human for help when it gets stuck.

In short: The AI agents of today are less like "Jarvis" (the fully autonomous AI from Iron Man) and more like a very helpful, slightly nervous assistant who double-checks everything with their boss before sending an email. And that's exactly what makes them successful in the real world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →