Quantifying the Expectation-Realisation Gap for Agentic AI Systems
This paper quantifies the significant gap between pre-deployment expectations and post-deployment realities of Agentic AI systems across software engineering and clinical domains, revealing that workflow friction and verification burdens often negate anticipated productivity gains and necessitate structured planning frameworks with realistic benefit projections.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are buying a new, high-tech self-driving car. The salesperson promises it will cut your commute time in half, drive perfectly, and let you nap the whole way. You buy it, excited for the future. But when you actually start driving it, you find yourself constantly taking the wheel, arguing with the GPS, and fixing the car's mistakes. Instead of saving time, your commute takes longer than before.
This is exactly what is happening with Agentic AI (smart computer programs that try to do complex jobs for us). A new paper by Sebastian Lobentanzer investigates why the "hype" about AI doesn't match the "reality" on the ground.
Here is the breakdown of the paper in simple, everyday terms:
1. The Big Mismatch: The "Expectation vs. Reality" Gap
The paper calls this the Expectation–Realisation Gap.
- The Promise: Companies and vendors say, "This AI will save you 5 minutes per task!" or "It will make you 20% faster!"
- The Reality: When people actually use the tools, they often end up slower, or the savings are tiny (like 30 seconds), or the tool makes mistakes that take longer to fix than doing the job yourself.
2. Three Real-World Examples (The "Test Drives")
The paper looked at three different areas where AI is being used:
Software Coding (The "Junior Mechanic" Problem):
- Expectation: Experienced programmers thought AI would help them finish code 24% faster.
- Reality: They actually finished 19% slower.
- Why? The AI wrote code, but it was often buggy or confusing. The programmers had to spend extra time reading, debugging, and fixing the AI's work. It was like hiring a fast but clumsy assistant who writes the report in 5 minutes but then you spend 20 minutes correcting their spelling and logic.
- Note: The AI did help beginners with simple tasks, but it confused the experts with complex, real-world jobs.
Medical Notes (The "Time-Shifting" Illusion):
- Expectation: Doctors were told AI "scribes" would save them 5 minutes per patient visit by listening and typing notes automatically.
- Reality: The savings were often less than 1 minute, or non-existent.
- Why? The AI sometimes wrote the wrong medical details. Doctors had to stop and carefully check every sentence. Also, the time saved during the day often just got pushed to the evening, where doctors stayed late to fix the notes. It wasn't saved time; it was just moved time.
Medical Decisions (The "Confident but Wrong" Problem):
- Expectation: AI tools claimed to predict diseases (like sepsis) with 80%+ accuracy.
- Reality: When tested in real hospitals, the accuracy dropped to about 63%.
- Why? The AI was tested in a "lab" with perfect data. In the messy real world, it missed two-thirds of the sick patients. It was like a weather app that is perfect in a simulation but fails to predict rain in your actual city.
3. Why Does This Happen? (The Three Culprits)
The paper identifies three main reasons why the AI promises fail:
- The "Friction" Factor: AI doesn't work in a vacuum. It has to fit into your existing messy workflow. If the AI doesn't talk nicely to your other software, or if you have to click five extra buttons to use it, that friction eats up all the time savings.
- The "Trust but Verify" Tax: You can't just let the AI do the work and walk away. You have to check its work. The time spent checking the AI often cancels out the time the AI saved you.
- The "One Size Fits None" Trap: AI doesn't help everyone equally.
- Analogy: Think of a new GPS app. A new driver might love it and get lost less. But a veteran driver who already knows every shortcut might find the GPS annoying and slower.
- In the studies, AI helped less experienced workers the most, but often slowed down the experts who already had optimized their own systems.
4. The Solution: Stop Guessing, Start Measuring
The author argues that we need to stop making vague promises like "AI will make us faster." Instead, we need a Structured Plan (called the "Agentic Automation Canvas").
This plan asks:
- Exactly what are we trying to save? (Time? Money? Errors?)
- Who will actually benefit? (The new hires or the veterans?)
- How much time will we spend checking the AI's work? (Subtract this from the savings!)
- What happens if it fails? (Do we have a backup plan?)
The Bottom Line
AI is a powerful tool, but it's not magic. The paper warns that if we keep buying AI based on sales pitches and "lab results," we will keep getting disappointed. To make AI work, we need to be realistic, measure the actual results (including the time spent fixing mistakes), and understand that it helps some people more than others.
In short: Don't buy the "self-driving car" just because the brochure says it's perfect. Test drive it in your own neighborhood, count the time you spend taking the wheel, and only then decide if it's worth the price.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.