← Latest papers
🤖 AI

Measuring AI Ability to Complete Long Software Tasks

This paper introduces the "50%-task-completion time horizon" metric to quantify AI capabilities by comparing them to human effort, finding that frontier models can now complete tasks in 50 minutes that humans typically take 50 minutes to finish, with a doubling trend suggesting AI could automate month-long software tasks within five years.

Original authors: Thomas Kwa, Ben West, Joel Becker, Amy Deng, Katharyn Garcia, Max Hasin, Sami Jawhar, Megan Kinniment, Nate Rush, Sydney Von Arx, Ryan Bloom, Thomas Broadley, Haoxing Du, Brian Goodrich, Nikola Jurkov
Published 2026-07-14
📖 5 min read🧠 Deep dive

Original authors: Thomas Kwa, Ben West, Joel Becker, Amy Deng, Katharyn Garcia, Max Hasin, Sami Jawhar, Megan Kinniment, Nate Rush, Sydney Von Arx, Ryan Bloom, Thomas Broadley, Haoxing Du, Brian Goodrich, Nikola Jurkovic, Luke Harold Miles, Seraphina Nix, Tao Lin, Chris Painter, Neev Parikh, David Rein, Lucas Jun Koba Sato, Hjalmar Wijk, Daniel M. Ziegler, Elizabeth Barnes, Lawrence Chan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you're watching a race between a human and a robot, but instead of running on a track, they are both trying to solve a giant, ever-growing pile of software puzzles. For years, we've been trying to measure how good these AI robots are using tricky tests, but those tests often feel like artificial video game levels that don't really tell us how fast the robot can actually work in the real world.

To fix this, a team of researchers from METR decided to measure something much more intuitive: How long can an AI work before it starts to fail?

They called this the "Time Horizon." Think of it like a battery life for intelligence. If an AI can solve a puzzle that takes a human 10 minutes to finish, its time horizon is 10 minutes. If it can handle a task that takes a human 4 hours, its horizon is 4 hours. The researchers wanted to find the specific point where an AI has a 50% chance of getting the job done right.

The Big Discovery: The Exponential Explosion

The team gathered a massive collection of 170 different tasks, ranging from tiny, 1-second checks (like spotting a specific file name) to massive, 8-hour research projects. They hired skilled human experts to time how long it took them to finish these tasks. Then, they let 12 different AI models, from the old GPT-2 (released in 2019) to the brand-new o3 (released in 2025), try to solve them.

Here is the wild part: The AI's "battery life" is doubling every seven months.

In 2019, the best AI could only reliably handle tasks that took a human about 2 seconds to finish. By 2025, the top models (like o3) could handle tasks that take a human about 110 minutes (nearly two hours) with a 50% success rate.

The researchers found that this growth isn't just a little bit faster; it's an exponential curve. It's like watching a snowball roll down a hill, getting twice as big every time it passes a certain tree. They measured this trend from 2019 to 2025 and found the "doubling time" to be approximately 207 days (roughly seven months).

Why Are They Getting Better?

The paper suggests that the AI isn't just getting "smarter" in a vague way; it's getting better at specific, practical skills. When the researchers looked at why the older models failed and the newer ones succeeded, they found the newer models were:

  • Better at using tools: They know how to use a calculator or a code editor without breaking them.
  • Better at fixing mistakes: If an old AI made a mistake, it would often just repeat the same error over and over. The new models can spot the error, say "Oops," and try a different path.
  • Better at logical reasoning: They can follow a chain of thought without getting lost.

The "Messy" Caveat: Real Life is Harder

However, the paper is very careful not to say the robots have won the war yet. The researchers explicitly argue that these tests might be too "clean."

Imagine the AI is playing a video game where the rules are written in a clear manual, the enemies are predictable, and there are no sudden power outages. The researchers call real-world work "messy." In the real world, tasks often have unclear instructions, require talking to other people, or involve resources that run out.

When the team tested the AI on tasks that were "messier" (less structured, more chaotic), the AI's performance dropped significantly. The paper suggests that while the AI is getting great at the "clean" puzzles, it might still struggle with the chaotic, unpredictable nature of a real software engineer's day. They found no evidence that the AI's improvement slows down on messy tasks yet, but they warn that the gap between "clean test" and "messy reality" is a big unknown.

What Does This Mean for the Future?

The researchers did some math to see where this curve leads. If the trend of doubling every seven months continues, they predict that within 5 years (sometime between mid-2028 and mid-2031), AI systems might be able to autonomously complete software tasks that currently take a human one month (about 167 work hours) to finish.

They emphasize that this is a projection, not a guarantee. It depends on whether the AI keeps improving at this exact speed and whether those "clean" test results apply to the "messy" real world. If the trend holds, we could see AI agents that can essentially work as full-time employees on complex projects in the near future. But if the trend slows down or the "messiness" of real life proves too difficult, that timeline could change.

The Bottom Line

The paper doesn't claim AI has solved everything. Instead, it gives us a new ruler to measure progress. It shows that for the last six years, AI's ability to handle long, complex tasks has been growing at a terrifyingly fast pace, doubling its "work day" roughly every seven months. While there are still big gaps in how well they handle messy, real-world chaos, the trajectory suggests that the day an AI can work a full month-long shift on its own might be closer than we think.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →