Frontier AI Forecasting Has a Measurement Problem: An Audit of Progress Evidence
This paper audits the measurement foundations of frontier AI forecasting and finds that critical data gaps, benchmark inconsistencies, and source concentration undermine simple trend-based predictions, arguing that defensible forecasts must instead be grounded in explicit, versioned measurement systems with transparent protocols and links.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
We live in an era where the capabilities of artificial intelligence are often predicted with the precision of a weather forecast. Experts look at how fast computers are getting, how much data they are fed, and how well they perform on standardized tests to guess when machines will reach human-level reasoning or even surpass it. These predictions usually take the form of a specific date on a calendar, suggesting that by a certain year, an AI will be able to solve a complex problem or perform a specific task. The logic seems straightforward: if you can measure how fast a car is going, you can calculate when it will arrive at its destination. But this approach relies on a hidden assumption—that the ruler used to measure the speed is the same today as it was yesterday, and that the speedometer is working correctly for every single car on the road.
A new analysis by Fabricio F. Costa challenges this assumption. The paper does not argue that predicting the future of AI is impossible, but rather that the current way we try to do it is built on a shaky foundation. The author conducted a rigorous audit of the public records used to make these forecasts, treating the data not as a smooth line on a graph, but as a collection of individual events, measurements, and sources. The investigation reveals that the connection between the resources used to build these systems and their actual performance is far more fragmented than popular forecasts suggest. The study finds that the evidence supporting these predictions is often incomplete, inconsistent, and heavily reliant on a single source of information, making the specific dates often cited in the news more of a guess than a scientific conclusion.
To understand the scope of the problem, one must first look at how these forecasts are typically constructed. Researchers usually try to link two things: the amount of computing power used to train a model and the results it achieves on a benchmark test. A benchmark is simply a standardized set of questions or tasks designed to measure a specific skill, like answering trivia or solving math problems. The idea is that as you give a model more computing power, its score on these tests goes up in a predictable way. If you can map that relationship, you can project forward to see when the score will hit a certain threshold. However, for this projection to be valid, you need to have data points where both the computing power and the test score are known for the same system.
When the author examined the public record of sixty-two major artificial intelligence systems, a significant gap appeared. While information about the computing power used was available for many systems, it was missing for a large number of the most advanced, closed-source models. Conversely, the performance scores from a specific, widely cited measurement program called METR were available for many closed models but were completely absent for the open-source systems. This created a situation where the two pieces of data needed to draw a line—the "fuel" and the "speed"—rarely existed together for the same machine. In the entire set of sixty-two systems, only seven had both the computing estimate and the performance score recorded. This means that the smooth curves often seen in forecasts are actually drawn through just a handful of data points, while the rest of the landscape is a blank space.
The problem deepens when you consider that the tests themselves are not static. Just as a ruler might be replaced by a new one with slightly different markings, the benchmarks used to measure AI are constantly being updated, revised, or replaced to prevent the models from simply memorizing the answers. The audit looked at how these different versions of tests relate to one another. It found that when you try to compare scores from an older version of a test to a newer version, the relationship is not always a simple, straight shift. Sometimes, the new test measures something slightly different, or the difficulty changes in a way that makes the scores jump or drop unexpectedly. The study tested this by looking at how systems performed on two versions of a popular test and found that the connection between the old scores and the new ones depended heavily on the mathematical method used to compare them. In some cases, the new version seemed to show a much steeper improvement than the old one, while in others, the improvement looked flat. This suggests that a rise in a score might not always mean the AI has gotten smarter; it might just mean the test has changed.
Another critical issue identified is the concentration of the evidence. The vast majority of the specific data points used to track AI progress come from a single measurement program. Out of seventy-one major quantitative events recorded in the audit, fifty-two came from just one source. This is like trying to understand the weather patterns of an entire continent by relying on the thermometer readings from a single city. While that city might be accurate, it does not tell you if the rest of the continent is experiencing the same conditions. The audit also noted that many of these measurements come from laboratory releases rather than real-world field tests. In a lab, the conditions are controlled, but in the real world, the AI interacts with unpredictable environments, different users, and varying levels of support. The paper argues that relying on a single source and a single type of test creates a false sense of certainty, where the confidence in a forecast is high simply because the data looks consistent, even though it is not truly independent.
The author also looked at the power of the statistical tests used to link these different measurements. The analysis showed that the current data sets are often too small to detect anything other than very large changes. If the relationship between computing power and performance changed by a small amount, the current records would likely miss it entirely. The study calculated that to reliably detect a small shift in how these systems improve, researchers would need many more systems to be tested on both the old and new benchmarks simultaneously. Without this larger, more diverse set of data, any forecast that claims to know exactly when a specific milestone will be reached is essentially guessing.
The paper does not suggest that we should stop trying to measure AI progress. Instead, it calls for a more honest and complex approach. Rather than trying to squeeze all of AI's capabilities into a single number or a single date, the author proposes a "portfolio" of measurements. This would involve tracking different things at the same time: how much energy and money it takes to run a model, how reliable it is over long periods, how it performs in real-world tasks, and how it handles safety risks. It would also require a system where different research groups test the same models using different methods to ensure that the results are not just an artifact of one specific lab or one specific test. The goal is to move away from the idea of a single, perfect ruler and toward a collection of tools that can measure different aspects of intelligence with their own uncertainties.
Ultimately, the study concludes that a defensible forecast about the future of artificial intelligence must be transparent about what it is measuring and how it is measuring it. It must admit where the data is missing, where the tests have changed, and where the evidence comes from. A prediction that simply gives a date without explaining the measurement system behind it is not a scientific claim; it is a guess. The paper urges the field to treat the measurement system itself as part of the forecast, acknowledging that the date is only as meaningful as the ruler used to find it. Until the community can build a record where computing power, performance, and real-world outcomes are consistently and independently measured across many different systems, the most precise part of any forecast will be the date on the calendar, while the least precise part will be what is actually being measured on the way there.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.