Finite Evidence, Infinite Horizons: The Artificial Age Score as a Case Study in Longitudinal AI Measurement
This paper argues that finite longitudinal AI evaluation records, exemplified by the Artificial Age Score (AAS), cannot determine infinite-horizon system properties or convergence without explicit stipulated continuation models, as demonstrated through counterexamples and stress tests showing that bounded observations alone yield only conditional, model-based inferences rather than definitive conclusions.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the rapidly evolving world of artificial intelligence, researchers are constantly trying to measure how well these systems learn, remember, and perform over time. Just as a teacher might track a student's grades across a semester to predict future success, scientists use longitudinal scores to monitor AI systems as they face new questions, updates, and changing environments. These scores are meant to be more than just a snapshot of a single moment; they are intended to tell a story about the system's journey, its stability, and its ability to keep working effectively as time passes. However, a fundamental challenge lurks in this approach: we only ever have a finite record of what has happened so far. We have a history of past tests, but we do not have the future. The critical question is whether a pattern observed in a limited number of past tests can truly tell us anything definitive about how the system will behave indefinitely into the future, or if we are simply reading too much into a short story.
A new study by Seyma Yaman Kayadibi at Victoria University tackles this problem directly, using a specific measurement tool called the Artificial Age Score as a case study to expose a hidden gap in how we interpret AI performance. The research does not argue that these scores are useless; rather, it demonstrates that a score calculated from a finite history of observations cannot, on its own, prove that a system will continue to exist, remain stable, or perform well forever. The study shows that a system could appear perfectly stable for twenty sessions and then fail immediately after, or it could continue perfectly, and the mathematical score calculated from those first twenty sessions would look identical in both scenarios. The paper argues that any claim about the infinite future of an AI system is not a consequence of the data we have collected, but rather a result of extra assumptions we make about how the system will continue.
To illustrate this, the researchers examined a real-world dataset involving an AI system tested over seventy-four weekly releases. They compared two versions of the system: one that operated in a closed environment and another that could search the internet for answers. The data showed that the version with internet access consistently performed better, resulting in a lower "burden" score, which indicates fewer errors. The researchers broke down exactly why this difference occurred, finding that the improvement came from better performance in both multiple-choice questions and open-ended text generation. However, when they tried to predict the future performance of these systems using simple rules—such as assuming the next score would be the same as the last one, or that it would be the average of the last four—they found that every single prediction rule made mistakes. The rule that worked best for one version of the system did not work best for the other, proving that there is no single, universal law hidden inside the scores that dictates the future.
The study also explored a deeper mathematical reality: for any finite sequence of scores, it is possible to imagine an infinite number of different futures that fit that sequence perfectly. One future might show the system improving forever, another might show it failing immediately, and a third might show it oscillating between good and bad performance. All of these different futures would produce the exact same score for the past twenty sessions. The researchers proved that without adding extra rules about how the system behaves over time, the score itself cannot distinguish between these possibilities. They also looked at how the passage of real time affects these measurements. If you count every test as one unit of time, you might get a different picture of the system's total workload than if you measure by the actual hours and days between tests. The paper shows that these different ways of measuring time can change the interpretation of how much "exposure" or stress the system has endured, yet neither method can predict what happens next.
Ultimately, the paper concludes that we must be much more careful about how we talk about the long-term reliability of artificial intelligence. A score that looks stable today is not a guarantee of stability tomorrow. To make claims about the future, researchers must explicitly state the assumptions they are making about the system's lineage, its environment, and its rules of operation. The study proposes a new way of reporting results that clearly separates what was actually observed, what was assumed to make a prediction, and what remains unknown. By treating the future as a conditional possibility rather than a mathematical certainty derived from past data, the research offers a more honest and rigorous framework for understanding the limits of what we can know about the artificial age. The findings serve as a reminder that while we can measure the past with precision, the future remains a separate domain that requires its own set of rules and assumptions to navigate.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.