← Latest papers
📊 statistics

Measurement-bounded predictive inference in international large-scale assessment: how the plausible-value ceiling governs attainable accuracy and the stability of predictor rankings in PISA 2022

This study demonstrates that the measurement precision ceiling inherent in PISA 2022 plausible values fundamentally limits attainable predictive accuracy and destabilizes predictor rankings across domains and education systems, necessitating that performance comparisons and model interpretations be grounded in these computable measurement bounds.

Original authors: Jonas A. Mandalunes

Published 2026-08-28
📖 5 min read🧠 Deep dive

Original authors: Jonas A. Mandalunes

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Every year, millions of teenagers around the world take a massive test called PISA, designed to see how well they understand math, reading, and science. The results are used by governments and educators to judge which schools and countries are doing the best job. In recent years, researchers have started using these test scores to build computer models that try to predict a student's success based on their background, such as their family's income, the books in their home, or their school's climate. These studies often produce a ranked list of the most important factors, suggesting that if we want to improve scores, we should focus on the top items on that list. However, there is a hidden problem with this approach. The test scores used in these studies are not direct observations of a student's ability. Instead, they are statistical estimates, created by drawing ten different possible numbers for each student from a cloud of possibilities. Because these numbers are estimates, they contain a certain amount of built-in noise, or uncertainty, that no amount of data can ever explain away. This uncertainty sets a hard limit on how accurately any model can predict a student's performance, a limit that changes depending on the subject being tested and the country where the test was taken.

A new study by Jonas Mandalunes investigates exactly how much this hidden noise limits our ability to learn from these massive tests. The researcher used the international database from the 2022 PISA cycle, which includes data from over 600,000 students across 80 different education systems. The study focused on four areas: mathematics, reading, science, and a newer addition called creative thinking. The core question was simple: how much of a student's score can a computer model actually explain, and does this limit change the list of factors that seem to matter most? To answer this, the researcher calculated a "ceiling" for each subject in each country. This ceiling represents the maximum possible accuracy any prediction model could ever achieve, simply because the test scores themselves are not perfect. The study found that this ceiling varies wildly. For creative thinking, the limit ranged from about 65 percent to nearly 90 percent depending on the country. For mathematics, the range was tighter but still varied, from about 81 percent to 93 percent. This means that in some places, no model can ever explain more than two-thirds of the differences in student scores, while in others, the limit is much higher.

The study went further to see if this limit affects the rankings of the factors that influence scores. The researcher ran the same computer models over and over again, each time using a different one of the ten possible score estimates for the students. In the most precisely measured subjects, the list of top factors stayed mostly the same. But in the less precise subjects, like creative thinking, the list changed dramatically. When the researcher switched from one score estimate to another, more than 40 percent of the top five factors on the list would swap places. This means that a study reporting a single list of "most important" factors is essentially showing just one random snapshot of a much wider, shifting reality. If a researcher had picked a different score estimate, they might have concluded that a completely different set of factors was driving student success.

Another critical finding concerned how the data was split up to test the models. Many studies randomly mix students from different schools into training and testing groups. The new research showed that this method creates a false sense of accuracy. Because students in the same school tend to be similar, a model that sees students from a specific school during training can learn patterns specific to that school, rather than learning about the students themselves. When the researcher forced the model to keep entire schools separate during training and testing, the predicted accuracy dropped significantly. In some cases, the drop was massive, revealing that previous studies had inflated their results by as much as 0.36 points simply because they did not account for the school structure. The size of this inflation depended on how many schools were in the study, not on the complexity of the computer model used.

The study also looked at why the precision of the scores differed so much between subjects. It found that the more a test relied on human graders to score open-ended answers, the less precise the scores tended to be, especially in very large countries. Creative thinking, which was scored entirely by humans reading written responses, showed the widest variation in precision. In contrast, mathematics, which is mostly scored by machines, was more consistent. This suggests that the way a test is administered and scored is just as important as the questions themselves when it comes to how much we can learn from the results.

The implications of these findings are straightforward for anyone trying to understand education data. You cannot compare the performance of a model in one country to another, or in one subject to another, without first knowing the precision ceiling for that specific context. A model that explains 45 percent of the variance in a country with a low ceiling is actually performing quite well, capturing most of what is possible to know. The same model in a country with a high ceiling might be performing poorly, leaving much of the signal unexplained. Furthermore, any list of important factors derived from these tests should be treated with caution. Because the rankings shift so much depending on which score estimate is used, a single list is not a stable fact but a temporary realization. The study concludes that researchers should always report the precision ceiling alongside their results and should test their models across all possible score estimates to see if their conclusions hold up. Without these checks, the rankings and comparisons that guide education policy may be built on a foundation of statistical noise rather than solid evidence.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →