When the adjustment fails: population-denominator sensitivity in coverage-adjusted PISA trends
This paper demonstrates that PISA's coverage-adjusted top-quarter trends are highly sensitive to the choice of population denominator, revealing significant discrepancies between reported data and World Population Prospects estimates that underscore the need for transparent provenance reporting and caution when synthetic weights trigger adjustments.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Every few years, a massive global experiment takes place in classrooms around the world. Hundreds of thousands of fifteen-year-olds sit down to take the same tests in math, reading, and science. This is the Programme for International Student Assessment, known as PISA. The results are more than just a report card for schools; they are a snapshot of how well a nation's young people are prepared for the future. Because the tests happen in cycles, researchers can look back and see if a country is getting better or worse over time. But there is a hidden complication in this comparison. The students who take the test are not always the same slice of the population from one year to the next. In some countries, the test might reach almost every fifteen-year-old, while in others, it might only reach a small fraction. When the group of students changes, the average score can shift simply because the mix of people changed, not because the students themselves learned more or less. To fix this, the organization that runs the tests uses a special mathematical adjustment. It tries to estimate what the top-performing students would have scored if the test had reached every single fifteen-year-old in the country, not just the ones who showed up. This adjusted number is supposed to be a fair way to compare progress across different years and different nations.
Two researchers, Jonas A. Mandalunes and Danica Jane S. Mandalunes, decided to see how fragile this adjustment really is. They focused on the most critical piece of the puzzle: the total number of fifteen-year-olds in each country. This number acts as the denominator, or the bottom part of the fraction, that determines how big the adjustment needs to be. The researchers asked a simple but profound question: what happens to the final results if you change that bottom number? They did not assume one source of population data was perfect and the other wrong. Instead, they treated two different ways of counting fifteen-year-olds as two different scenarios. One scenario used the numbers that national governments officially submitted to the test organizers. The other scenario used a widely respected global projection of population growth from the United Nations. By running the entire trend analysis twice—once with the official numbers and once with the global projections—they could see how much the story of a country's progress changed depending on which count was used.
The researchers looked at data from seventy-three economies between 2018 and 2022. They found that the official count of fifteen-year-olds used by the test organizers changed significantly in many places. On average, the difference in the coverage rate between the two years was nearly two percentage points, but in some cases, the change was massive, swinging by as much as twenty-seven points in one direction or forty-five in the other. When they broke down why these numbers changed, they discovered that in most cases, the shift came from the survey itself—meaning the test reached a different share of the enrolled students—rather than from a change in how many students were actually enrolled in school. This distinction is vital because it means a rising score might not mean more kids are going to school; it might just mean the test reached a different group of them.
The most striking finding was how sensitive the final trends were to this denominator. When the researchers swapped the official population count for the United Nations projection, the estimated change in test scores shifted dramatically. In mathematics, the average difference between the two scenarios was about 3.4 points. In reading, it was 3.5 points, and in science, 3.4 points. To put this in perspective, the test organizers publish a standard "link error" value to help readers understand the margin of uncertainty in their trends. For mathematics, that value is 2.24 points. The difference caused simply by changing the population count was larger than this standard margin of error in thirty-eight out of sixty-nine countries. In reading, it exceeded the margin in fifty-two countries, and in science, in fifty-one countries. This means that for a majority of the countries studied, the choice of which population number to use mattered more than the statistical uncertainty usually associated with the test itself.
The researchers also examined what happens when the math hits a wall. The adjustment procedure assumes that the test reached fewer students than actually exist in the country. But in a few specific cases, the math suggested the test reached more students than the official population count. When this happened, the procedure had to make a special fix: it simply set the extra weight to zero and reported the raw score of the students who took the test, effectively turning off the adjustment. This occurred in countries like Korea and Ireland in 2022. The researchers showed that this switch from an adjusted score to a raw score creates a sudden break in the data. It is not a smooth transition; it is a hard stop where the method changes entirely. They also noted that in some cases, the official population count was higher than the number of students the test claimed to have reached, which should be impossible if the definitions were perfectly aligned, suggesting that the definitions of who counts as a "fifteen-year-old" or who is "enrolled" might differ between the test organizers and the population statisticians.
The study did not conclude that one population count was the true one and the other a mistake. The researchers were careful to state that neither the official government numbers nor the United Nations projections are the absolute truth. The official numbers often lack details about exactly when the count was taken or what specific definition of "resident" was used. The global projections are models based on assumptions. The core problem is that the final trend score depends heavily on a number that is not fully explained or consistent. The researchers found that even when they tried to clean up the data by removing countries with known territorial disputes or inconsistent records, the differences remained. The gap between the two scenarios did not disappear; it just shifted slightly. This suggests that the sensitivity is a fundamental feature of the current method, not just a result of messy data.
In the end, the researchers argue that the way these trends are reported needs to change to be more honest about this uncertainty. They recommend that the test organizers publish exactly where the population numbers come from, including the date and the specific definition used. They suggest that every time a trend is reported, it should be accompanied by a range showing what the result would look like if a different, reasonable population count were used. This would not be a confidence interval in the statistical sense, but a clear statement of how much the result depends on the denominator. Finally, they propose that if the math breaks down because the numbers don't add up, the adjusted score should not be silently fixed or hidden; it should be flagged and withheld until the inputs are reconciled. The study shows that the story of how countries are improving or declining is real, but the specific numbers we use to tell that story are far more delicate than they appear. The trends exist, but their precise magnitude is tied to a choice of numbers that is currently made in the dark.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.