Substituting National-IQ and Harmonized-Learning Indicators: An Output-Specific Replication and Sensitivity Audit
This paper conducts a sensitivity audit comparing national-IQ and harmonized-learning indicators, finding high correlations and similar rank distributions but concluding that the datasets are not interchangeable or causally equivalent despite their statistical similarities.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the study of nations, researchers often try to measure the collective mental capacity of a country's people. One approach involves compiling scores from intelligence tests given to small groups of individuals and averaging them to create a national estimate. Another approach looks at the results of standardized exams taken by millions of schoolchildren in math, reading, and science. For years, these two different sets of numbers have been compared. When they move together—when countries with high test scores also have high estimated intelligence scores—some have assumed the two measures are essentially the same thing, interchangeable tools that can be swapped in and out of research without changing the story. This assumption is tempting because it simplifies the complex task of comparing nations. However, a high correlation between two lists of numbers does not automatically mean they tell the same story about every single country, nor does it guarantee they will lead to the same conclusions when used to predict economic success.
A recent audit conducted by independent researcher Georgios Kiminos challenges this assumption by treating the two measures not as identical twins, but as distinct instruments that need to be tested against specific questions. The study asks a practical, if unglamorous, question: if a researcher replaces a national intelligence score with a harmonized learning score, does the ranking of countries change? Does the list of top-performing nations stay the same? Does the link between these scores and a country's wealth hold up? To answer this, the researcher gathered data for exactly sixty-six economies where both types of scores were available. He then ran a series of checks, comparing how the countries lined up against each other, which nations appeared at the very top and bottom of the lists, and how strongly each list predicted the country's gross domestic product per capita in 2017.
The results show that while the two measures are closely related, they are not interchangeable. When the researcher compared the rankings of the sixty-six countries, the average country moved nearly ten spots up or down the list when switching from the intelligence score to the learning score. While the overall pattern remained similar, the specific order changed enough to matter. For instance, when looking at the top group of nations, the two lists agreed on only four out of the top seven countries. At the bottom of the list, they agreed on only two out of the seven lowest-ranked nations. This means that a country considered a leader in one system might be seen as merely average in the other, and a nation struggling in one might appear less dire in the other. The two measures are like two different maps of the same terrain: they show the same general shape of the landscape, but if you are trying to navigate to a specific destination, the route they suggest can differ significantly.
The study also examined how well these scores predicted a country's economic wealth. In this specific group of sixty-five countries, the learning scores showed a slightly stronger connection to wealth than the intelligence scores did. However, the difference was small enough that statistical tests could not confirm it was a real, reliable advantage rather than a fluke of the specific sample used. The researcher found that the result was sensitive to how the data was constructed. When the sample was changed to include different countries, or when the data was weighted to reflect population size rather than treating every country as equal, the size of the difference between the two measures shifted. In some variations, the learning scores looked much better at predicting wealth; in others, the gap narrowed or even reversed. Crucially, the study found that the direction of the difference remained consistent in the primary analysis, but the magnitude of that difference depended entirely on which countries were included and how the data was weighed.
This work serves as a cautionary tale for anyone using national data to make broad comparisons. The audit demonstrates that high correlation is not a magic wand that makes two different measures the same. Just because two lists of numbers move together does not mean they agree on the details, nor does it mean they will lead to the same policy decisions or research conclusions. The researcher did not declare one measure superior to the other, nor did he prove that one is invalid. Instead, the study highlights that the choice of which number to use matters. It shows that swapping one indicator for another can alter the ranking of nations, change which countries are identified as outliers, and shift the strength of their relationship with economic outcomes. The findings suggest that before replacing one measure with another, researchers must define exactly what they are trying to preserve—the specific ranking, the classification of top performers, or the link to an outcome—and test whether the new measure actually holds that ground. Without such a check, the substitution remains an unverified guess, and the conclusions drawn from it may be more fragile than they appear.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.