← Latest papers
📊 statistics

Feature Selection for Multidimensional Poverty Measurement: Taming Indicator Dimensionality

This paper demonstrates that machine learning-based feature selection can significantly reduce the computational and interpretive burden of multidimensional poverty measurement by identifying a small, non-redundant subset of indicators that maintains classification accuracy comparable to full indicator sets, as evidenced by an analysis of India's IHDS-II data.

Original authors: Kan Sun

Published 2026-09-23✓ Author reviewed ⓘ
📖 5 min read🧠 Deep dive

Original authors: Kan Sun

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Poverty has long been measured by a single number: how much money a person or family has. If their income falls below a certain line, they are counted as poor. But this narrow view misses the reality of hardship. A family might have enough to buy food but lack clean water, electricity, or access to a doctor. To capture this fuller picture, social scientists and policymakers have turned to multidimensional poverty measurement. Instead of looking at just one thing, they track dozens of different indicators—such as whether children go to school, if the home has a flush toilet, or if adults suffer from chronic illness. By combining these many signals, they can identify who is truly deprived in multiple ways at once. This approach is now the global standard for understanding poverty, used by organizations like the United Nations to guide aid and policy across more than a hundred countries. Yet, as these lists of indicators grow longer, often reaching into the hundreds, a new problem emerges. Collecting and processing so much data becomes incredibly expensive and slow, making it difficult for governments to monitor the poor effectively in real time.

A researcher at Fudan University in China, Kan Sun, set out to solve this practical bottleneck. The question was not whether we should measure many aspects of life, but whether we really need to measure every single one to get an accurate result. Sun asked if there was a way to trim the fat from these massive lists without losing the ability to correctly identify the poor. To find the answer, the researcher turned to tools usually reserved for artificial intelligence and machine learning. Specifically, the study used a set of techniques called "feature selection." In plain terms, these are methods that act like a sieve, sorting through a large pile of data to find the few items that carry the most important information and discarding the rest as redundant. The goal was to see if a small, carefully chosen group of indicators could do the same job as a much larger, unwieldy set.

The study tested these ideas using a massive survey of over 42,000 households in India, which tracked 15 different signs of deprivation across education, health, and living standards. The researcher applied five different machine-learning algorithms to this data. Each algorithm acted like a different judge, ranking the 15 indicators by how useful they were for predicting whether a household was poor. The results were striking. The algorithms unanimously agreed that a very small subset of indicators contained almost all the necessary information. At the very top of the list was female schooling. The data showed that knowing whether the most educated woman in a household had fewer than six years of education was enough to predict poverty with near-identical accuracy to the full set. If a woman in the house lacked this basic education, the household was almost certainly poor across multiple dimensions. If she had it, the household was almost certainly not poor.

Beyond female education, a few other factors like adult health and core living standards also proved highly informative. In contrast, the indicators that seemed important on paper turned out to be useless for sorting the poor from the non-poor. For instance, computer ownership was ranked at the very bottom. This was not because computers are unimportant, but because almost no one in the survey owned one; since nearly everyone was deprived of this item, it could not help distinguish between a poor family and a slightly less poor one. Similarly, electricity was so common that its absence was rare, making it a poor tool for sorting. The study found that by dropping the least useful indicators, the researchers could reduce the list from 15 items down to 8 without losing much accuracy; the system could still identify the poor with near-identical accuracy. However, accuracy does not fall sharply until fewer than 5 indicators remain. This finding held true even when the researchers changed the rules for what counted as "poor" or adjusted how much weight was given to different categories.

To ensure this was not just a quirk of the Indian data, the researcher repeated the exercise with a similar survey from South Africa. The results were nearly identical. In that country as well, education and health formed a small, powerful core that carried the weight of the entire measurement system, while a large number of living-standard indicators turned out to be redundant. The study confirmed that this pattern is not a fluke of one dataset but a structural feature of how poverty manifests in these regions.

It is crucial to understand what this study does and does not say. The research does not argue that we should stop measuring things like computer ownership or wall quality because they are unimportant for human well-being. The choice of which dimensions to include in a poverty measure is a moral and political decision, based on what a society values. This study does not make those choices. Instead, it acts as a practical tool for the people who have already made those choices. It shows them that once they have decided to measure 15 things, they might not need to collect data on all 15 to get a reliable count of the poor. By identifying which indicators are statistically redundant, the study offers a way to make poverty monitoring systems cheaper, faster, and easier to manage without sacrificing accuracy. The final decision on what to measure remains with policymakers, but they can now do so with a clear map of which indicators are essential and which are merely extra.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →