← Latest papers
📈 economics

Classification testing: A new framework for drawing qualitative conclusions from quantitative estimates

This paper introduces "classification testing," a new statistical framework that allows researchers to assign quantitative estimates to substantively relevant qualitative classes with controlled error rates or declare results inconclusive, arguing that this approach improves upon standard hypothesis testing by better adjudicating rival possibilities and exposing hypotheses to refutation.

Original authors: Andrew C. Eggers, Zikai Li

Published 2026-08-25
📖 5 min read🧠 Deep dive

Original authors: Andrew C. Eggers, Zikai Li

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the world of social science, researchers spend their time measuring things that cannot be seen directly: how a new policy changes a community, whether a specific message alters a person's vote, or if a treatment improves a patient's health. They collect data and crunch numbers to find the average effect of these actions. However, the final story they tell is rarely a long string of decimals. Instead, they translate these precise numbers into simple, qualitative judgments. They decide if an effect is real or non-existent, positive or negative, or so small it does not matter. To support these judgments, scientists rely on a standard method called hypothesis testing. This method acts like a gatekeeper, designed to protect the public from false alarms by ensuring that when a researcher claims a result is real, the odds of that claim being a mistake are kept very low.

For decades, this system has worked well for confirming a specific prediction, but it has a significant blind spot. It is built to prove that something is happening, but it struggles to prove that nothing is happening, or to decide between two competing possibilities with equal fairness. If a researcher sets out to prove a treatment works, the system allows them to declare victory if the evidence is strong enough. But if the evidence is weak, the system simply says the researcher failed to prove their point; it does not allow them to confidently declare that the treatment actually does nothing. This creates an uneven playing field where some conclusions are easy to reach while others remain out of reach, potentially skewing the entire scientific conversation toward only the most optimistic findings.

A new framework proposed by political scientists Andrew Eggers and Zikai Li offers a different way to navigate these uncertainties. They call it "classification testing." Instead of forcing a researcher to pick a single favorite theory to defend, this approach asks them to divide all possible outcomes into distinct categories based on what actually matters for the real world. For example, a researcher might divide the possibilities into three groups: the effect is positive and large enough to care about, the effect is negative and large enough to care about, or the effect is so small it is negligible. The goal is not just to see if a number is different from zero, but to sort the result into the right bucket.

The power of this new method lies in its ability to treat all these categories with the same level of statistical rigor. In the old system, a researcher could only confidently say "the effect is positive" if they rejected the idea that it was zero or negative. They could never confidently say "the effect is zero" or "the effect is negative" with the same level of certainty unless they had set up a very specific, often awkward, test beforehand. Classification testing removes this bias. It allows a researcher to look at the data and say, with controlled error rates, that the result falls into the "negligible" category, or the "positive and substantial" category, or that the data is simply too muddy to decide. It turns the scientific question from "Can I prove my hypothesis?" into "What is the most honest description of this result?"

To see how this works in practice, the authors applied their method to a well-known field experiment involving media consumption. The original study examined what happened when regular viewers of Fox News were exposed to CNN. The researchers had measured thirty-two different outcomes, ranging from changes in political preferences to shifts in trust in news sources. When they applied the standard testing method used in most social science journals, they found that sixteen of these outcomes were "significant" and sixteen were not. This left half the results in a state of uncertainty, where the researchers could not confidently say the effect was real or non-existent.

When the authors re-analyzed the same data using classification testing, the picture became much clearer. By sorting the results into categories of "positive and substantial," "negative and substantial," or "negligible," they were able to make confident, error-controlled statements about twenty-nine of the thirty-two outcomes. They could confidently say that for some measures, the effect was large and positive. For others, they could confidently say the effect was so small it was negligible. Even for the results that the old method labeled as "not significant," the new method often allowed them to conclude that the effect was truly negligible, rather than just unknown. The only time they had to say "inconclusive" was when the data was genuinely too noisy to support any firm claim.

This shift changes the incentives for how research is conducted. Under the old rules, researchers were often pressured to find a positive result to get published, leading them to design studies that were likely to confirm their expectations. The new framework encourages a more open approach. Because the method allows for confident conclusions that an effect is small or non-existent, researchers can study genuinely open questions without fear that a "null" result is a failure. It treats the possibility of a small effect as a valid, testable outcome rather than a dead end.

The authors argue that this approach does not make science less strict; in fact, it maintains the same high standards for avoiding false claims. It simply expands the range of questions that can be answered with confidence. By allowing scientists to declare that an effect is negligible with the same certainty they declare it is large, the method reduces the pressure to force data into a specific narrative. It suggests that the most valuable scientific conclusion is not always a dramatic discovery, but sometimes a clear, honest assessment that a particular intervention does not make a meaningful difference. This clarity, the authors suggest, is the key to building a more reliable and less biased understanding of how the social world works.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →