← Latest papers
📊 statistics

T_root by Default: A Reporting Standard for Independence in Contingency Tables

This paper argues that current statistical software defaults for contingency tables are fundamentally flawed due to poor calibration in sparse or heterogeneous data, and proposes replacing them with a new, single closed-form statistic called T_root that offers superior size control and power without requiring resampling methods.

Original authors: William J. Dwyer

Published 2026-09-03
📖 6 min read🧠 Deep dive

Original authors: William J. Dwyer

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the world of scientific research, data often comes in the form of simple grids, where researchers count how many people fall into different categories. Imagine a table comparing two groups, such as patients who received a new treatment versus those who received a standard one, and checking whether they recovered or did not. To decide if the treatment actually made a difference, scientists use a mathematical tool to calculate a number that tells them if the pattern they see is real or just a random fluke. For decades, the standard practice has been to rely on two specific tools for this job. One tool is famous for being "exact," meaning it calculates probabilities without making broad guesses, while the other is a classic method that works well when the data is plentiful and evenly spread out. However, a new analysis suggests that the way software currently chooses between these tools is flawed, leading to misleading conclusions in the vast majority of real-world studies. The problem is not that the tools are broken, but that the rules for using them are too rigid, causing researchers to either miss real effects or see patterns that do not exist.

A researcher at the University of Massachusetts Lowell, William J. Dwyer, has proposed a new standard to fix this reporting error. He argues that the current habit of automatically switching to the "exact" tool whenever a table has small numbers is a mistake. While that tool is precise in its calculations, it is overly cautious, often hiding real discoveries by making the evidence look weaker than it truly is. Conversely, the other common tool becomes too eager to find patterns when the data is uneven, sometimes shouting "discovery" when there is none. The author found that these two traditional methods are actually equally good at finding true effects when the data is balanced, but they both fail to report the correct level of certainty when the data is messy or sparse. Instead of forcing researchers to choose between a tool that is too shy and one that is too bold, the paper introduces a single, new statistic called T_root, which was developed and validated in a companion methods paper. This statistic works reliably across the entire spectrum of data types.

The study began by testing thousands of simulated tables to see how the old methods performed under different conditions. The results showed that the "exact" tool, despite its name, does not control error rates as promised. Because it can only produce specific, discrete steps of probability, it often ends up being much stricter than intended, effectively lowering the threshold for what counts as a discovery without telling anyone. In these simulations, the tool would frequently fail to flag a real connection, even though the data clearly showed one. At the same time, the alternative tool would often flag connections that were just random noise, especially when the groups being compared were of very different sizes. The research demonstrated that the choice between these two methods was never about which one was more powerful at finding truth; both found the truth at the same rate. The real issue was calibration: neither tool could be trusted to report the correct level of confidence across all types of tables.

To solve this, the author utilizes the new statistic T_root. This is a single, self-contained formula that adjusts itself based on the specific shape and size of the data table it is analyzing. Unlike the old methods, which require a complex decision tree to figure out which tool to use, T_root works the same way for every table, from small and empty grids to large and crowded ones. In extensive tests, this new method held its ground perfectly, reporting error rates that stayed very close to the intended target, even in the most difficult, uneven, or sparse scenarios. It also recovered the ability to detect real associations that the old "exact" tool had missed, effectively doubling the power to find true signals in the most problematic zones of data. The new method is fast, requires no complex adjustments, and returns a single, clear answer without needing to run thousands of computer simulations to get there.

The paper also looked at thousands of real-world tables from public datasets to see how often the current system fails. The findings were striking: roughly 94 percent of the tables analyzed fell into a category where the default software settings would produce an untrustworthy result. In these cases, the standard tools were either too conservative or too liberal, leading to incorrect conclusions about whether a relationship existed. One specific example involved a study of burn wounds, where the standard software said there was no significant link between a cleaning treatment and the need for surgery. The new method, however, detected a clear link, changing the conclusion from "no effect" to "significant effect" and providing a more accurate range for how strong that effect was. This kind of flip in verdict is not a rare anomaly but a common occurrence in the data scientists use every day.

The proposed solution is to change the default setting in statistical software. Instead of asking users to choose between different tests or relying on outdated rules about small numbers, the software should automatically report the new T_root statistic for almost every table. There is only one exception: if a table is so small and empty that it collapses into a shape too simple to analyze, the software should then fall back to the traditional "exact" method. This creates a simple, honest standard: use the new tool for everything, and only use the old one in the extreme corner where no other method can work. The author also recommends that reports include not just a single number, but also a measure of how well the test performed in that specific situation, ensuring that readers know exactly how much confidence to place in the result.

This work reframes a long-standing problem in statistics not as a puzzle of which tool to pick, but as a failure of reporting standards. The research shows that the confusion between "exact calculation" and "exact error control" has led to decades of misleading results. By adopting a single, calibrated statistic that works across the board, the scientific community can ensure that the conclusions drawn from data are both powerful and honest. The paper concludes that the era of complex decision trees and conflicting defaults should end, replaced by a straightforward rule that reports the truth, regardless of how messy the data might be.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →