Weighted k-Sample Kolmogorov-Smirnov, Cramer-von Mises, and Anderson-Darling Tests for Assessing Covariate Balance
This paper extends weighted Kolmogorov-Smirnov, Cramer-von Mises, and Anderson-Darling tests to assess covariate balance across arbitrary numbers of groups (k ≥ 2), demonstrating that while omnibus power decreases with more groups, the accompanying post-hoc pairwise procedure remains a sensitive tool for detecting specific imbalances, with all methods implemented in Stata.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to solve a mystery: who really did it? In the world of science, specifically when researchers try to figure out if a new medicine or policy actually works, they play a similar game. They compare a group of people who got the treatment against a group who didn't. But here's the catch: for the comparison to be fair, the two groups must be twins in every way that matters—same age, same health history, same habits. If they aren't twins, the results are just a mess of confusion. This is called checking for "covariate balance."
Usually, scientists just check if the average age or average income is the same. But averages can be tricky. Imagine two groups where the average age is 30 in both. In one group, everyone is exactly 30. In the other, half are 10 and half are 50. The averages match, but the groups are totally different! To catch these sneaky differences, scientists use special "distributional tests." Think of these as high-tech scanners that don't just look at the average height of a crowd, but scan the entire shape of the crowd to see if the groups are truly identical. For a long time, these scanners only worked when comparing two groups at a time. But what if you have three, four, or more groups to compare? That's the puzzle this paper tackles.
The paper, written by Ariel Linden, introduces a new way to use three famous statistical scanners—the Kolmogorov–Smirnov, Anderson–Darling, and Cramér–von Mises tests—to compare any number of groups (two or more) at once, even when the data is weighted (like giving more importance to certain people). The author also adds a clever "post-hoc" tool, which is like a magnifying glass that helps you pinpoint exactly which groups are different if the big scan says something is wrong.
Here is the story of what they found. First, the author built these new multi-group scanners and tested them in a virtual world using computer simulations. They created thousands of fake scenarios with three groups to see if the scanners would make mistakes. The good news? The scanners were very honest. When the groups were actually identical, the scanners rarely cried "foul," keeping their error rate right where it should be (around 5%).
However, the simulations revealed a surprising twist. When the scientists compared three groups instead of two, the big "omnibus" scanner (the one that checks all groups at once) became slightly less sensitive. It was harder for the scanner to spot a single weird group hiding among two normal ones. The author explains this isn't because the weird group became less weird; it's because the scanner had to look at more groups, which made the "normal" range of what it expects to see much wider. It's like trying to hear a whisper in a quiet room versus a whisper in a noisy stadium; the whisper is the same, but the background noise of the stadium makes it harder to detect.
Because of this, the paper suggests a smart strategy: if you suspect one specific group is different, don't just rely on the big group scan. Instead, use the new "post-hoc" tool to compare the groups two-by-two. This pairwise tool doesn't suffer from the "noise" of the extra groups and is much better at catching the specific culprit.
The paper also confirmed that the three different scanners still have their own special superpowers, just as they do when comparing two groups. The Kolmogorov–Smirnov test is the best at spotting differences right in the middle of the data (like if everyone in one group is slightly taller than the others). The Anderson–Darling test is the champion at spotting differences at the very edges or "tails" (like if one group has a few extremely sick people that the others don't). The Cramér–von Mises test is a solid all-rounder, performing similarly to Anderson–Darling when differences are spread out everywhere.
To show this works in the real world, the author applied these methods to data from a heart failure disease management program. They had three groups: people who didn't join, people who got phone calls, and people who got remote monitoring. Before adjusting the data, the scanners screamed that the groups were totally different. But after applying a special weighting method to balance them out, the scanners said, "All clear!" The groups were now statistically twins.
In short, this paper gives scientists a new toolkit to compare multiple groups fairly. It warns that checking all groups at once might miss a single oddball, but it provides a specific follow-up tool to find that oddball. It confirms that different tests are better for different types of differences, and it proves that these methods work well even when the data is complex and weighted. The author is confident in these findings based on their simulations, though they note that future work could explore even more complex group setups.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.