Examining the Efficacy of Coarsened Exact Matching as an Alternative to Propensity Score Matching
This paper argues that Coarsened Exact Matching (CEM) is not a superior alternative to Propensity Score Matching (PSM) because it suffers from residual confounding due to its inexact nature, fails to outperform PSM in reducing imbalance when accounting for random variation, and becomes unstable and inefficient in high-dimensional settings, whereas PSM remains more robust to model misspecification.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to figure out if a new medicine works. The gold standard is a Randomized Clinical Trial (RCT), where you flip a coin to decide who gets the medicine and who gets a placebo. Because it's random, the two groups are perfectly balanced: the people who got the medicine are, on average, just as healthy, old, and wealthy as those who didn't.
But in the real world, we often can't flip coins (it's too expensive or unethical). So, we look at observational data (like hospital records). Here, the "medicine group" might be sicker or older than the "control group" to begin with. This is called confounding. To fix this, statisticians use "matching" techniques to pair up sick people from the medicine group with similar sick people from the control group, trying to recreate that perfect coin-flip balance.
For a long time, the standard tool for this was Propensity Score Matching (PSM). Recently, a new tool called Coarsened Exact Matching (CEM) became very popular. Many researchers claimed CEM was the "superior" version of PSM.
This paper, written by Fei Wan, argues that CEM is actually not the superhero everyone thinks it is. In fact, for many situations, PSM is the better, more reliable tool.
Here is the breakdown of why, using simple analogies:
1. The "Fuzzy Photo" vs. The "Blurry Snapshot"
The Claim: CEM is often called "Exact Matching," but it isn't exact. It's actually "Fuzzy Matching."
- The Analogy: Imagine you are trying to match people based on their height.
- PSM is like taking a high-resolution photo of everyone's height and finding the person who is closest in height to your target. Even if they aren't exactly the same height, the math ensures that if you look at the whole group, the average height is perfectly balanced. Any small differences are just random noise (like a slight blur in a photo) that cancels itself out when you have enough people.
- CEM is like putting everyone into buckets: "Short," "Medium," and "Tall." You then match people who are in the same bucket.
- The Problem: Inside the "Medium" bucket, you might have someone who is 5'6" and someone who is 5'10". They are in the same bucket, but they are very different. This creates a systematic error (residual confounding) that doesn't go away, no matter how many people you add to the study. It's like trying to compare apples and oranges just because they are both in the "Fruit" basket.
2. The "Tightrope" of Complexity
The Claim: CEM falls apart when you have too many variables (the "Curse of Dimensionality").
- The Analogy: Imagine you are trying to find a perfect twin for a person based on 10 different traits (height, weight, eye color, shoe size, favorite pizza, etc.).
- CEM tries to put everyone into boxes based on all 10 traits at once. As you add more traits, the boxes get so specific that most of them end up empty. You lose so much data (people who don't fit in any box) that your results become unstable and unreliable. It's like trying to find a needle in a haystack, but the haystack keeps shrinking until there's nothing left.
- PSM doesn't try to match on all 10 traits directly. Instead, it calculates a single "score" (a summary of all 10 traits) and matches people based on that score. This avoids the empty boxes problem and keeps the data stable, even when you have many variables.
3. The "Guessing Game" (Model Dependence)
The Claim: CEM forces you to make perfect guesses about the data later on, while PSM is more forgiving.
- The Analogy:
- CEM is like a puzzle where you force the pieces together by sanding them down (coarsening) so they fit. Because you sanded them, the pieces don't fit perfectly anymore. To get the right picture, you now have to use a very complex, perfect formula to "fix" the gaps. If your formula is even slightly wrong, the whole picture is ruined.
- PSM is like a puzzle where the pieces fit naturally. Even if you use a simple, slightly imperfect formula to look at the result, the picture still looks mostly correct because the pieces were balanced to begin with. PSM is "robust," meaning it doesn't break as easily when you make a small mistake in your math.
4. The "Ruler" Problem
The Claim: Previous studies said PSM was bad at balancing groups because they used the wrong measuring stick.
- The Analogy: Some researchers claimed PSM was failing because they used a ruler that only measured the distance between people, ignoring the direction.
- The author argues that in PSM, if one person is slightly taller and another is slightly shorter, these random differences cancel each other out.
- The old measuring tools (like Mahalanobis distance) treated these random ups and downs as "errors" that were getting worse.
- The author proposes a new way to measure (Standardized Mean Differences) that accounts for these random ups and downs. When you use this new ruler, PSM looks much better at balancing groups than CEM.
The Bottom Line
The paper concludes that CEM is not a magic bullet.
- If you have a small number of variables, CEM might work okay, but it still has hidden flaws.
- If you have many variables (which is common in real-world data), CEM causes you to lose too much data and introduces systematic errors that are hard to fix.
- PSM remains the more robust, reliable choice because it handles randomness better, requires less perfect math to get a good result, and doesn't throw away as much data.
The author warns that many doctors and researchers have been using CEM thinking it's the "exact" solution, but it's actually an "inexact" shortcut that can lead to biased results if not handled with extreme care.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.