The effect of collinearity and sample size on linear regression results: a simulation study
This simulation study demonstrates that while collinearity primarily reduces precision in small samples without causing bias under correct model specification, it significantly amplifies bias and degrades inference under misspecification, arguing against the mechanical application of fixed VIF thresholds without considering sample size and potential omitted-variable bias.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Question: When is "Too Much Similarity" a Problem?
Imagine you are trying to figure out how much a specific ingredient (let's say, salt) affects the taste of a soup. You have a recipe with many ingredients: salt, pepper, garlic, onions, and carrots.
In statistics, this is called linear regression. You want to know the true effect of the salt.
However, sometimes ingredients are very similar to each other. For example, if you always add salt and pepper in a fixed ratio (whenever you add a pinch of salt, you add a pinch of pepper), it becomes hard to tell which one is actually making the soup salty. In statistics, this is called collinearity.
To measure this "similarity," researchers use a score called the VIF (Variance Inflation Factor).
- VIF = 1: The ingredients are totally unique (no similarity).
- VIF = 10 or higher: The ingredients are very similar (high collinearity).
The Old Rule: For a long time, scientists have followed a rigid rule of thumb: "If the VIF is above 4 or 10, throw one of those ingredients out of the recipe immediately." They do this regardless of how big their cooking pot is.
The New Study: This paper asks: Does that rule actually make sense? Does it matter if we are cooking a tiny cup of soup or a massive industrial vat?
The Experiment: Cooking with Different Pot Sizes
The researchers ran a massive computer simulation (a "digital kitchen") to test this. They cooked thousands of batches of soup with two main variables:
- The Pot Size (Sample Size): Ranging from a tiny cup (100 servings) to a massive industrial vat (100,000 servings).
- The Similarity (Collinearity): Ranging from totally unique ingredients (VIF=1) to ingredients that are almost identical twins (VIF=50).
They then checked four things:
- Accuracy: Did they get the right answer?
- Confidence: How sure were they of that answer?
- Precision: Was the answer tight and specific, or a wide, fuzzy guess?
- Bias: Did they accidentally leave out a key ingredient that changed the flavor?
The Surprising Findings
Here is what they discovered, broken down by pot size:
1. The Tiny Pot (Small Sample Size: ~100 people)
The Problem: In a tiny pot, even a little bit of similarity between ingredients causes chaos.
- The Analogy: Imagine trying to guess the exact amount of salt in a single teaspoon of soup. If the salt and pepper are slightly mixed up, your guess will be all over the place.
- The Result: Even a low similarity score (VIF < 2) made the results unreliable. The "confidence intervals" (the range of possible answers) became so wide that the study lost its power.
- The Lesson: In small studies, you cannot ignore even mild similarity.
2. The Massive Vat (Large Sample Size: ~50,000+ people)
The Good News: In a huge pot, the rules change completely.
- The Analogy: Imagine trying to guess the salt in a swimming pool full of soup. Even if the salt and pepper are mixed up perfectly, the sheer volume of soup makes the average taste incredibly clear. The "noise" of the similarity gets drowned out by the sheer amount of data.
- The Result: Even with extreme similarity (VIF = 50), the results remained accurate and precise. The confidence intervals stayed tight.
- The Lesson: In massive datasets, you can often ignore high VIF scores. They don't hurt your results.
3. The "Missing Ingredient" Trap (Bias)
This is the most critical warning in the paper.
- The Scenario: What happens if you decide to fix the "similarity problem" by throwing out an ingredient (like the pepper) just to make the math easier?
- The Result: If that ingredient actually mattered (even if it was similar to the salt), removing it creates a bias.
- The Analogy: If you throw out the pepper because it's too similar to the salt, but the pepper actually adds a spicy kick, your soup will taste wrong. The more similar the ingredients were, the worse the flavor gets when you remove one.
- The Lesson: Removing variables just to lower the VIF score can actually make your results more wrong, not less.
The Takeaway: Stop Using Rigid Rules
The paper concludes that the old rule of "If VIF > 10, delete the variable" is broken.
- Don't be a robot: You cannot apply the same rule to a small study and a massive study.
- Context is King:
- If you have a small study, even small similarities are dangerous.
- If you have a huge study, even extreme similarities are usually harmless.
- Don't throw things away: If you remove a variable just to fix a VIF score, you might introduce a new, bigger problem (bias) that ruins your conclusion.
The Bottom Line: Instead of blindly following a number on a chart, researchers need to look at their sample size. If they have a lot of data, they can likely keep all their variables, even if they are similar. If they have little data, they need to be very careful, but they still shouldn't just delete variables without thinking about the consequences.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.