← Latest papers
📊 statistics

Detecting and Quantifying Circularity Bias When Composite Quality Ratings Are Used as Covariates: A Methodological Demonstration Using CMS Hospital Excess Readmissions Data

This paper demonstrates that including composite quality ratings as covariates in regression models introduces a circularity bias that systematically attenuates the estimated effects of structural predictors, a methodological hazard confirmed through CMS hospital readmission data and Monte Carlo simulations.

Original authors: Ali Akram

Published 2026-06-30
📖 5 min read🧠 Deep dive

Original authors: Ali Akram

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: Don't Use the Answer Key to Grade the Test

Imagine you are a teacher trying to figure out if Student A (a for-profit hospital) performs differently than Student B (a non-profit hospital) on a specific math test (hospital readmissions).

You want to be fair, so you decide to adjust the scores based on how "smart" the students are overall. You pull out a report card that gives them a Star Rating based on their grades in Math, Science, History, and Art.

Here is the trap: The "Math Grade" is actually part of that Star Rating.

If you include the Star Rating in your calculation to "adjust" for the students' overall ability, you are accidentally using the Math Grade to explain the Math Grade. It's like asking, "How well did they do on the Math test?" and then saying, "Well, they got a high Star Rating, which includes their Math score, so they must have done okay."

This creates a circular logic loop. Because the Star Rating already contains the answer you are trying to measure, using it as a "control" makes the difference between the two students look much smaller than it actually is. You end up underestimating the real gap.

What the Author Did

The author, Ali Akram, investigated this problem using real data from the U.S. government (CMS) regarding hospital readmissions.

  1. The Setup: He looked at 2,446 hospitals. He wanted to see if For-Profit hospitals had higher readmission rates (more patients coming back) than Non-Profit hospitals.
  2. The Mistake: Many researchers try to be thorough by adding the "CMS Overall Star Rating" to their math models to "adjust" for quality.
  3. The Problem: The Star Rating is built partly from the readmission numbers. So, adding it to the model is like putting the answer key on the right side of the equation.
  4. The Result:
    • Without the Star Rating: For-profit hospitals had a readmission rate that was 0.0196 higher than non-profits. This is a clear, noticeable difference.
    • With the Star Rating: The difference shrank to 0.0113.
    • The Takeaway: By including the Star Rating, the researchers accidentally hid nearly half (48%) of the real difference. The "adjustment" made the problem look less severe than it was.

The "Magic Trick" to Detect This

The author proposes a simple test to catch this error, which he calls a Diagnostic:

  • Step 1: Run your analysis without the Star Rating.
  • Step 2: Run the analysis with the Star Rating.
  • Step 3: Compare the results.

If the number you are interested in (the difference between for-profit and non-profit) suddenly shrinks significantly when you add the Star Rating, you have found the "circularity bias." It's a red flag telling you: "Stop! You are using the outcome to explain the outcome."

The Simulation: Proving It's Not a Fluke

To prove this wasn't just a weird coincidence with this specific data, the author ran a computer simulation (a "Monte Carlo" test).

  • He created a fake world where he knew the true answer: For-profit hospitals were definitely 20 points worse.
  • He then ran the math with and without the fake "Star Rating."
  • Result: When he added the fake Star Rating, the math automatically cut the true difference in half (down to 10 points).
  • Conclusion: This proves the bias is a mathematical certainty, not just a fluke of the hospital data.

Another Finding: Geography Doesn't Matter Much

The paper also looked at whether hospitals in different parts of the country (North, South, East, West) performed differently.

  • The Statistic: The math said the differences between regions were "statistically significant" (meaning they weren't just random noise).
  • The Reality Check: The author used a tool called "variance decomposition" (think of it as a pie chart of where the differences come from).
  • The Result: 98.9% of the differences happened inside the regions (between individual hospitals), and only 1.1% happened between the regions.
  • The Lesson: Even though the math says "Region matters," in the real world, it barely matters at all. A hospital's location explains almost none of the variation in readmission rates.

The Bottom Line

  1. Don't use the Star Rating as a control variable if you are studying the specific metric (like readmissions) that is inside that Star Rating. It will make your results look weaker and less important than they really are.
  2. Use the "Two-Model Test": Always check if adding a composite rating changes your main result drastically. If it does, you are likely falling into a circular trap.
  3. Context Matters: Just because a result is "statistically significant" doesn't mean it's "important." In this case, geography was statistically significant but practically irrelevant.

Important Note: The author clarifies that removing the Star Rating doesn't make the study a "perfect" proof of cause-and-effect. There are still other factors (like patient health) that aren't accounted for. However, removing the Star Rating gives a cleaner, more accurate description of the relationship between ownership and readmissions, free from this specific mathematical error.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →