← Latest papers
📊 statistics

Propensity score adjustment when errors in achievement measures inform treatment assignment

This paper introduces a propensity score adjustment method that accounts for measurement error in noisy achievement data to improve subgroup balance and reduce bias in evaluating interventions aimed at closing achievement gaps, as demonstrated through simulations and a Texas summer learning loss initiative.

Original authors: Joshua Wasserman, Ben B. Hansen, Michael R. Elliott

Published 2026-06-16
📖 5 min read🧠 Deep dive

Original authors: Joshua Wasserman, Ben B. Hansen, Michael R. Elliott

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a school district trying to figure out if a new program (let's call it "Summer School Plus") actually helps students learn. To do this fairly, you need to compare schools that used the program with schools that didn't. But here's the catch: schools only get picked for the program if they have big "achievement gaps" (big differences in test scores between different groups of students).

The problem is that test scores are like noisy snapshots. If a school has a tiny group of students (say, only 5 or 6 kids) in a specific demographic, their average test score for that year might be wildly high or low just by pure luck. Maybe those 5 kids just had a great day, or maybe they were all sick. This "noise" doesn't reflect their true, long-term ability.

The authors of this paper argue that if you try to match schools based on these noisy, lucky (or unlucky) snapshots, you might end up comparing apples to oranges. You might match a school that got picked because of a lucky fluke with a school that actually has a real problem.

The Two New Tools

The authors propose two new ways to fix this "noise" before you try to match the schools. Think of these as tools to see through the static on a radio to hear the real music.

1. The "Regression Calibration" (RC) Method: The Crystal Ball
Imagine you have a blurry photo of a person. You know they are wearing a red hat, but the face is fuzzy. The RC method is like using a crystal ball (specifically, a statistical model called a Hierarchical Linear Model) to guess what the person really looks like based on the blurry photo and other clear details you have (like their height or hair color).

  • How it works: Instead of using the noisy test score directly, the method uses the crystal ball to predict the student's "True Score" (what they would get if they took the test 1,000 times). It then uses this predicted "True Score" to find matching schools.
  • The Result: It smooths out the luck factor, helping you find schools that are truly similar in their underlying ability, not just in their lucky test day.

2. The "Maximum Likelihood" (ML) Method: The Smart Filter
This method is a bit more sophisticated. Imagine you are trying to sort a pile of mixed-up puzzle pieces. Some pieces are clear, and some are covered in mud. The ML method doesn't just try to guess what's under the mud; it builds a complete picture of how the mud got there in the first place.

  • How it works: It creates a mathematical model that understands both the true ability of the students and the "mud" (the measurement error) that got on the test scores. It calculates the probability of a school being picked for the program by mathematically "filtering out" the noise.
  • The Result: This method is particularly good at handling the "luck" factor. It creates a list of schools to compare that overlap much better than the old methods. It's less likely to throw away a school just because its test score was a fluke.

Why This Matters (The "Less is More" Idea)

Usually, in statistics, you want to use all the data you have. But the authors found that when you have small groups of students, using the raw, noisy data actually makes your comparison worse. It's like trying to balance a scale with a feather that keeps blowing in the wind; the scale tips wildly.

By using their new methods to estimate the "True Score" instead of the "Noisy Score," they found:

  • Better Matches: They could find more schools to compare. The old methods often had to throw away schools because their noisy scores made them look too different. The new methods realized those schools were actually similar once you accounted for the noise.
  • Fairer Results: When they tested these methods on a real program in Texas (ADSY) designed to stop summer learning loss, the new methods created groups that were much more balanced.
  • Handling Missing Data: In many states, if a school has fewer than 5 students in a group, the government hides the test score to protect privacy. The old methods couldn't use these schools at all. The new methods can still estimate the "True Score" for these schools using the data they do have, allowing them to be included in the study.

The Bottom Line

The paper shows that when you are dealing with small groups of students, ignoring the "noise" in test scores leads to bad comparisons. By using these new statistical "filters" (RC and ML) to guess the true ability behind the noisy scores, researchers can make fairer comparisons, keep more schools in their studies, and get more accurate answers about whether a school program actually works.

They tested this with computer simulations and a real-world example in Texas, and in both cases, their new "filters" worked better than the standard way of doing things.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →