← Latest papers
💻 computer science

Hidden Errors in Big Data: The Case of Property Records

This paper audits prominent brokered property datasets and reveals systematic errors and coverage gaps that bias key measures of economic inequality and property tax regressivity, underscoring the critical need for open administrative data and greater transparency regarding data provenance.

Original authors: Evelyn Smith, Emma Harvey, Jacob Goldin, Daniel E. Ho

Published 2026-08-03
📖 7 min read🧠 Deep dive

Original authors: Evelyn Smith, Emma Harvey, Jacob Goldin, Daniel E. Ho

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Invisible Glitch in the World's Biggest Data Library

Imagine you are trying to solve a giant puzzle, but instead of using the actual pieces, you are using a photocopy of a photocopy that someone else made. This is the world of "Big Data." Today, scientists, governments, and even the artificial intelligence systems that run our cities rely on massive piles of information to make decisions. They use this data to figure out everything from how to treat diseases to how to tax your house. But here is the catch: most of this data isn't collected directly by the scientists. It is bought from "data brokers"—companies that gather information from thousands of different places, clean it up, and sell it as a neat, ready-to-use package.

The big question this paper asks is: What if the photocopy is wrong?

When you buy a dataset, you are trusting that the broker did their job perfectly. But what if they missed some pieces? What if they copied a number wrong? Or what if they used a weird trick to guess a missing number, and that trick made the whole picture look distorted? This is known as the "Big Data Paradox": the more data you have, the more confident you feel, but if that data is flawed, your confidence is actually a trap. You might get a very precise answer that is completely misleading. This paper dives into one specific, high-stakes corner of this problem: the data used to track how much houses sell for and how much tax they owe. If this data is broken, it could mean that the rich are paying less tax than they should, or that the poor are paying too much, and nobody would even know it until it's too late.

The Great House-Price Heist: When Data Brokers Get It Wrong

In this study, the authors acted like digital detectives, auditing two of the biggest "data brokers" in the United States: Cotality and ATTOM. These companies are the giants of the property world; they sell information about millions of homes to banks, governments, and researchers. The authors decided to check their work against the "ground truth"—the original, official records kept by the Cook County government in Illinois (which includes Chicago). Think of the government records as the original, signed receipt from the store, and the brokers' data as the copy you get from a third-party app.

The detectives found that the brokers' copies were far from perfect. They discovered two main types of errors that were messing up the math:

1. The "Missing Piece" Problem (Coverage Errors)
Imagine you are counting the apples in a basket, but the person who made the list forgot to write down 12% to 15% of the apples. That's a "coverage error." The authors found that for every 100 real house sales that happened, the brokers were missing about 12 to 15 of them. Why? Because the brokers sometimes got confused about what counted as a "house" or a "real sale." They might have accidentally labeled a house sale as a "foreclosure" (which is a different category) or missed a sale entirely because the address was written slightly differently. This meant that the data wasn't just a little off; it was missing a huge chunk of the population, making the picture of the housing market incomplete.

2. The "Guessing Game" Problem (Imputation Errors)
Sometimes, the brokers didn't have the actual sale price of a house. Instead of saying "I don't know," they tried to guess it using a math trick. They looked at the "transfer tax" (a small fee paid when a house is sold) and tried to work backward to figure out the price. But here is where it went wrong: different towns have different tax rates. If the broker guessed the wrong town's tax rate, their math would be off by a huge amount.

The authors found that for about 1% to 2% of the house sales they checked, the broker's price was wildly different from the real price—sometimes off by hundreds of thousands of dollars! Even stranger, the errors weren't random. The brokers often made the exact same mistake on the exact same houses. For example, if a house sold for $300,000, a broker might accidentally record it as $200,000 (two-thirds of the real price) or $600,000 (double the real price). It was like a stamping machine that kept pressing the wrong number over and over again.

The Domino Effect: Why This Matters for Your Wallet

You might think, "So, a few house prices are wrong. Who cares?" The authors show that this matters a lot because these numbers are used to calculate something called property tax regressivity.

Think of property tax like a group dinner where everyone is supposed to pay a share based on how much they ordered. "Regressivity" is a fancy word for when the people who ordered the cheap meal end up paying a bigger percentage of the bill than the people who ordered the expensive steak. In the real world, studies often show that poorer neighborhoods are taxed at a higher effective rate than wealthy ones.

The authors used the brokers' data to calculate this tax fairness. The result? The brokers' data gave a completely different answer than the real government data.

  • Cotality's data made it look like the tax system was fairer than it really was (underestimating the unfairness).
  • ATTOM's data made it look like the tax system was more unfair than it really was (overestimating the unfairness).

Because the brokers' data had these hidden errors, the "truth" about who is paying what shifted dramatically depending on which data you bought. If a government official used the wrong data, they might think the tax system is fine when it's actually hurting poor families, or vice versa.

The "Copy-Paste" Conspiracy

One of the most surprising discoveries was that the two big brokers, Cotality and ATTOM, were making the same mistakes. They missed the same houses, and they made the same weird math errors on the same transactions. It turns out that these companies often share data with each other. It's like if two news stations both got their breaking news from the same wire service, and that wire service had a typo. Both stations would report the same wrong story, and nobody would realize it was a mistake because they both agreed on it.

The authors found that for the houses where the brokers got the price wrong, 99.8% of the time, both brokers had the exact same wrong number. This is a big problem for scientists. Usually, if two different sources agree, you think, "Great, this must be true!" But in this case, their agreement was a trap. They were both wrong for the same reason.

The Bottom Line

This paper doesn't just point out a few typos; it reveals a crack in the foundation of how we understand the economy. The authors conclude that we cannot blindly trust the "neat packages" of data sold by brokers. Because these companies don't always tell us how they clean or guess their numbers, we are flying blind.

The solution? We need open data. Governments already have the original, correct records (the "ground truth"). The authors argue that instead of relying on expensive, opaque brokers, researchers and policymakers should use the free, public data directly. If we can't see how the data was made, we can't trust the conclusions we draw from it. In a world where AI and big data are running the show, the most important thing we can do is make sure the data isn't just big—it's also right.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →