Auditing Recorded Predictive Lead Service-Line Classifications Against Physical Verification: A Statewide Study of New York
This statewide study of New York reveals that while most utilities' predictive lead service-line models align with physical verifications, New York City's system shows a statistically significant discrepancy where a model classifies over 43,000 addresses as "Known Other" despite physical excavations confirming lead on nearly 121,000 addresses, alongside evidence that some public-side determinations were erroneously copied from customer-side model outputs.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the United States, a federal rule called the Lead and Copper Rule Revisions requires every community water system to create a complete map of its service lines, the pipes that carry water from the street to individual homes. The goal is to find and replace pipes made of lead, a toxic metal that can cause severe health problems, especially in children. For decades, utilities knew where lead pipes were because they had old paper records or because they had dug them up to look. But for millions of addresses where no records existed, the new rules allowed water companies to use computer models to guess the material of the pipe instead of digging. These models analyze data like the age of the house and the neighborhood to predict whether a pipe is likely lead, copper, or plastic. The idea was to save money and time, as digging up every single pipe is incredibly expensive. However, because these models are guesses, the rules require that utilities check their accuracy by physically inspecting a sample of the pipes they classified. The public interest hinges on whether these computer guesses are reliable enough to trust with public health, or if they are hiding a dangerous number of lead pipes under a layer of statistical confidence.
A researcher named Muhammad Sarmad Sohail decided to test this system by looking at the public records filed by water utilities across New York State. New York is unique because it publishes a detailed list for every single address, showing exactly how the utility determined the pipe material: whether they dug it up, checked a paper record, or used a computer model. Sohail did not try to prove that the computer models were wrong in a general sense; instead, he looked at the specific numbers the utilities filed to see if they made sense. He focused on the 153 localities in New York that used these computer models for at least 100 addresses. His first step was a simple check: he counted how many different answers the models gave. In a healthy system, a model should sometimes say "lead," sometimes "copper," and sometimes "unknown," depending on the specific house. But in 75 of those localities, covering nearly 126,000 addresses, the model gave only one single answer for every single house. It was as if a weather forecaster predicted rain for every single day of the year, regardless of the season or the location.
This lack of variety is not automatically a sign of fraud, but it is a warning sign that something is wrong. Sohail then compared these "one-answer" lists against the physical inspections that the same water utilities had performed in those same towns. In most cases, the single answer the model gave matched what the workers found in the ground. However, in seven specific places, the model's single answer directly contradicted what the workers saw. The most striking example was New York City. The city used a computer model to classify 43,215 addresses. For every single one of those addresses, the model said the pipe was "Known Other," a category that means the pipe is definitely not lead, even though the specific material is unknown. The city had found lead pipes in thousands of other places using different methods, and it had used the "unknown" category for over 120,000 other addresses. Yet, for this massive group of 43,000 homes, the model never once suggested a pipe might be lead or even unknown. It was a perfect, unbroken record of "not lead" for a group of houses that included many built before 1940, a time when lead pipes were common.
The contradiction was not just a statistical fluke; it was confirmed by the city's own workers. In the five boroughs of New York City, workers had physically dug up and inspected tens of thousands of pipes, finding lead in significant numbers. If the computer model had been working correctly, it should have flagged some of those 43,000 addresses as potentially lead. Instead, it cleared them all. The same strange pattern appeared in East Rochester, a small village 550 kilometers away with a completely different water company. There, the model also gave a single "not lead" answer for every house it classified, even though workers had dug up pipes in that village and found lead in nearly 10% of them. The researcher also discovered that the city's public records for these pipes were not created independently. The data showed that the city had simply copied the computer guess made for the customer's side of the pipe and pasted it onto the public side, even though the public pipe is a different pipe owned by the city. This suggests the determination was not a fresh analysis but a copy-paste error that erased any uncertainty the model might have had.
To understand how many lead pipes might be hiding in that group of 43,215 addresses, the researcher looked at the age of the buildings. The model had been applied mostly to newer homes, which naturally have fewer lead pipes. However, even after adjusting for the age of the buildings, the numbers did not add up. In older buildings, where other methods found lead, the model found nothing. By using the physical inspection results from the same neighborhoods as a guide, the researcher estimated that between 1,150 and 1,450 of those 43,215 addresses likely have lead pipes that were incorrectly labeled as safe. This estimate relies on the assumption that the workers who dug up the pipes were not biased, but even with this uncertainty, the gap between the model's perfect record and the reality of the ground is too large to ignore.
The study does not claim that computer models can never work, nor does it say that every utility is failing. It found that in many places, the models produced a mix of answers that matched the physical reality. The problem was specific to the way the data was reported in these seven locations. The researcher argues that the issue is not necessarily the computer code itself, but how the results were handled before they were published. It is possible that the model did produce a range of answers, but a computer program or a human process filtered out anything that wasn't a clear "safe" result before it was filed with the state. This would mean the public record shows a clean, confident list that hides the uncertainty the model actually had. The state's own rules require utilities to provide a confidence interval—a measure of how sure they are—and to have a plan for checking the results. The records filed by these utilities showed none of that.
This audit was possible only because New York publishes the method used for every single address. In most other states, the public does not know whether a pipe classification came from a shovel or a spreadsheet. The researcher suggests that requiring this level of detail nationwide would cost nothing but would allow anyone to check if a utility is hiding uncertainty behind a single, unchanging answer. The findings serve as a warning that when a model produces a perfect, unvarying result across thousands of homes, it is likely not a sign of a perfect model, but a sign that the process has lost its ability to see the truth. The data shows that in the largest case, the filing did not meet the conditions the state set for using that specific label. The researchers conclude that inventories should record the uncertainty of the model, not just its final guess, and that regulators should treat a list with zero variation as a red flag that needs immediate investigation.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.