← Latest papers
💻 computer science

Predicting the Inspector, Not the Plant: Regulatory Targeting and Explanation Instability in Pharmaceutical Inspection-Outcome Models

This study demonstrates that machine learning models predicting pharmaceutical inspection outcomes primarily capture the FDA's own regulatory targeting processes rather than inherent manufacturing quality, resulting in unstable feature explanations and the risk that deploying such models would amplify existing inspection biases.

Original authors: Phani Kumar Balagam

Published 2026-09-16
📖 5 min read🧠 Deep dive

Original authors: Phani Kumar Balagam

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Every year, the United States Food and Drug Administration sends teams of inspectors to pharmaceutical factories to ensure the medicines people take are safe and made correctly. These inspections are critical; they are the primary way regulators catch problems before they reach patients. When an inspector finds a serious violation, the factory receives a formal warning, and the FDA may take further legal action. Because these inspections are expensive and time-consuming, regulators and scientists have long wondered if they could use computers to predict which factories are most likely to fail. The idea is simple: if a computer could look at a factory's history and flag the risky ones, inspectors could focus their limited time where it is needed most. This field of study, known as regulatory science, tries to turn years of public records into a clear signal of danger. But to do this, one must understand that the data itself is not just a record of factory behavior; it is also a record of the regulator's own choices about where to look.

A new study by independent researcher Phani Kumar Balagam takes a hard look at this idea. The researcher built a computer model using nearly forty thousand real inspections conducted between 2008 and 2026. The model was designed to predict whether a factory would receive a serious violation. The results were surprisingly good at first glance: the model could distinguish between safe and risky factories with a high degree of accuracy, performing about four times better than random guessing. However, when the researcher peeled back the layers to see exactly what the computer was learning, a different story emerged. The model was not just learning about the factories; it was learning about the FDA's own inspection strategy.

The most striking discovery was that the single most powerful clue the model used was not a detail about the factory's equipment or its workers, but a simple code indicating which specific FDA program had assigned the inspection. This single piece of information, which tells us whether the visit was part of a routine check, a clinical trial audit, or an investigation into unapproved drugs, was just as important as the entire factory's history of past violations, warnings, and citations combined. In fact, the researcher found that removing the factory's entire history of misconduct from the model reduced its accuracy significantly, while removing the program code caused a similar drop. This suggests that the computer was not simply identifying "bad" factories. Instead, it was identifying the types of visits the FDA had already decided were likely to go badly. The model was effectively learning the regulator's intuition about where to look, rather than the factory's actual quality.

The study also revealed that the definition of "risky" is not fixed; it changes based on policy. During the pandemic in 2020, the FDA had to cancel many routine inspections abroad and focus its limited resources on domestic factories it already suspected of problems. As a result, the rate of serious violations found in the United States doubled, while the rate for the same type of inspections in other countries remained flat. The computer model, if left unchanged, would have failed to predict this shift because it was trained on old data where the rules were different. The target itself had moved, driven by a change in government policy rather than a sudden collapse in manufacturing quality.

Perhaps the most unsettling finding concerned the explanations the model provides. When data scientists build these tools, they often ask the computer to explain which factors drove its decision, hoping to give inspectors a clear reason for the warning. The researcher tested these explanations by running the model many times with slightly different data samples to see if the answers stayed the same. They found that the explanations were unstable. The computer would point to one factor as the main cause of risk in one run, and a completely different factor in the next, even though the overall prediction remained the same. The standard methods used to check these explanations were not reliable enough to trust in a high-stakes environment.

The paper concludes that while such a model can technically predict inspection outcomes, it is dangerous to use it to prioritize future inspections. If regulators use the model to send inspectors to the factories it flags, they will simply be sending them to the same places they already visit, creating a self-fulfilling loop. The model would confirm its own bias rather than finding new problems. The researcher argues that these tools are not just models of manufacturing quality, but models of the regulatory process itself. Until the system can account for the fact that the regulator's choices shape the data, using these predictions to decide where to send inspectors would amplify existing targeting rather than improve safety. The study serves as a reminder that in the complex world of drug safety, the most accurate number is not always the most useful one, and understanding the source of a prediction is just as important as the prediction itself.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →