Testing Claimed Associations of Prenatal Acetaminophen with Neurodevelopmental Disorders: Re-analysis Using Bias-Correction Methods
This reanalysis challenges the claim of "strong evidence" linking prenatal acetaminophen exposure to neurodevelopmental disorders by demonstrating that, after applying standard meta-analytic and bias-correction methods, the association for autism spectrum disorder attenuated to non-significance and the evidence for other disorders was inconsistent, leaving only a modest, uncertain association with ADHD.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Pregnant women often face a difficult choice when fever or pain strikes: take a common over-the-counter medication to feel better, or avoid it for fear of harming the developing baby. For decades, acetaminophen has been the standard recommendation for this dilemma, viewed as the safest option. However, in recent years, a growing body of research has suggested a troubling link between taking this medication during pregnancy and neurodevelopmental disorders in children, specifically attention-deficit/hyperactivity disorder and autism spectrum disorder. These conditions affect how the brain develops, influencing behavior, learning, and social interaction. When a large review of existing studies claimed there was "strong evidence" for this link, it sparked intense debate, leading to federal health advisories and a re-evaluation of medical guidelines. The core question became whether the observed connections were real biological effects or merely the result of how the studies were counted and analyzed.
A team of researchers at Harvard Medical School and Stanford University decided to look behind the curtain of that original review. They did not collect new data or conduct new experiments. Instead, they took the exact same set of forty-six studies that the original review had examined and applied a different, more rigorous mathematical approach to them. Their goal was to see if the "strong evidence" verdict would hold up when the studies were weighed correctly and when the analysis accounted for the ways scientists sometimes unintentionally skew results. They treated the original review's conclusion not as a final fact, but as a hypothesis to be stress-tested.
The researchers began by fixing a fundamental counting error in the original work. In the initial review, some studies were counted multiple times because they reported several different sub-analyses, while larger studies with fewer sub-analyses were counted less. This gave disproportionate weight to smaller studies with many sub-questions. The new team corrected this by ensuring each study counted only once for each specific outcome. When they did this simple adjustment, the link between the medication and autism spectrum disorder vanished. The statistical evidence that once suggested a connection disappeared, leaving a result that was indistinguishable from random chance.
Next, the team applied eight different statistical tools designed to detect and correct for publication bias. These methods act like a filter, asking whether the studies that made it into the review were the only ones that found a positive result, while studies finding no effect were hidden away in file drawers. They also tested the "credibility ceiling," a concept that acknowledges that no single observational study can be 100% certain about a cause-and-effect relationship because of hidden factors that researchers cannot measure. When they applied these filters, the link to autism remained non-existent. The connection to other neurodevelopmental disorders also fell apart, with the statistical signal reversing direction in some cases, suggesting the original findings were likely artifacts of the analysis rather than real-world truths.
The link to attention-deficit/hyperactivity disorder was the only one that showed some resilience, but even it was fragile. While a small signal persisted under most methods, it lost statistical significance when the researchers assumed a modest level of uncertainty, a standard requirement for trusting observational data. Furthermore, when the team converted all the different types of measurements used in the studies into a single, consistent scale, the evidence for this link weakened significantly. The funnel plot, a visual tool used to spot bias, revealed that the original analysis had hidden a significant asymmetry; once the data was harmonized, the plot showed that many missing studies with negative results were likely unreported, a classic sign of publication bias.
The researchers also investigated why the original review had rated certain studies as "high quality." They discovered a circular logic in the scoring system: the review gave higher scores to studies that found larger effects, effectively penalizing the very studies that used the most robust methods to rule out family-based confounding. Studies that compared siblings to control for shared genetics and environment, which are the strongest tools for proving causality, were scored lower because they found weaker or no associations. When the new team filtered the data to include only the studies the original review deemed "best," the signal for the medication's harm actually grew stronger, confirming that the scoring system was biased toward finding an effect rather than finding the truth.
Ultimately, this re-analysis suggests that the claim of "strong evidence" linking prenatal acetaminophen use to neurodevelopmental disorders does not survive a rigorous quantitative re-examination of the same underlying data. The association with autism appears to be non-existent once the studies are weighted correctly, and the link to other disorders is not robust. While a faint, method-dependent signal for attention-deficit/hyperactivity disorder remains, it is not strong enough to support the original conclusion of a likely relationship. The findings indicate that the observed connections in the literature are likely the result of how the studies were combined and the limitations of observational research, rather than a definitive biological harm. The authors caution that while the evidence for harm is weak, the decision to use medication during pregnancy remains complex, involving a balance of risks that must be weighed carefully by patients and doctors.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.