← Latest papers
📊 statistics

When the record makes the result: exposure-linked recording of a continuous outcome can manufacture a spurious treatment effect in real- world evidence

This study demonstrates that exposure-linked differential recording of continuous outcomes, such as visually estimated blood loss, can fabricate large, robust-looking treatment effects in real-world evidence through systematic measurement error rather than true clinical differences, necessitating specific diagnostic checks and reliance on objective measures to avoid spurious conclusions.

Original authors: Chia-Wei Chen, Shih-Peng Mao

Published 2026-08-06
📖 7 min read🧠 Deep dive

Original authors: Chia-Wei Chen, Shih-Peng Mao

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Hidden Trap in the Data: When "Guessing" Creates Fake Science

Imagine you are a detective trying to solve a mystery using a giant stack of police reports. You want to know if a new type of lock prevents burglaries better than the old kind. In the world of modern science, researchers do something very similar. Instead of police reports, they use "Real-World Evidence"—massive digital records from hospitals, like electronic medical records, to see if new treatments actually work. This is a powerful tool because it looks at real patients in real life, not just people in a controlled lab.

But there is a catch. Sometimes, the people writing the reports aren't using a ruler; they are using their eyes to guess. If a doctor has to guess how much blood a patient lost, or how big a wound is, they might write down a number that feels "right" to them, rather than the exact truth. Scientists call this "estimation." Usually, we think these guesses just add a little bit of noise, like static on a radio. But what if the static isn't random? What if the way the doctor guesses changes depending on which medicine the patient took? That is the scary question this paper asks. It explores a hidden trap where the way we write down the results can accidentally invent a miracle cure that doesn't actually exist.

The Case of the "Magic" Medicine That Wasn't

In this study, researchers looked at a massive collection of 8 years of hospital records from a single hospital in Taiwan. They were investigating a drug called carbetocin, used to help stop bleeding after a baby is born. Here is the twist: the drug wasn't covered by insurance, so patients had to pay for it themselves. This meant the doctors knew exactly who got the drug and who didn't, creating a perfect "natural experiment" to see if the drug worked.

The researchers focused on one specific outcome: how much blood the mother lost during a vaginal delivery. In many hospitals, this number isn't measured with a scale; it is visually estimated by the doctor or nurse. They look at the blood and write down a number, like "150 mL" or "200 mL."

When the researchers ran the numbers the "standard" way, the results looked incredible. The data suggested that the drug reduced blood loss by a massive 62%. The confidence in this number was so high that it looked like a slam-dunk scientific discovery. The treated group seemed to have lost almost no blood at all, while the control group lost a lot.

But the authors of this paper didn't stop there. They started playing detective, looking for clues that something was fishy.

The "Round Number" Clue

The first thing they noticed was a strange pattern in the numbers, like a fingerprint left at a crime scene. In the group that took the drug, about 60% of the recorded blood loss values were exactly 30 mL or 50 mL. In contrast, only 2.4% of the control group had these exact low numbers.

It turns out that doctors often have a habit of rounding their guesses to nice, round numbers (like 50 or 100). But here, the doctors seemed to have a secret rule: if the patient took the drug, they almost always wrote down "50 mL," even if the actual loss was higher. It was as if the doctors had a template that said, "If they took the magic pill, write '50'." This is called "digit preference" or "value heaping." The treated group wasn't just losing less blood; they were being recorded as losing less blood because of a recording habit.

The "Time Travel" Test

To be sure this wasn't just a fluke, the researchers checked if the pattern changed over time. If the drug really worked, the effect should be consistent year after year. And it was! From 2019 to 2025, the median blood loss for the drug group was pinned exactly at 50 mL every single year. Meanwhile, the control group hovered between 150 mL and 200 mL.

This is a huge red flag. Real biological effects don't usually freeze at the exact same number for seven years in a row. That kind of rigidity suggests a recording rule, not a biological miracle. The researchers also tried to fix the data using standard statistical tricks (like adjusting for the year or the patient's health), but the "62% reduction" stayed exactly the same. This proved that the problem wasn't that the groups were different; the problem was that the way the results were written down was different.

The "Threshold" Reality Check

The most damning evidence came when they looked at the numbers through a different lens. They asked: "How many women actually had a dangerous amount of bleeding (Postpartum Hemorrhage), defined as losing 500 mL or more?"

When they looked at this binary "yes or no" question, the magic disappeared.

  • Drug group: 3.2% had heavy bleeding.
  • Control group: 4.2% had heavy bleeding.
  • Result: There was no significant difference.

The "62% reduction" only existed in the messy, estimated numbers. When they looked at the actual threshold for danger, the drug didn't seem to help at all. In fact, at the very extreme end (losing 1,000 mL or more), the drug group actually had slightly more cases than the control group, which is the opposite of what a miracle drug should do.

The Simulation: Proving the Illusion

To prove that this recording habit alone could create such a fake result, the researchers built a computer simulation. They created a fake world where the drug did absolutely nothing (the true effect was zero). They programmed the computer to record the "treated" group's blood loss with the same bias: pinning low values to 30 or 50 mL, while leaving the control group more natural.

The result? The computer generated a fake "62% reduction" that looked just as impressive and statistically significant as the real data. This confirmed that differential recording alone was enough to manufacture a massive, convincing treatment effect out of thin air.

The Takeaway: Trust the Ruler, Not the Guess

This paper isn't saying the drug definitely doesn't work; it's saying that the "62% reduction" found in the estimated blood loss numbers is likely a lie created by how doctors wrote down the data.

The authors conclude that when we use real-world data, we have to be careful about outcomes that are just "guessed" or "estimated." If the way we record the data changes based on who got the treatment, we can accidentally invent a cure.

The lesson for the future:

  1. Check the "Rounding": Before believing a big result from estimated data, check if one group is rounding their numbers to 50 or 100 while the other doesn't.
  2. Look at the "Real" Thresholds: Don't just trust the average number. Check if the treatment actually helps people cross the line into danger (like the 500 mL bleeding threshold).
  3. Use Objective Measures: Whenever possible, use measurements that machines make (like a hemoglobin blood test) rather than human guesses.

In short, a large, precise-looking effect in a study based on human guesses might just be a reflection of the doctor's pen, not the patient's body. The "magic" was in the notebook, not in the medicine.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →