State-dependent non-identifiability of the reproduction number under adaptive behavior: an empirical characterization from COVID-19 mobility
This paper demonstrates that the basic reproduction number () is fundamentally non-identifiable from epidemic trajectories alone because it conflates pathogen biology with adaptive human behavior, yet empirical analysis of US COVID-19 mobility data reveals that while this behavioral component is structurally characterizable as state-dependent, its practical impact is modest due to a negative correlation between risk responsiveness and behavioral saturation across jurisdictions.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Imagine you are trying to predict how a crowd of people will move through a maze. In the world of disease modeling, scientists use a special number called the basic reproduction number, or (pronounced "R-naught"). Think of as a scorecard for a virus: it tells you how many new people one sick person will infect if everyone else in the crowd is healthy and has no idea a virus is coming. For a long time, scientists treated this number like a fixed property of the virus itself, similar to how a car's top speed is a fixed feature of its engine.
But here's the catch: people aren't cars. When a virus starts spreading, humans notice. They get scared, they stay home, they wear masks, and they stop hugging. This is called adaptive behavior. The big question for scientists has been: If people change their behavior because of the virus, does that change the virus's "score"? Does the number we calculate today still work if the virus changes tomorrow? This paper dives into that exact puzzle, using the massive amount of data from the COVID-19 pandemic to see if we can actually separate the virus's biology from our own human reactions.
The Great "Who Did It?" Mystery
Imagine you are watching a video of a car slowing down as it approaches a stop sign. You want to know: Did the driver hit the brakes because they saw the sign, or did the car just run out of gas? In the world of epidemics, the "car slowing down" is the number of new infections dropping. The "brakes" are people staying home (behavior), and the "gas" is the virus's natural ability to spread (biology).
For years, scientists tried to figure out how much of the slowdown was the virus losing steam and how much was people hitting the brakes. But they were stuck. They could see the car slowing down, but they couldn't see the driver's foot on the pedal. They had to guess.
This paper, however, had a superpower: Google Mobility Reports. During the pandemic, these reports tracked exactly how much people were moving around in real-time. It was like having a camera inside the car showing the driver's foot. The author, Fabio Sanchez, used this data to finally watch the "foot" move.
The Big Discovery: The "Magic Trick" of the Data
The paper's main finding is a bit of a mind-bender. It turns out that if you only look at the history of the epidemic (the "factual" path), you cannot tell the difference between two very different stories.
Story A: The virus is super strong, but people are terrified and staying home, which is why cases dropped.
Story B: The virus is weak, and people didn't change their behavior at all, which is why cases dropped.
Both stories fit the data perfectly. They produce the exact same curve of infections. This means that the "true" is not identified. You can't pin it down just by looking at the past. The number we usually calculate (the "apparent ") is actually a mix of the virus's strength and our fear, and the data can't tell us which part is which.
The paper calls this an "observational equivalence class." Imagine a magician pulling a rabbit out of a hat. If you only see the rabbit, you don't know if the magician actually had a rabbit, or if they used a trick. The data is the rabbit; the trick is the hidden mix of biology and behavior.
What Happens When We Change the Rules?
So, if we can't tell the difference in the past, does it matter? The paper says yes, but only if we try to predict the future.
The author ran simulations to see what would happen if the virus changed—say, if it became 40% more contagious (a "counterfactual" scenario).
- If we assume people don't change their behavior (the "exogenous" view), the model predicts a huge, terrifying wave of infections.
- If we assume people react to the new danger (the "endogenous" view), they stay home even more, and the wave is much smaller.
The gap between these two predictions is what the paper calls "what deletes." It's the part of the future that the standard number hides. The paper found that this hidden gap is state-dependent. It's biggest when the epidemic is at a "medium" level of severity.
- If the epidemic is tiny, people don't react, so the gap is small.
- If the epidemic is already a total disaster, people are already staying home as much as humanly possible (they are "saturated"), so they can't react any more, and the gap is small again.
- The gap is biggest in the middle, where people are scared enough to react but haven't hit the "floor" of staying home yet.
The Surprising Twist: Why the Gap Was Small in Reality
You might think, "Okay, so the gap is real, but how big is it in the real world?" The author checked this across all 50 US states and Washington D.C.
Here is the surprising part: The gap was actually quite small in practice.
Why? Because of a weird coincidence in human behavior. The paper found that the two things needed to make the gap huge—high sensitivity to risk (people reacting strongly) and high "headroom" (people having room to move less)—never happened at the same time.
- In places where people were very sensitive to the risk (they reacted strongly), they had already reduced their movement to the absolute minimum. They had no "headroom" left to react further.
- In places where people still had "headroom" (they were still moving around a lot), they simply didn't react much to the risk.
It's like a thermostat that is either already turned all the way down, or it's broken and won't turn down at all. You never get the situation where the thermostat is sensitive and has room to go lower. Because these two conditions didn't overlap, the "hidden" error in the number turned out to be modest for the first wave of COVID-19.
The Takeaway for the Future
So, what does this mean for us?
- Describing the past is safe: If you just want to say how bad the virus was during the time we watched it, the standard number is fine.
- Predicting the future is tricky: If you want to know what happens if a new, stronger variant shows up, the standard is incomplete. It assumes people won't change their behavior, which is rarely true.
- The "Peak" matters most: The paper suggests that while the total number of people who get sick (the "attack rate") might not change much based on behavior, the height and timing of the peak (how many people get sick at once) changes a lot. This is crucial for hospitals. If people react to a new wave, the peak might be lower and later, giving hospitals a chance to cope. If they don't react, the peak could be a crushing wall.
In short, the paper teaches us that the virus's "score" isn't just a number written in stone; it's a conversation between the virus and our fear. We can't always tell who is speaking louder just by looking at the history, but if the conversation changes, the outcome changes too. And luckily for the first wave of COVID-19, the conversation didn't get too chaotic because people either had nothing left to give or didn't have much to give in the first place.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.