The Likelihood Ratio Wall: Structural Limits on Accurate Risk Assessment for Rare Violence
This paper argues that pretrial risk assessment tools face a fundamental "Likelihood Ratio Wall" where the rarity of violent re-offenses and structural biases like over-policing mathematically prevent high-confidence individualized predictions, rendering current instruments prone to high false-positive rates that cannot be fixed by recalibration.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a judge trying to decide whether to keep a defendant in jail before their trial. You have a new "Risk Tool" that scans their history and gives them a label: "High Risk for Violence."
The tool is supposed to tell you: "If we let this person go, there is a good chance they will commit a violent crime."
This paper argues that, mathematically, this tool is almost always lying to you when it says someone is "High Risk" for rare violent crimes. Even if the tool looks impressive on paper, the reality is that most people it flags as "dangerous" would actually have stayed peaceful if released.
Here is the breakdown of why this happens, using simple analogies.
1. The "Needle in a Haystack" Problem (The Base Rate)
Violent re-arrests are very rare. In most places, only about 2 to 5 out of every 100 released defendants commit a new violent crime.
Imagine you are a security guard at a massive stadium with 10,000 people. You are looking for one person who is secretly carrying a weapon (the "violent offender").
- The "Risk Tool" is a metal detector.
- The problem is that the metal detector is not perfect. It sometimes beeps for harmless things like belt buckles or keys (false alarms).
Because the "bad guy" is so rare (1 in 10,000), even a very good metal detector will ring for innocent people far more often than it rings for the actual threat. The paper calls this the Likelihood Ratio Wall.
2. The Wall You Can't Climb
The authors prove a mathematical rule: To be right 50% of the time (meaning, if the tool says "High Risk," there is a coin-flip chance the person is actually dangerous), the tool needs to be incredibly good at spotting the needle.
- The Reality: Current tools are like a metal detector that beeps for belt buckles 10 times for every actual gun it finds.
- The Requirement: To be right half the time, the tool would need to be 30 to 50 times better at spotting the gun than the belt buckle.
The Analogy: Imagine trying to find a specific grain of sand on a beach. If your "sand finder" picks up 100 grains of sand for every 1 grain of gold, and gold is incredibly rare, you will end up holding a bucket full of sand thinking you found gold. The tool isn't "broken"; it's just that the gold is too rare for the tool to work well.
The paper shows that current tools are nowhere near good enough to clear this wall. When they flag someone as "High Risk," they are usually wrong.
3. Why "Tweaking" the Tool Doesn't Help (Recalibration)
You might think, "Can't we just adjust the tool's settings to make it more accurate?"
The authors say no. They prove that changing the numbers (recalibration) is like repainting the metal detector. You can make it look fancier or change the volume of the beep, but you cannot change the fact that it still beeps for belt buckles.
If the underlying data (the history of arrests, prior charges, etc.) doesn't contain a strong enough signal to separate the violent from the non-violent, no amount of math tricks can create that signal out of thin air. The "Wall" is a structural limit, not a software bug.
4. The "Surveillance Ceiling" (Why Some Groups Get Flagged More Often)
The paper also points out a cruel irony: Over-policing makes the tool worse for certain groups.
Imagine two neighborhoods, Neighborhood A and Neighborhood B.
- Neighborhood A is policed normally.
- Neighborhood B is over-policed. The police are there constantly, checking everyone's pockets and records.
Even if people in Neighborhood B are no more violent than people in Neighborhood A, they accumulate more "risk markers" (like prior arrests or failed court appearances) simply because they are watched more closely.
The Analogy:
- In Neighborhood A, a person might get a "risk marker" only if they actually do something wrong.
- In Neighborhood B, a person gets a "risk marker" just for being in the wrong place at the wrong time, or for a minor infraction that police in Neighborhood A would have ignored.
The paper proves that because Neighborhood B has more "false markers" (innocent people flagged by the system), the tool becomes less accurate for them. The tool hits a "Surveillance Ceiling" where it simply cannot be precise, no matter how smart the algorithm is. It ends up flagging innocent people in Neighborhood B at a much higher rate, making the "High Risk" label even less trustworthy for them.
5. The "Number Needed to Detain" (The Human Cost)
The authors translate these math problems into a real-world cost: How many people do we have to lock up to stop one violent crime?
- If the tool is right 11% of the time (which is what the math says happens at current performance levels), it means you have to detain 9 people to prevent 1 violent crime.
- This implies that 8 of those 9 people were innocent of future violence and were deprived of their liberty unnecessarily.
The Bottom Line
The paper concludes that for rare violent crimes, the data we currently have (criminal records, arrest history) is not strong enough to support high-confidence decisions about locking people up.
When a judge sees a "High Risk for Violence" label, they are being told the person is likely to be dangerous. But the math says that label is actually wrong about 8 or 9 times out of 10.
The Recommendation:
The authors suggest that if we keep using these tools, we must be honest about the uncertainty. Instead of just saying "High Risk," the reports should say:
"This person is flagged as high risk, but based on current data, there is only an 11% chance this flag is correct. To prevent one violent crime, we would likely detain 9 people who would not have committed a crime."
The paper argues that until we have data that is 30 to 50 times better than what we have now, we should stop treating these "High Risk" labels as proof that someone is dangerous.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.