Compliant But Unsatisfactory: The Gap Between Auditing Standards and Practices for Probabilistic Genotyping Software
This paper critiques the ASB 018 auditing standard for probabilistic genotyping software, arguing that its vague design creates a gap between compliance and effectiveness, allowing audits to appear satisfactory while failing to establish necessary restrictions on software use in the U.S. criminal legal system.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Picture: The "Seal of Approval" That Isn't Sealed
Imagine a world where a new, high-tech magic box is used to solve crimes. This box takes DNA samples from a crime scene and tells a judge, "There is a 99% chance this suspect was at the scene."
Because this box is so powerful, the legal system needs to make sure it works correctly before letting it testify in court. So, a group of experts created a Rulebook (called ASB 018) to tell laboratories how to test these boxes. The idea was: "If a lab follows these rules, the box is safe to use."
The Problem: The researchers in this paper found that laboratories can follow the Rulebook perfectly on paper, but the box might still be broken, dangerous, or misleading. They call this "Compliant but Unsatisfactory." It's like a restaurant passing a health inspection because they have a fire extinguisher, even though they are serving rotten food.
The Five Ways the System Fails (The "Loopholes")
The researchers looked at the Rulebook and five real-life reports from labs. They found five major gaps where the labs "cheated" the spirit of the rules while following the letter of the law.
1. The "Robot vs. Human" Gap (Audit Scope)
- The Goal: The Rulebook says you must test the whole system, including the human pressing the buttons.
- The Reality: The labs only tested the machine. They pretended the human operator was perfect.
- The Analogy: Imagine testing a self-driving car. The Rulebook says, "Test the car and the driver." But the lab only tested the car on a perfect, empty track. They didn't test what happens when a confused human driver tries to take the wheel or inputs the wrong destination. In real life, humans make mistakes, and if the machine doesn't account for that, it can give a wrong answer.
2. The "Fake Menu" Gap (Inputs)
- The Goal: You should test the machine with food that looks like what you actually serve (real crime samples).
- The Reality: The labs tested the machine with "convenient" samples that were easy to make, not samples that looked like messy, real-world crime scenes.
- The Analogy: A chef wants to prove their soup is good. The Rulebook says, "Test it with ingredients you actually use." But the chef only tested it with perfect, store-bought tomatoes. They didn't test it with the bruised, weird-shaped tomatoes they actually get from the farm. The soup might taste great with perfect tomatoes but turn to sludge with real ones.
3. The "Vague Goal" Gap (Standards)
- The Goal: Before testing, you must define exactly what "success" looks like.
- The Reality: The labs didn't define success. They just said, "The results looked good."
- The Analogy: Imagine a teacher grading a student. The Rulebook says, "Define what an 'A' looks like." But the teacher just says, "I think this essay is an A because it feels right." Without a clear rubric (e.g., "Must have 5 paragraphs"), anyone can claim they passed. The labs used words like "high" or "intuitive" instead of hard numbers, so they could claim the machine worked even when it didn't.
4. The "Small Sample" Gap (Performance)
- The Goal: You need to test the machine on a huge variety of difficult cases to be sure it works.
- The Reality: The labs tested it on a tiny, easy group of samples.
- The Analogy: To prove a parachute works, you shouldn't just jump off a 1-foot step. You need to jump from a plane. The labs only jumped off the step (easy samples) and claimed, "See? It worked!" They didn't test the parachute on the high jumps (complex, messy DNA samples) where it might actually fail.
5. The "No Red Lines" Gap (Judgment)
- The Goal: If the machine makes a mistake, you must draw a line in the sand and say, "We will never use the machine on cases like this again."
- The Reality: The labs saw the mistakes but didn't draw the line. They just said, "The machine made a weird noise, but a smart human can fix it."
- The Analogy: A pilot's plane has a warning light that flashes when the engine is failing. The Rulebook says, "If the light flashes, you must stop flying." The lab said, "The light flashed, but we think the pilot is smart enough to ignore it and keep flying." They didn't ban the bad cases; they just hoped the humans would catch the error.
Why Did This Happen? (The Design Flaws)
The paper argues that the Rulebook itself was poorly designed. It wasn't malicious, but it was too loose.
- Too Much "Maybe": The Rulebook used words like "consider," "address," and "evaluate." These are soft words. You can "consider" a problem by thinking about it for one second and then ignoring it. The labs did the bare minimum to check the box.
- Siloed Rules: The Rulebook treated different parts of the test as separate items. It said, "Test the machine" and "Test the human" as two different tasks. The labs did them separately, so they never saw how the human mistakes broke the machine.
- No "Who": The Rulebook didn't clearly say who had to do the testing. Sometimes, the company that built the machine helped the lab test the machine. It's like asking the car manufacturer to test their own brakes.
The Takeaway
The researchers aren't saying the DNA software is bad. They are saying the safety net is full of holes.
If we want these tools to be fair and safe in court, we can't just have a checklist that says "Did you do it?" We need a checklist that says "Did you do it well enough to protect innocent people?"
The Solution?
- Be Specific: Don't say "Consider the human." Say "Test the human making 10 specific mistakes."
- Be Clear: Define "Success" with numbers, not feelings.
- Include Everyone: Don't just ask the labs what they want. Ask the defense lawyers, the victims, and the scientists what they need to feel safe.
In short: A rulebook that allows you to pass without actually being safe is no rulebook at all.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.