Beyond Accuracy: Evaluating Posterior Fidelity of Diffusion Inverse Solvers
This paper addresses the limitation of existing Diffusion Inverse Solvers benchmarks that focus solely on reconstruction accuracy by introducing a systematic study of posterior fidelity and proposing a ground-truth-free metric, score-based Kernel Stein Discrepancy (score-KSD), to evaluate how well generated samples capture the target posterior distribution in both controlled simulations and real-world inverse problems.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Problem: The "Accuracy Trap"
Imagine you are trying to solve a mystery, but the clues you have are blurry and incomplete. In science and engineering, this is called an inverse problem. You see the result (like a blurry photo or a noisy signal) and you have to guess what caused it.
Recently, scientists started using a new type of AI called a Diffusion Inverse Solver (DIS). Think of these solvers as a team of detectives. When they look at the blurry clues, they don't just guess one single answer; they generate a whole cloud of possible answers to show how uncertain they are.
The Trap:
Until now, we've only been judging these detectives by how close their single best guess is to the real answer. We use a score called "Accuracy" (like PSNR).
- The Paper's Claim: This is a trap. A detective might get a high accuracy score just by getting lucky and guessing a point very close to the truth, even if their entire cloud of guesses is wrong. They might miss the real solution entirely or only guess one tiny possibility, ignoring other valid ones.
- The Analogy: Imagine a weather forecaster.
- Forecaster A predicts "It will rain exactly at 2:00 PM." It rains at 2:00 PM. High Accuracy. But they missed that it might also rain at 3:00 PM or 4:00 PM.
- Forecaster B predicts "It will rain anytime between 2:00 PM and 5:00 PM." It rains at 2:00 PM. Lower Accuracy (because their single point guess was less precise), but their Uncertainty was perfect. They captured the whole truth.
- The paper argues we need to stop only praising Forecaster A and start checking if Forecaster B's "cloud of guesses" actually matches the real weather patterns.
The Solution: A New "Truth Detector" (Score-KSD)
The problem is: How do you check if a detective's cloud of guesses is right if you don't know the real answer? In real-world science (like medical scans), we often don't have the "ground truth" to compare against.
The authors created a new tool called Score-KSD.
How it works (The Metaphor):
Imagine the "Truth" is a hidden landscape with hills and valleys. The "Score" is like a compass that always points uphill toward the most likely places.
- The Compass: The paper shows that even if we don't know the exact map (the true answer), we can build a compass using the physics of the problem and the AI's training. This compass points toward the "right" areas.
- The Test: We take the cloud of guesses generated by the AI and check: "Do these guesses follow the compass?"
- If the guesses are scattered randomly or stuck in the wrong valley, the compass will spin wildly. The Score-KSD number will be high (bad).
- If the guesses cluster perfectly where the compass points, the Score-KSD number will be low (good).
The Key Innovation: This tool doesn't need to see the "Ground Truth" (the real answer). It only needs to check if the AI's guesses are consistent with the rules of the universe (the physics) and the AI's own training.
What They Found
The authors tested this new tool on many different problems, from fixing blurry photos to reconstructing MRI scans and CT scans.
- Accuracy Truth: They found that the AI methods with the "sharpest" single images (highest accuracy) often had terrible "clouds of guesses." They were confident but wrong about the uncertainty.
- The "Trap" Confirmed: Some methods looked great on standard charts but failed the Score-KSD test. They were producing "off-posterior" samples—guesses that looked good but were statistically impossible given the physics of the problem.
- A New Standard: The Score-KSD metric successfully identified which AI methods were actually capturing the full range of possibilities (the true uncertainty) and which ones were collapsing into a single, potentially misleading guess.
Summary
- Old Way: "Look how close this single guess is to the real thing!" (Misses the bigger picture of uncertainty).
- New Way: "Does this whole group of guesses follow the rules of physics and probability, even if we don't know the real answer?"
- Result: The paper proves that being "accurate" doesn't mean you understand the uncertainty. Their new tool, Score-KSD, is a way to check if an AI is being honest about what it knows and doesn't know, without needing a cheat sheet (ground truth) to compare against.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.