Local Intrinsic Dimension Unveils Hallucinations in Diffusion Models
This paper proposes that structural hallucinations in diffusion models stem from instabilities driven by high local intrinsic dimension (LID) on the model-induced manifold, introducing a method called Intrinsic Quenching (IQ) to deflate LID and effectively reduce these anomalies, particularly improving anatomical consistency in medical imaging.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a Diffusion Model as a master sculptor who has spent years studying thousands of statues. Their job is to create new statues from a block of marble (noise) by chipping away pieces until a perfect form emerges. Usually, they do a great job. But sometimes, they get confused and create a statue with six fingers, three eyes, or a hand that looks like a flower. In the AI world, these mistakes are called hallucinations.
This paper argues that these mistakes happen because the sculptor gets lost in a "foggy" part of their mental map. The authors propose a new way to fix this by measuring how "wobbly" the sculptor's path is and then gently smoothing it out.
Here is the breakdown of their discovery and solution:
1. The Problem: The Sculptor's "Wobbly" Path
The authors suggest that when a Diffusion Model creates a hallucination (like a six-fingered hand), it isn't just a random error. It's happening in a specific part of the model's "mental landscape" (called a manifold) where the geometry is unstable.
Think of the model's training data as a smooth, flat valley where real hands live.
- Normal Generation: The model walks smoothly along the valley floor, creating perfect hands.
- Hallucination: The model steps onto a shaky, unstable ledge. Here, the ground seems to have extra, invisible dimensions. Because the ground is wobbly, the model accidentally invents extra fingers or misaligned eyes.
The authors call this instability Local Manifold Instability (LMI). It's like trying to walk on a trampoline that is stretching in weird directions; you might end up bouncing into a shape you didn't intend.
2. The Discovery: Measuring the "Wobble"
To find these wobbly spots, the authors looked at a concept called Local Intrinsic Dimension (LID).
- The Analogy: Imagine a flat sheet of paper (2D) floating in a 3D room. If you look at a tiny spot on the paper, it feels flat (2D). But if the paper is crumpled or if the model is hallucinating, that spot might suddenly feel like it has extra, unnecessary directions to move in (like 3D or 4D).
- The Finding: The paper shows that when the model is about to make a mistake (like adding an extra finger), the "dimension" of the space it is walking through suddenly inflates. It's as if the model is trying to walk in a direction that shouldn't exist.
They tested this by creating a filter that looks for these "inflated dimensions." They found that this filter is actually better at spotting bad hands than previous methods that looked at how the image changed over time.
3. The Solution: "Intrinsic Quenching" (IQ)
Once they knew where the mistakes happened (in the wobbly, high-dimensional spots), they created a fix called Intrinsic Quenching (IQ).
- The Metaphor: Imagine the model is a car driving down a road. When it hits a bumpy, unstable patch (a hallucination), the car starts swerving.
- The Fix: IQ acts like a smart suspension system. As the car approaches the bumpy patch, it detects the instability. Instead of letting the car swerve, it gently pushes the car back onto the smooth, stable part of the road. It "deflates" the extra, unnecessary dimensions, forcing the model to stick to the logical, real-world rules (like "hands only have five fingers").
They call it "Quenching" because it's like cooling down a hot, unstable metal to make it solid and stable again.
4. Does it Work?
The authors tested this on many different tasks:
- Drawing Hands: They showed that IQ drastically reduced the number of six-fingered hands compared to other methods.
- Medical Imaging: They tested it on reconstructing brain CT scans (images used by doctors). In this case, a hallucination could look like a fake tumor or a missing brain part, which is dangerous. IQ helped ensure the reconstructed images were anatomically correct without inventing fake diseases.
- Faces and Animals: It also improved the quality of generated faces and animal pictures, making them look more natural.
Summary
The paper claims that AI hallucinations happen because the model gets lost in "wobbly" parts of its own math. By measuring this wobble (using Local Intrinsic Dimension) and gently pushing the model back to stable ground (using Intrinsic Quenching), they can stop the AI from inventing impossible things like six-fingered hands or fake tumors. It's a way of teaching the AI to stay on the "solid ground" of reality.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.