Denoising Models Develop Human-Like Perceptual Illusion Representations Across Architectures
This paper demonstrates that denoising models across various architectures develop internal "perceptual phantom" representations of human brightness illusions that correlate with psychophysical models and causally influence internal processing, yet remain undetectable in the final output.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are looking at a painting where a gray square sits on a white background, and another identical gray square sits on a black background. Even though the two squares are exactly the same color, your brain tricks you into seeing the one on the black background as much brighter. This is a visual illusion, a glitch in our perception that actually reveals how our brains are wired to make sense of the world. For a long time, scientists wondered if the artificial brains we build—called deep neural networks—could be tricked the same way. They found that when these AI models try to clean up noisy photos or generate new images, they often make the same "mistakes" as humans, seeing brightness where there is none. But here is the big mystery: Is the AI actually seeing the illusion in its own mind, or is it just faking it by producing the right picture at the very end? It's like asking if a magician is actually holding a rabbit in their pocket, or if they just happen to pull a rabbit out of a hat that looks like it has one.
This paper dives deep inside the "mind" of these AI models to find out. The researchers tested a bunch of different AI models designed to clean up images (called denoising models) and watched what happened inside their layers as they processed these tricky optical illusions. They discovered something fascinating and slightly spooky: the AI models do develop an internal representation of the illusion, but it's a ghost. The signal that tells the AI "this looks bright" gets created, it travels through the network, and it even influences how the AI thinks about the image. However, by the time the AI finishes its job and spits out the final picture, that ghost signal has completely vanished. The AI is haunted by the illusion internally, but it never lets it show up in the final result. The authors call these invisible internal signals "perceptual phantoms."
The Ghost in the Machine
To understand how the researchers found these ghosts, imagine the AI model as a massive, multi-story factory. Raw, noisy images enter at the bottom, and clean, perfect images come out the top. In between, there are dozens of floors (layers) where the image gets processed, transformed, and refined. The researchers didn't just look at the final product; they installed cameras on every single floor to see what the workers (the AI's neurons) were doing.
They used a special set of images known for creating visual tricks, like the "GVIL" dataset, where two identical colored patches look different because of their surroundings. They fed these images into the AI and watched the activity levels in different parts of the factory. What they found was that the illusion didn't show up everywhere. Instead, the "ghost" signal was strongest in a specific, narrow hallway in the middle of the factory, known as the "bottleneck." This is where the AI compresses all the information it has gathered.
Here is the twist: The researchers found that this internal signal is very real. When they measured how much the AI's activity changed between the two illusion patches, the difference was huge—about 0.66 on a scale where 0.5 is considered a "medium" effect. This means the AI was definitely processing the illusion differently than a normal image. In fact, the strength of this internal signal matched human perception so well that when they compared the AI's activity to a mathematical model of human vision (called FLODOG), the results lined up with a correlation of 0.78. That's a very strong match, suggesting the AI is thinking about brightness in a way that is surprisingly similar to how a human brain does.
The Phantom Property
But then, the researchers asked the million-dollar question: If the AI is thinking about the illusion so strongly, why doesn't the final picture look different? To answer this, they played a game of "spot the difference" with the factory. They took the internal "ghost" signals from the illusion images and tried to inject them into the processing of a normal image. They expected the final picture to shift and show the illusion.
It didn't.
No matter which type of AI model they used—whether it was a standard "U-Net" factory or a newer "Transformer" style factory—the injected ghost signals disappeared before they could reach the exit. The final images looked exactly the same as if the ghost had never been there. The researchers call this a "perceptual phantom": a representation that is active and doing work inside the machine but is completely invisible to the outside world.
To prove this wasn't just a fluke, they tried to break the factory. They found the specific "channels" (the wires carrying the ghost signal) and cut them out (a process called ablation). When they cut these illusion-sensitive wires, the internal signal dropped by nearly 44%. This proved that these specific wires were indeed carrying the illusion. But here is the kicker: when they cut these wires, the final image barely changed at all. In fact, cutting these "illusion" wires actually made the final picture more stable than cutting random wires. It's as if the factory has a self-correcting mechanism that catches the ghost and quietly deletes it before it can mess up the final product.
Why the Ghost Exists
The researchers also wanted to know: Is this a quirk of the specific AI architecture, or is it something about how these models are trained? They compared the "denoising" models (which are trained to clean up noise) with "discriminative" models (which are trained just to recognize what an image is, like a cat or a dog).
The results were clear: The ghost only appears in the denoising models. The discriminative models, even if they have the exact same factory layout, didn't show the strong illusion signal. This suggests that the act of trying to reconstruct or clean an image is what forces the AI to develop these human-like illusions. It's as if the training process teaches the AI to pay attention to context and surroundings, just like humans do, but then the AI's own internal rules decide that this context shouldn't change the final output.
The study also looked at different types of AI, including a massive model called DiT-XL/2. Even in this giant model, the illusion signal peaked in the middle layers but vanished by the end. The researchers found that the AI's self-correcting dynamics—its way of re-checking its own work—seem to be the reason the ghost disappears. It's like the AI has a strict editor that says, "No, that brightness doesn't belong here," and scrubs it out before the story is published.
The Takeaway
So, what does this all mean? It turns out that AI models are not just mimicking human behavior; they are actually experiencing a version of human perception internally. They build up a rich, detailed understanding of visual tricks, complete with the same biases we have. But, unlike us, they seem to have a filter that keeps these biases from changing their final output.
This discovery is a bit like finding out that a robot has a vivid dream about flying, but when it wakes up and builds a house, it builds a perfectly normal, grounded house. The dream was real to the robot, but it didn't affect the house. The researchers suggest that this "perceptual phantom" phenomenon might be common in many AI systems. It means that if we only look at what an AI produces, we might be missing a huge part of what it actually knows and how it thinks. The AI might be seeing things we can't see, even if it never shows them to us.
The study didn't just suggest this; they measured it, cut the wires to prove it, and injected signals to test it. They found that across nine different models, the illusion signal was strong, causal, and then completely invisible at the output. It's a reminder that in the world of AI, what you see is not always what you get, and sometimes the most interesting things are the ghosts that stay hidden inside the machine.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.