Parametric neural control differentiates top neural network models of primate visual cortex
This study introduces axis-aligned feature accentuation as a causal method to reveal that, despite similar accuracy in predicting natural image responses, leading deep neural network models diverge significantly in their ability to control primate visual cortex firing, with performance better predicted by input gradient spatial frequency structure than by adversarial robustness.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
The human brain is a vast, intricate machine for making sense of the world, and at the front of this machine lies the visual cortex, a region dedicated entirely to processing what we see. For decades, scientists have tried to build computer programs that mimic this biological machinery. These programs, known as deep neural networks, are trained on millions of photographs until they learn to recognize objects, faces, and scenes with a skill that rivals our own. When researchers test these programs, they often find that the most advanced models predict how neurons in a monkey's brain will fire when shown a natural image with startling accuracy. This success led to a comforting assumption: if different computer models can all predict brain activity so well, they must all be using the same internal logic to understand the visual world. It seemed as though the path to understanding the brain had finally been found, and that all these successful models had converged on a single, correct way of seeing.
However, a new study challenges this comforting conclusion, suggesting that high accuracy in prediction does not necessarily mean the models have truly grasped the brain's logic. The researchers, working with five macaques, set out to test whether these computer models were merely guessing correctly or actually understanding the specific rules that govern how neurons respond. They created a new method to probe the models, moving beyond simple picture recognition to a more direct form of interaction. Instead of just showing the models images and checking if they guessed the right brain response, the scientists used the models to generate specific, controlled changes to an image. They took the internal settings of each model and turned them into a set of instructions for how to tweak an image to make a neuron fire more or less. This process, which the researchers call axis-aligned feature accentuation, allowed them to create thousands of unique test images designed to push the neural response in a specific direction.
The team generated over 27,500 of these custom stimuli from ten leading vision models and presented them to the monkeys in a closed-loop experiment. This setup allowed the researchers to target specific areas of the visual cortex, from the early stages where simple shapes are processed to the higher levels where complex objects are recognized. The results were surprising. Even though the ten models all performed equally well when predicting responses to standard, natural photographs, they fell apart when asked to control the neurons with their custom-made stimuli. Most of the models failed to produce the predicted changes in neural firing. Their internal maps of the visual world, while good enough for guessing, were fundamentally misaligned with the actual tuning of the brain cells. The models were not capturing the precise details of how neurons respond to the world; they were simply finding statistical shortcuts that worked for natural images but broke down under closer scrutiny.
There was one notable exception to this failure. Two of the models, which had been trained using a specific technique called adversarial training, showed a consistent advantage. These models were better at controlling the neural firing, suggesting their internal parameters were closer to the brain's own. Yet, the study found that being robust against these adversarial tricks was not the whole story. The ability to control the neurons was not strongly linked to how well a model resisted artificial attacks. Instead, the key factor was something more fundamental: the spatial frequency structure of the input gradient. In plain terms, this refers to the specific pattern of pixels that influence a model's decision. The models that worked best were those where the distribution of pixels influencing the encoding axis matched the way the brain actually processes visual information.
This research establishes a new, more rigorous way to test how well computer models align with the brain. By moving from passive prediction to active control, the scientists demonstrated that high accuracy on standard tests can hide deep flaws in how a model represents the visual world. The findings suggest that while current models are impressive, most have not yet converged on the true biological parameterization of natural image space. Only by testing whether a model can causally drive neural activity with precise, accentuated features can we know if it truly understands the visual world the way a primate brain does.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.