← Latest papers
🧬 biology

Spatial organization of faces and bodies in single neurons

This study demonstrates that inferotemporal cortex neurons encode faces and bodies not as translation-invariant templates, but through spatially localized subunits that preserve the geometric arrangement of facial features, thereby supporting both object identity and configural processing.

Original authors: Margaret Livingstone, Akshay Jagadeesh, sohrab najafian, Binxu Wang, Michael Arcaro

Published 2026-08-19
📖 5 min read🧠 Deep dive

Original authors: Margaret Livingstone, Akshay Jagadeesh, sohrab najafian, Binxu Wang, Michael Arcaro

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

The human brain is a master of recognition. It can identify a friend's face in a crowd, a familiar tree in a forest, or a specific car in a parking lot, regardless of where those objects appear in our vision or how large they seem. For decades, scientists studying the part of the brain responsible for this high-level vision, known as the inferotemporal cortex, believed that the neurons here acted like perfect, position-invariant templates. The prevailing idea was that once a neuron learned what a face looked like, it would fire whenever that face appeared, whether it was in the center of your vision, off to the side, or upside down. This theory suggested that the brain's final stage of visual processing had smoothed out all the messy details of location to focus purely on identity.

However, recognizing a face is a delicate business. It relies not just on seeing the parts, like eyes or a mouth, but on understanding exactly how those parts are arranged relative to one another. If you scramble the features of a face, it becomes unrecognizable, even if every single part is still there. This tension between needing to know "what" an object is and "where" it is located has long puzzled researchers. If the brain truly ignores location to recognize objects, how does it preserve the precise geometry that makes a face a face? A new study from Harvard Medical School and the University of Pennsylvania challenges the old view, suggesting that the brain's face-recognition cells are not blind to location, but are instead deeply sensitive to it, preserving the spatial map of the face within the neuron itself.

To investigate this, the researchers turned to macaque monkeys, whose visual systems are remarkably similar to our own. They implanted tiny electrode arrays into the monkeys' brains, specifically targeting regions known as "face patches," where neurons are highly selective for faces. The team then conducted a series of experiments to see how these neurons reacted when the monkeys looked at images. In the first experiment, they showed the monkeys pictures of upright faces and the exact same pictures turned upside down. They flashed these images at different spots across the monkeys' field of view, moving them around in a grid pattern. If the neurons were truly position-invariant, the center of their sensitivity should have remained the same regardless of whether the face was upright or inverted.

The results were striking. When the researchers mapped the sensitivity of these neurons, they found that the center of the neuron's attention shifted depending on the orientation of the face. Specifically, when the monkeys looked at an upside-down face, the neurons seemed to be looking higher up in the visual field than when the monkey looked at an upright face. This shift was not random; it was a consistent, measurable movement. The researchers realized this was not because the neurons were confused by the inversion, but because they were zeroing in on a specific feature: the eyes. In an upright face, the eyes are at the top. In an upside-down face, the eyes are at the bottom. The neurons were firing most strongly when the eyes landed in a specific spot relative to the neuron's own position. When the face was flipped, the eyes moved, and the neuron's "sweet spot" appeared to move with them. This suggested that the neurons were not seeing the whole face as a single, unchanging blob, but were instead tracking the specific location of its parts.

To test this further, the team showed the monkeys isolated parts of the face and body, such as just an eye, just a mouth, a whole face, or a whole body, flashing them at various locations. They found a clear, orderly pattern in how the neurons responded. Neurons that preferred eyes tended to be most active when the eyes were in the upper part of the visual field. Neurons that preferred mouths or bodies were more active when those features were lower down. This created a kind of internal map within the neuron's receptive field, where the top was reserved for eyes and the bottom for mouths or bodies. It was as if the neuron contained a miniature, spatially organized model of a face, rather than just a generic detector for "face-ness."

The researchers then took this a step further by showing the monkeys thousands of natural photographs, ranging from landscapes to portraits, and used a sophisticated computer model to decode what the neurons were actually seeing. This model broke down the neuron's response into smaller, distinct components. Instead of finding one single reason why a neuron fired, the model revealed that each neuron was actually a combination of several smaller, specialized detectors working together. One part of the neuron might be tuned to look for eyes in the upper visual field, while another part looked for hands or bodies in the lower field. These components were distinct and spatially separated, confirming that the neuron's sensitivity was not a uniform blanket but a mosaic of specific feature-location pairs.

These findings fundamentally change how we understand high-level vision. The study shows that the brain does not discard spatial information to recognize complex objects. Instead, it preserves the geometry of what it sees. The neurons in the face patches are not translation-invariant templates that ignore where an object is; they are spatially localized feature detectors that maintain the relationship between parts. This means that the brain's ability to recognize a face holistically likely emerges from the combined activity of many neurons, each holding a piece of the spatial puzzle. The brain does not need to throw away the "where" to know the "what"; it keeps both, weaving them together to create the seamless perception we experience every day. This work suggests that the brain's strategy for recognizing complex objects is far more intricate and spatially precise than previously thought, relying on a structured, map-like organization that mirrors the very shapes it seeks to understand.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →