The semantic geometry of action: convergence across foundation models, human behaviour and the brain
This study demonstrates that foundation models organize naturalistic actions within a semantic geometry that converges with human behavioral judgments and cortical representations, outperforming traditional video encoders in predicting both human similarity ratings and brain activity.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Human beings are masters at reading the world in motion. When we watch someone else, we do not just see a body moving through space; we instantly understand what they are trying to do, why they are doing it, and what their next move might be. This ability to interpret dynamic action is a cornerstone of human intelligence, allowing us to navigate complex social situations and cooperate with others. For decades, scientists have tried to build artificial intelligence that can do the same. While modern computer programs can now identify actions with high accuracy—telling the difference between a person running and a person walking—researchers have long wondered if these machines actually understand the meaning behind the movement in the same way humans do. Do they organize these actions in their internal "minds" based on the same logic we use, or are they simply matching patterns of pixels without grasping the underlying concept?
A team of researchers from the Institute of Automation in China has now taken a significant step toward answering this question. They set out to map the hidden structure of how both humans and advanced artificial intelligence organize the concept of human action. By comparing the internal "geography" of action in large language models, multimodal models that can see and read, and the human brain, they discovered a surprising convergence. The study suggests that when these artificial systems are trained on vast amounts of data, they spontaneously develop a way of thinking about human actions that aligns closely with human perception and even mirrors the way the human brain processes these same movements.
To explore this, the researchers treated artificial intelligence models like participants in a psychological experiment. They presented these models with thousands of short video clips and text descriptions of people doing various things, from cooking and gardening to fighting and dancing. The models were asked to perform a simple but revealing task: given three different actions, they had to identify which one was the "odd one out." For example, if shown a video of someone skiing, someone drawing, and someone climbing a rock wall, the model had to decide which action was most different from the other two. By analyzing millions of these choices, the researchers could reconstruct the invisible mental map the models were using to sort these actions. They did the same with a large group of human participants, creating a baseline of how people naturally group and distinguish different behaviors.
The results revealed that both text-only models, which only read descriptions of actions, and multimodal models, which watched the actual videos, developed internal maps that were strikingly similar to the human map. These maps were not random; they organized actions into low-dimensional spaces where similar behaviors sat close together and different ones were far apart. Crucially, these artificial maps predicted human choices better than traditional video analysis tools that rely on simple visual features. The study found that the models were not just memorizing the videos; they were capturing the abstract goals and social contexts that humans use to make sense of movement. For instance, the models learned to group "socializing" actions together and separate them from "solitary" tasks, just as people do.
To prove that these internal maps were truly meaningful and not just statistical tricks, the researchers tested them in two rigorous ways. First, they used the models' internal maps to guide a video generation system. By adjusting a single number in the model's internal representation, they could smoothly change a video's content. If they increased the value for a "nature-related" dimension, an indoor scene would transform into an outdoor one; if they boosted a "sports" dimension, the activity would shift toward athletic competition. This demonstrated that the models had learned distinct, controllable concepts of action that could be manipulated to create new, coherent visual scenes.
Second, and perhaps most significantly, the researchers tested whether these artificial maps could predict activity in the human brain. They took the fixed maps learned by the models and applied them to a massive dataset of brain scans from people watching thousands of different action videos. They found that the models' internal organization successfully predicted which parts of the human brain would light up when viewing specific actions. The text-based models, which had learned their action concepts purely from language descriptions without ever seeing a video, were particularly good at predicting activity in higher-order brain regions responsible for understanding complex social and goal-directed behaviors. This suggests that the abstract concepts of action, learned purely through language, are sufficient to capture the core structure of how the human brain represents movement.
The study also highlighted interesting differences between the types of models. While the video-watching models captured fine-grained visual details, the text-only models were surprisingly effective at organizing actions into broad, meaningful categories. In some brain regions, the text-based models actually outperformed the video models in predicting human neural responses. This implies that the conceptual understanding of action—knowing what a person is trying to achieve—might be more important for human-like perception than the specific visual details of how the movement looks.
Ultimately, this research provides a powerful new way to understand the relationship between artificial and biological intelligence. It shows that when machines are exposed to enough data, they can develop a shared semantic geometry with humans, organizing the world of action in a way that is functionally similar to our own. The findings suggest that the path to truly intelligent machines may not require us to explicitly program them with human rules, but rather to let them discover these structures on their own through exposure to the richness of human experience. The study does not claim that machines have consciousness or feelings, but it does demonstrate that they can build a structural understanding of human behavior that is deeply aligned with our own, offering a promising foundation for future interactions between people and machines.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.