← Latest papers
🤖 AI

Grounded world models in biological organisms and future embodied AI

This paper argues that future embodied AI should shift from passive, language-centric training to biologically inspired, grounded world models built through active environmental interaction, intrinsic dynamics, and social engagement to achieve robust, open-ended intelligence.

Original authors: Giovanni Pezzulo, Davide Nuzzi, Marco D'Alessandro, Riccardo Proietti, Roberto Bottini, Paul Cisek

Published 2026-07-16
📖 8 min read🧠 Deep dive

Original authors: Giovanni Pezzulo, Davide Nuzzi, Marco D'Alessandro, Riccardo Proietti, Roberto Bottini, Paul Cisek

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Great Brain-Builder: Why Robots Need to Touch Before They Talk

Imagine you are trying to teach a super-smart robot how to understand the world. You have two main ways to do it. The first way is like handing the robot a billion books and a million videos, then saying, "Read everything, memorize the patterns, and guess what comes next." This is how most of today's most famous artificial intelligence works. It's like a student who has read every dictionary but has never actually held an apple, felt its weight, or tasted its sweetness. It knows the words for "apple," "heavy," and "crunchy," but it doesn't really know what an apple is.

The second way is how living things—like you, me, or even a tiny worm—learn. We don't start by reading a manual. We start by moving, bumping into things, feeling hungry, and figuring out what happens when we push a ball or pull a string. Our brains build a map of the world based on these real-life experiences first. Only later do we learn to put labels (words) on those experiences. This paper is about the difference between these two approaches. It asks a big question: If we want to build robots that can truly understand us and navigate our messy, physical world, should we keep stuffing them with books, or should we let them go out and play? The authors suggest that the "play first, talk later" method might be the secret sauce that current robots are missing.


The Paper's Big Idea: From "Passive Reader" to "Active Explorer"

This paper, written by a team of scientists from Italy and Canada, argues that the way we are currently building Artificial Intelligence (AI) is backwards compared to how nature built us.

Right now, the hottest trend in AI is "Generative AI." These are systems that learn by reading massive amounts of text and watching videos. They are incredibly good at predicting the next word in a sentence or the next frame in a video. The paper calls this a "passive training regime." Imagine a student sitting in a chair, staring at a screen, watching a thousand videos of people catching balls, but never actually throwing or catching one themselves. They might learn the statistics of how balls move, but they haven't learned the feeling of the ball hitting their hand. The authors point out that in these systems, language (words) is the foundation. The AI learns the rules of language first, and then tries to attach pictures and actions to those words later.

The paper suggests that biological organisms (like humans, dogs, and even worms) do the exact opposite. We build "grounded world models" first. This is a fancy way of saying our brains create a mental map of how the world works based on our own bodies moving through it. We learn that fire is hot because we feel the heat, not because someone told us "fire is hot." Only after we have this physical, grounded understanding do we layer language on top of it. For us, words are just labels for things we already know through experience.

Five Ways Our Brains Build Better Maps

To prove their point, the authors look at five specific "circuits" or parts of the brain that help living things build these grounded maps. They use these examples to show what is missing in today's robots.

1. The GPS and the Concept Map
Think about how you navigate your city. You don't just turn left when you see a red sign; you have a mental map. Your brain has a special system (in the hippocampus and entorhinal cortex) that acts like a GPS. It uses "grid cells" that fire in a hexagonal pattern, like a honeycomb, to help you know exactly where you are.
But here's the cool part: this same GPS system doesn't just work for physical streets. The paper explains that your brain uses this same "grid" to navigate ideas. When you think about a story, or compare two different concepts, or even figure out social relationships, your brain is essentially "walking" through a mental map. The paper suggests that in biology, the ability to navigate physical space came first, and the ability to navigate abstract ideas was built on top of it. Current AI tries to learn abstract ideas directly from text, skipping the physical "walking" part.

2. The "What Can I Do?" Detector (Affordances)
When you look at a chair, you don't just see "a wooden object with four legs." You see "something I can sit on." When you see a cup, you see "something I can grasp." The paper calls these "affordances." It's a direct link between what you see and what you can do with it.
Your brain is constantly running a competition: "Can I climb that rock? Can I grab that apple?" It's not just describing the world; it's calculating possibilities for action. The paper argues that for a robot to truly understand a sentence like "Be careful, that cup is fragile," it needs to have a grounded model of what "fragile" feels like when you hold it. It needs to know that if you squeeze too hard, the cup breaks. Current AI knows the word "fragile," but it might not have the internal "muscle memory" to understand the risk.

3. The Curiosity Engine
Have you ever felt that itch to explore something new just because it's interesting? That's not just a human quirk; it's a biological drive. The paper talks about how our brains have a "meta-model" that tracks how much we don't know. When we are curious, our brain releases chemicals (like norepinephrine) that make us more alert and ready to learn.
This is different from a robot that just waits for a teacher to give it a new dataset. Biological agents are "intrinsically motivated." They seek out new experiences to fill in the gaps in their own mental maps. The paper suggests that for AI to learn effectively, it shouldn't just wait for data; it should have an internal drive to go out and find the things it doesn't understand yet.

4. The Body's Alarm System (Allostasis)
Your brain isn't just thinking about the outside world; it's constantly monitoring your inside world, too. It tracks your hunger, thirst, temperature, and energy levels. This is called "allostatic control."
Imagine your brain as a thermostat that doesn't just react to the cold, but predicts you will get cold and tells you to put on a jacket before you even shiver. This system creates a sense of "self." It tells you what is good for you (food) and what is bad (pain). The paper argues that this internal feeling of "I need this" or "I want that" is the foundation of all motivation. Current AI systems usually have goals that humans give them (like "win this game"). They don't have their own internal "hunger" or "fear" that drives them to act.

5. The "Did I Do That?" Filter
When you wiggle your finger, your brain knows exactly what that will look like. It sends a copy of the "move" command to your eyes so it can cancel out the movement you caused. This helps you ignore your own movements and focus on things moving on their own (like a bird flying by).
This is how we know the difference between "me" and "the world." If you touch your nose, you know you did it. If someone else touches your nose, it's a surprise. The paper explains that this ability to distinguish self-generated actions from external events is crucial for understanding cause and effect. Without this, a robot might not understand that it is the one breaking the cup, or that it is the one causing the noise.

What This Means for the Future

The authors are careful not to say that current AI is "broken" or that biology is "perfect." They admit that the "passive reading" method has been incredibly successful for things like writing and chatting. However, they suggest that if we want robots to truly understand the physical world, to interact with humans naturally, and to learn on their own, we might need to change our approach.

They propose that future AI should be built more like a biological organism:

  • Start with the body: Let the AI learn through movement and interaction first, not just by reading text.
  • Build the map from the ground up: Let the AI develop a physical understanding of objects and space before trying to teach it complex language.
  • Add a sense of self: Give the AI internal drives (like curiosity or a need to maintain balance) so it learns actively, not just passively.
  • Learn socially: Just as humans learn by interacting with others, future AI might need to learn by hanging out with humans and other agents, building a "shared" understanding of the world.

The paper concludes that while we don't know exactly which parts of biology are essential and which are just lucky accidents of evolution, ignoring the way living things learn is a big risk. If we want AI that can truly navigate our world, we might need to stop treating it like a library and start treating it like a curious explorer.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →