← Latest papers
💬 NLP

From Human Cognition to Neural Activations: Probing the Computational Primitives of Spatial Reasoning in LLMs

This paper investigates the internal mechanisms of spatial reasoning in multilingual LLMs by decomposing the task into three cognitive primitives and employing mechanistic analyses, revealing that while models encode transient and fragmented spatial information that can causally influence behavior, they lack robust, general-purpose spatial representations and instead rely on context-dependent pathways that exhibit mechanistic degeneracy across languages.

Original authors: Jiyuan An, Liner Yang, Mengyan Wang, Luming Lu, Weihua An, Erhong Yang

Published 2026-03-30
📖 5 min read🧠 Deep dive

Original authors: Jiyuan An, Liner Yang, Mengyan Wang, Luming Lu, Weihua An, Erhong Yang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very smart robot friend who has read almost every book ever written. You ask it, "If a cat is on the mat, and the mat is under the table, where is the cat?" The robot answers correctly: "Under the table."

For a long time, we assumed the robot was doing this by building a 3D mental map in its "brain," just like a human does. But this new paper asks a scary question: Is the robot actually seeing the room in its mind, or is it just guessing based on the words it's heard before?

The authors decided to stop just asking the robot "What's the answer?" and started performing "brain surgery" to see what was actually happening inside its circuits while it thought.

Here is the breakdown of their findings, using some everyday analogies:

1. The Three "Mental Gym" Exercises

To test the robot, the researchers didn't just give it random puzzles. They designed three specific "gym workouts" based on how human brains handle space, but stripped of any real-world context (no cats or tables, just abstract shapes and directions):

  • The Detective Game (Relational Reasoning): "A is left of B, B is above C. Where is A relative to C?"
    • The Test: Can the robot connect the dots to build a consistent picture?
  • The Spin Class (Perspective Transformation): "You are facing North. Turn right, then turn around. Which way are you facing now?"
    • The Test: Can the robot mentally rotate its viewpoint without getting dizzy?
  • The GPS Walk (Spatial Program Execution): "Start at (0,0). Walk 3 steps forward, 2 steps right, then flip over."
    • The Test: Can the robot keep a running tally of its position as it moves?

They ran these tests in English, Chinese, and Arabic to see if the robot's "brain" worked differently depending on the language.

2. The "Flashlight" Discovery (Where the thinking happens)

The researchers used a "flashlight" (a technique called probing) to shine a light into the robot's layers of neurons while it was solving these puzzles.

  • The Finding: They found that the robot does create a mental map, but it's like a flickering candle.
  • The Metaphor: Imagine the robot's brain is a long hallway.
    • In the middle of the hallway (the middle layers), the robot builds a very clear, bright 3D model of the room. It knows exactly where everything is.
    • But as the information travels toward the end of the hallway (the final layers where the answer is spoken), the lights go out. The 3D model fades away.
    • By the time the robot gives its answer, it has forgotten the map and is relying on a "gut feeling" or a linguistic guess based on the words it just processed.

3. The "Swiss Army Knife" vs. The "Specialized Tool"

The researchers also looked at the specific tools the robot used to solve these problems.

  • The Finding: The robot doesn't have one "Spatial Brain" that handles all space. Instead, it has different, disconnected tools for different jobs.
  • The Metaphor:
    • To solve the Detective Game, it uses a specific set of gears.
    • To solve the Spin Class, it uses a completely different set of gears.
    • To solve the GPS Walk, it uses a third set.
    • If you break the gears for the "Spin Class," the robot can still solve the "Detective Game." This means it doesn't have a unified, general understanding of space like a human does. It's just a collection of specialized tricks.

4. The "Language Barrier" Surprise

They tested the robot in three languages.

  • The Finding: The robot got the right answers in all three languages, but the way it got there was different.
  • The Metaphor: Imagine two people trying to get to the same destination.
    • The English speaker takes a path through the park.
    • The Arabic speaker takes a path through the subway.
    • They both arrive at the same spot (the correct answer), but they took completely different routes. This suggests the robot isn't using a universal "space language" inside its brain; it's just translating the problem into whatever language-specific path it knows best.

The Big Conclusion

So, do these AI models have "spatial intelligence"?

Yes and No.

  • Yes: They can build a mental map. They aren't just guessing randomly.
  • No: That map is fragile. It disappears before the robot speaks. It's not a solid, permanent structure like a human's cognitive map. It's more like a temporary sketch that gets erased the moment the robot has to give its final answer.

The Takeaway:
Current AI models are like brilliant actors who can memorize a script and pretend to be a navigator perfectly. But if you ask them to navigate a room they've never seen, or if you change the script slightly, they might stumble because they don't actually own the map; they just know the lines.

To make AI truly smart about the physical world, we need to teach them to keep that "mental map" lit up all the way to the end, not just in the middle of the process.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →