← Latest papers
🧬 biology

Brain alignment of reasoning and action representations from vision-language and action models during naturalistic gameplay

This study demonstrates that fine-tuning foundation models for action (LAMs) versus reasoning (VLMs) during naturalistic gameplay leads to distinct brain alignment patterns, where LAMs show asymmetric, action-specialized representations in frontal-motor regions while both models outperform reinforcement learning baselines with gains scaling along the cortical processing hierarchy.

Original authors: Subba Reddy Oota, Anant Khandelwal, Khushbu Pahwa, Satya Sai Srinath Namburi, Tanmoy Chakraborty, Bapi S. Raju, Manish Gupta

Published 2026-05-20
📖 5 min read🧠 Deep dive

Original authors: Subba Reddy Oota, Anant Khandelwal, Khushbu Pahwa, Satya Sai Srinath Namburi, Tanmoy Chakraborty, Bapi S. Raju, Manish Gupta

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine you are watching a group of people play a classic arcade video game (like Pac-Man or Space Invaders) while inside a giant MRI machine. This machine acts like a super-sensitive camera that takes pictures of their brains every few seconds, showing which parts light up as they see the screen, think about their next move, and press the buttons.

The goal of this paper is to see if modern AI computers "think" about the game in a way that looks similar to how human brains do.

Here is the breakdown of the study using simple analogies:

1. The Players: Three Types of "Minds"

The researchers compared three different types of "minds" trying to understand the game:

  • The Old-School Robot (RL Baselines): These are traditional AI agents trained only to win by trial and error. They are like a dog learning to press a lever for a treat; they know what to do, but they don't really "understand" the game world or the rules.
  • The Generalist Observer (VLMs): These are modern AI models (Vision-Language Models) that have read the internet and watched millions of videos. They are like a very smart tourist who can describe the game, explain the rules, and talk about the graphics, but they haven't necessarily been trained to play it perfectly.
  • The Action Specialist (LAMs): These are similar to the Generalist, but they have been specifically "fine-tuned" or trained on massive amounts of data about how to control computers and play games. They are like a professional gamer who has practiced the controls until they are second nature.

2. The Experiment: The "Prompt" Trick

The researchers didn't just let the AI look at the screen. They gave them specific instructions, or "prompts," to see how that changed their thinking:

  • The "Action" Prompt: "What is your next move, and why?" (Focuses on doing).
  • The "Reasoning" Prompt: "What is happening in the game, what are the goals, and what are the threats?" (Focuses on understanding).

They then compared the AI's internal "thoughts" (digital representations) to the human brain scans to see which AI matched the human brain better.

3. The Big Discoveries

Discovery A: The New AI is Much Closer to Humans
The "Old-School Robots" were okay, but the modern AI models (both the Generalist and the Action Specialist) were much better at matching human brain activity.

  • Analogy: Imagine trying to guess what a human is thinking. The old robot guesses based on simple patterns. The new AI guesses based on a rich understanding of the story, the objects, and the goals. The new AI's "guess" matches the human brain's actual activity much more closely.

Discovery B: The "Thinking" Boost Happens in the "Planning" Parts of the Brain
When the researchers used the "Action" or "Reasoning" prompts, the AI's match with the human brain got even better. But this improvement wasn't equal everywhere.

  • Analogy: Think of the brain like a city. The "Early Visual Cortex" is the front gate where people first see the game. The "Frontal-Parietal" areas are the city hall where planning and decision-making happen.
  • The prompts helped the AI match the human brain twice as much in the "City Hall" (planning/decision areas) compared to the "Front Gate" (visual areas). This suggests that when humans play, they aren't just seeing pixels; they are heavily engaged in planning and reasoning, and the AI models capture this best when asked to think about it.

Discovery C: The "Symmetry" vs. "Asymmetry" Surprise
This is the most interesting finding. Even though the Generalist (VLM) and the Action Specialist (LAM) were equally good at predicting the brain activity overall, they were doing it in very different ways.

  • The Generalist (VLM) is Balanced: When you ask it to "Reason" or to "Act," it uses its brain power roughly equally for both. It's like a Swiss Army knife that is equally good at opening a bottle or cutting a rope.
  • The Action Specialist (LAM) is Biased: When you ask the Action Specialist to "Reason," it actually struggles to add new value. It seems that its training to "Act" has already absorbed so much information that asking it to "Reason" separately is redundant.
  • Analogy: Imagine the Action Specialist is a chef who has memorized the entire recipe book. If you ask them, "What ingredients do you need?" (Reasoning), they just repeat what they already know about cooking (Action). The "Reasoning" part of their brain becomes redundant because their "Action" training already covers it. In the human brain, this showed up as the Action Specialist's "Reasoning" signal actually getting in the way of the prediction in the planning areas of the brain.

Discovery D: The "Thinking" Trace is Weaker
The researchers also looked at a newer type of AI that "thinks out loud" (generating a chain of thought before answering). Surprisingly, the part of the AI that generated the "thinking" text was less like the human brain than the part that just gave the final answer.

  • Analogy: It's like watching someone solve a puzzle. The final move they make matches how a human would move their hand. But the internal monologue they say while solving it ("Hmm, maybe I should try this...") doesn't match the human brain's activity as well as the final decision does.

Summary

The paper shows that modern AI models are becoming much better at simulating how humans think and plan during games. However, simply making an AI better at doing things (Action) changes its internal structure in a specific way: it makes "Reasoning" and "Action" overlap so much that they become indistinguishable in the brain's planning centers. This helps scientists understand that while AI and humans might reach the same "score" (prediction accuracy), the way their internal "brains" are organized can be fundamentally different.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →