← Latest papers
🧬 biology

Task-conditioned probing of instruction-tuned multimodal LLMs: Region-specific brain alignment patterns under naturalistic stimuli

This study demonstrates that instruction-tuned multimodal large language models exhibit significantly stronger and more task-specific alignment with human brain activity during naturalistic movie watching compared to non-instruction-tuned and in-context learning models, suggesting that instruction tuning organizes representations around functional task demands rather than surface semantics.

Original authors: Subba Reddy Oota, Khushbu Pahwa, Prachi Jindal, Satya Sai Srinath Namburi, Maneesh Singh, Tanmoy Chakraborty, Bapi S. Raju, Manish Gupta

Published 2026-05-21
📖 5 min read🧠 Deep dive

Original authors: Subba Reddy Oota, Khushbu Pahwa, Prachi Jindal, Satya Sai Srinath Namburi, Maneesh Singh, Tanmoy Chakraborty, Bapi S. Raju, Manish Gupta

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

The Big Idea: Teaching Computers to "Think" Like Brains

Imagine you have a group of very smart AI assistants (Multimodal Large Language Models, or MLLMs). These AIs can watch movies and listen to audio. The researchers wanted to know: Do these AIs process information the same way human brains do?

To find out, they didn't just ask the AIs questions; they looked at the "brainwaves" (fMRI scans) of real humans watching the same movies. They then tried to predict those brainwaves using the AI's internal thoughts. The goal was to see which AI "brain" matched the human brain best.

The Main Characters: Two Types of AIs

The study compared two types of AI assistants:

  1. The "Raw" AI (In-Context Learning): Think of this as a brilliant student who has read millions of books but has never been given a specific test. If you show them a movie and ask, "What's happening?", they answer based on their general knowledge. They are smart, but they might just be repeating what they've seen before.
  2. The "Trained" AI (Instruction-Tuned): Think of this as the same student, but they have just finished a rigorous training camp where they practiced specific tasks: "Summarize this," "Find the emotion," "Describe the sound." They are now experts at following specific instructions.

The Experiment: The Movie Night

The researchers set up a "movie night" with four human volunteers. While they watched clips from popular movies (like The Wolf of Wall Street), their brains were scanned.

Then, they fed those same movie clips into the AIs.

  • The Raw AI just looked at the video.
  • The Trained AI was given 13 different "instructions" for the same video, such as:
    • "What are the main events?"
    • "Describe the emotions."
    • "What is the central conflict?"
    • "Identify the sounds."

The researchers then asked: Which AI's internal "thoughts" could best predict what the human brain was doing at that exact moment?

The Surprising Results

Here is what they discovered, using some simple metaphors:

1. The "Trained" AI is a Better Match for the Human Brain
The "Trained" (Instruction-Tuned) AIs were significantly better at predicting human brain activity than the "Raw" AIs.

  • Analogy: Imagine trying to guess what a friend is thinking. The "Raw" AI is like a friend who just lists random facts about the movie. The "Trained" AI is like a friend who is actively trying to understand the story and the feelings, just like your brain does. The Trained AI's "thoughts" aligned with the human brain about 9% to 20% better than the others.

2. The "Raw" AI is Just Parroting Words; The "Trained" AI is Doing the Work
The researchers found a key difference in how these AIs think:

  • The Raw AI is very sensitive to the exact words you use. If you say "Describe the video" vs. "Tell me about the video," the Raw AI's internal brain changes a lot because the words are different. It's like a parrot mimicking sounds.
  • The Trained AI ignores the specific wording and focuses on the job you want done. Whether you say "Summarize" or "Tell the story," its internal brain stays focused on the task.
  • The Takeaway: Human brains also seem to care more about the task (understanding the story) than the exact words used to ask for it. The Trained AI mimics this human-like behavior better.

3. Different Tasks Light Up Different Parts of the Brain (and the AI)
The study showed that when you give the AI a specific job, it activates different "departments" in its brain, just like humans.

  • Analogy: Think of the brain as a city with different districts.
    • When the AI is asked to detect sounds, its "Audio District" lights up, matching the human auditory cortex.
    • When asked to understand a story, its "Language District" lights up, matching the human language areas.
    • When asked to recognize objects, its "Visual District" lights up.
  • The "Raw" AI didn't switch districts as clearly. The "Trained" AI knew exactly which part of the city to use for the specific job.

4. The "Deep" Layers Match the "Deep" Brain
AI models have layers, like a multi-story building.

  • Bottom floors (Shallow layers): These handle simple things like edges, colors, and basic sounds.
  • Top floors (Deep layers): These handle complex ideas like stories, emotions, and reasoning.
  • The Finding: The researchers found that the AI's bottom floors matched the human brain's sensory areas (eyes/ears), and the AI's top floors matched the human brain's thinking areas. This "hierarchy" was much clearer in the Trained AI.

The "Video vs. Audio" Twist

The study also looked at AIs designed specifically for video and AIs designed specifically for audio.

  • Video AIs: These were the stars of the show. They matched human brain activity very well across the whole brain.
  • Audio AIs: These did okay, but they mostly matched the human brain's hearing centers. They didn't match the rest of the brain as well as the video models did.
  • Conclusion: Currently, AI models that can "see" and "hear" together (Video) are much closer to human brain function than models that only "hear" (Audio).

Summary

This paper proves that when we teach AI models to follow specific instructions (like a human teacher would), they stop just "reciting" data and start "thinking" in a way that looks remarkably like how our own brains work. They organize their thoughts by task rather than by wording, and they activate specific brain regions depending on what they are asked to do. This makes them powerful tools for scientists trying to understand how the human brain processes the complex world of movies and stories.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →