← Latest papers
💬 NLP

Interpreting Brain Responses to Language with Sparse Features from Language Models

This paper introduces Augmented Sparse Encoding Models, which utilize hierarchically-organized sparse autoencoder features and surprisal to interpret 7T fMRI data, revealing that the human language cortex aligns with the most general information captured by artificial language models rather than arbitrary representations.

Original authors: Michael A. Lepori, Kendrick Kay, Greta Tuckute

Published 2026-06-08
📖 5 min read🧠 Deep dive

Original authors: Michael A. Lepori, Kendrick Kay, Greta Tuckute

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine trying to understand how the human brain works when we speak or listen. For a long time, scientists have tried to map this by looking at brain scans while people listen to sentences. At the same time, computers have become incredibly good at understanding language (these are called "Language Models" or LMs).

Usually, when scientists compare the brain to these computer models, they run into a problem: they are comparing one "black box" (the brain) to another "black box" (the computer). They can see that the two match up, but they can't explain why or what specific parts of the computer are doing the same thing as the brain.

This paper introduces a new way to look at this, called Augmented Sparse Encoding Models. Here is a simple breakdown of what they did and what they found, using some everyday analogies.

The New Tool: Turning a Smoothie into a Fruit Salad

Think of a standard computer language model as a smoothie. It takes all the words in a sentence and blends them into a single, dense, complex mixture. It's powerful, but you can't easily pick out the individual strawberries or bananas to see what they are doing.

The researchers used a new tool called a Sparse Autoencoder (SAE). Think of this as a machine that takes that smoothie and separates it back into a fruit salad. Instead of one blended mix, it breaks the information down into thousands of tiny, distinct "features" (like a specific piece of strawberry, a specific piece of banana, or a specific piece of ice).

  • Why do this? Because these individual pieces are easier to understand. Some pieces might represent "people," others "emotions," and others "questions."

They also added a specific ingredient to their mix: Surprisal. This is a measure of how "surprised" the computer is by a word. If a sentence says "The cat sat on the... mat," the computer isn't surprised. If it says "The cat sat on the... toaster," the computer is very surprised. This measures how hard it is for the brain to process the sentence.

The Experiment

The researchers had eight people listen to 200 different sentences while inside a super-powerful MRI machine (7T fMRI). They looked at tiny spots in the brain (called voxels) to see which ones lit up.

They used their new "fruit salad" tool to predict which brain spots would light up for which sentences.

What They Found

1. The "Fruit Salad" works just as well as the "Smoothie"
They found that using the separated features (the fruit salad) predicted brain activity just as accurately as using the blended smoothie. But the big win is that they could now see what the brain was actually responding to.

2. Two Different Types of Brain Spots
They discovered that the brain treats "difficulty" and "meaning" differently:

  • The "Difficulty" Spots: Some parts of the brain (mostly in the front) only care about how hard the sentence is to process. If a sentence is confusing or surprising, these spots light up. The researchers found that for these spots, you don't even need the complex "fruit salad" features; just knowing how "surprising" the sentence is (Surprisal) explains almost everything.
  • The "Meaning" Spots: Other parts of the brain (mostly in the back/temporal areas) care about what the sentence is about. These spots need the detailed "fruit salad" features to understand if the sentence is about concrete things (like a dog) or abstract ideas (like freedom).

3. Discovering a Hidden Group: The "People" Spots
The researchers found a group of brain spots that didn't fit into the usual categories of "easy/hard" or "concrete/abstract." They called these "Ghost" voxels because they were hard to explain before.
Using their new tool, they realized these spots are specifically tuned to people. They light up when sentences talk about relationships, pronouns (he, she, they), or people doing things. Interestingly, these spots are located in areas of the brain usually associated with social thinking, not just language.

4. The Brain Loves the "General" Features
The computer model they used has a hierarchy of features, like a library with a "General" section and a "Very Specific" section.

  • The "General" section has broad concepts (like "people" or "questions").
  • The "Specific" section has tiny, weird, or very niche details.
    The researchers found that the human brain mostly ignores the tiny, weird details. It relies almost entirely on the broad, general features. It seems that when we understand language, our brains are looking for the big picture concepts, not the microscopic quirks of the computer's internal math.

The Big Picture

This study shows that the human brain and artificial intelligence aren't just randomly matching up. They are aligned because both systems rely on the same general, high-level concepts to understand language.

By breaking the computer's "black box" into understandable pieces, the researchers could finally say: "Ah, this part of the brain lights up because it's thinking about people," or "This part lights up because the sentence was confusing." It turns a mystery into a clear map of how we process the words we hear.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →