The Capacity of Thought: Benchmarking Llama 3.2 in Semantic fMRI Neural Language Decoding and Improving the Huth Encoding-Model Baseline
This paper presents two complementary studies on fMRI-based language decoding: one that improves the traditional Huth encoding-model baseline to achieve higher METEOR and BLEU scores, and another introducing fMRIFlamingo, which reveals that high-capacity language models like Llama 3.2 can produce misleadingly high decoding metrics driven by their internal priors rather than actual neural signal when evaluated without rigorous blind controls.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine your brain is a bustling, noisy radio station broadcasting a secret story. Scientists have been trying to build a "thought decoder" that can tune into this station, catch the signal, and write down the story word-for-word. For a long time, the best way to do this was like using a very specific, old-school map (a model called GPT-1) to guess what the radio was saying.
Two teams of researchers at UC Santa Cruz decided to test two different ways to upgrade this decoder. They wanted to see if they could use a super-smart, modern AI (Llama-3.2) to read thoughts directly, or if they should just make the old map better.
The "Better Map" Upgrade
First, they took the old, reliable method and gave it a serious tune-up. Think of the old method as trying to listen to a radio with only 10,000 tiny antennas. The researchers added 5,000 more antennas (bringing the total to 15,000) to catch more of the signal. They also swapped out the old dictionary for a slightly newer one (GPT-2) to help guess the next word in the story, and they used super-fast computer chips (GPUs) to do the math 10 times faster.
The Result: This upgrade worked! By adding more antennas and a better dictionary, they improved their ability to guess the right words. They went from getting about 13.4% of the "score" right to 14.9%. It wasn't a magic fix, but it was a solid, measurable improvement. It proved that if you just give the old method a little more data and a better helper, it gets better at reading the brain's story.
The "Super-Brain" Experiment
Next, they tried something wild. They built a new machine called fMRIFlamingo. Instead of using a map to guess the story, this machine tried to plug the brain's radio signal directly into a giant, frozen super-AI (Llama-3.2). The idea was: "If we feed the brain signal into this super-smart AI, maybe the AI will just know what the person is thinking."
At first glance, it looked like a huge success. When they asked the AI to pick the right word from a list of 100 options (a ranking task), it selected the correct one 42.86% of the time. That sounds amazing, right? Much better than random guessing (which would be 1%).
But here is the twist: The researchers didn't just stop there. They played a trick. They turned off the brain signal completely—giving the AI a "blank" signal instead of the real brain data. They asked the same ranking question again.
The Shocking Discovery: The AI selected the correct answer 42.86% of the time again.
It turns out, the AI wasn't reading the brain at all. It was just using its own massive memory of how language works to guess the answer. It was like a student taking a test who doesn't need to read the question because they already memorized the answer key. The "brain signal" was actually getting in the way, slightly making the AI's guesses worse, not better.
The Big Lesson
The paper shows that just because a super-smart AI can guess the right word, it doesn't mean it's reading your mind. In fact, these big AIs are so good at guessing based on their own training that they can hide the fact that they are failing to read the brain.
The researchers found that:
- The old method, when improved, actually works better at decoding the story than the fancy new direct method when looking at full story reconstruction. While the new method did show a small improvement in picking single words correctly from a window (10.3% accuracy), the "Better Map" method achieved much higher scores when trying to reconstruct the actual narrative.
- The new method (fMRIFlamingo) is mostly relying on its own "language prior" (its internal knowledge of words) rather than the brain signal. When they tested it with a "blind control" (zeroing out the brain data), the ranking performance didn't drop, proving the brain data wasn't the hero.
- You can't trust a "success" in brain decoding unless you test it blindly. If you don't check if the AI is just guessing based on its own knowledge, you might think you've cracked the code of the mind when you haven't.
In short, the researchers showed that while we can make the old "map" method slightly better, simply plugging a brain into a giant AI doesn't automatically unlock thoughts. The AI is too smart for its own good, often guessing the right answer for the wrong reasons. To truly decode thoughts, we need to be much more careful and rigorous than just looking at a high score.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.