← Latest papers
💬 NLP

Vectors from Larger Language Models Predict Human Reading Time and fMRI Data More Poorly when Dimensionality Expansion is Controlled

This study demonstrates that when controlling for dimensionality expansion using untrained large language models, the predictive power of trained LLM vectors for human reading time and fMRI data exhibits inverse scaling, with the added value of training diminishing to zero at just a few billion parameters.

Original authors: Yi-Chien Lin, Hongao Zhu, William Schuler

Published 2026-09-01
📖 6 min read🧠 Deep dive

Original authors: Yi-Chien Lin, Hongao Zhu, William Schuler

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

For decades, scientists have watched computers learn to read and write with a speed that feels almost biological. These machines, known as large language models, are trained on vast libraries of human text until they can predict the next word in a sentence with startling accuracy. Because they handle language so well, researchers began to wonder if these machines were not just mimicking us, but actually mirroring the way our own brains process language. The prevailing hope was that as these models grew larger and more complex, they would become better and better at predicting how long it takes a human to read a word, or how our brains light up when we hear a story. This idea suggested that the architecture of these digital minds might be a map to the architecture of human thought, implying that any failure to match human data was simply a matter of needing more computing power.

However, a new study from researchers at The Ohio State University and the University of California, San Diego, challenges this optimistic view. They found that when they carefully controlled for the sheer size of the data these models produced, the relationship between model size and human-like performance changed in a specific way: the additional benefit of training disappeared. Instead of the training process making larger models significantly better than their untrained counterparts, the largest models performed no better than untrained versions of the same size. The study suggests that the apparent success of these massive systems was largely an illusion created by the sheer volume of information they could access, rather than a genuine understanding of language. When the researchers stripped away this advantage, the training that made these models so powerful seemed to offer no extra benefit over a completely untrained version of the same machine.

To understand how they reached this conclusion, one must first look at how these models are typically tested. Researchers often feed a story into a language model and compare the model's internal state to human data, such as how long a person pauses on a specific word or how their brain responds to a sentence. In previous studies, it appeared that larger models, which have more internal variables to work with, consistently predicted human behavior better. This led to a theory called the "quality-power" hypothesis: the bigger the model, the more it resembles the human mind. But the researchers suspected a hidden flaw in this logic. When a model gets bigger, it doesn't just get smarter; it also produces a much larger number of internal signals, or vectors, to analyze. It is like giving a student a test with twice as many questions; even if the student knows nothing, having more questions gives them more chances to guess the right answer by luck. The researchers realized that previous studies might have been measuring the benefit of having more questions, rather than the benefit of having a better brain.

To solve this puzzle, the team designed a series of experiments that separated the effect of size from the effect of intelligence. They gathered data from five different families of language models, ranging from small ones with a few hundred million parameters to massive ones with 66 billion parameters. They tested these models against real human data, including reading times from people reading natural stories and brain scans from people listening to those same stories. In their first experiment, they replicated the earlier findings: as the models got bigger, their predictions of human behavior did seem to improve. This confirmed that the "quality-power" effect was visible on the surface.

But then, they introduced a crucial control. They took the same models and looked at them before they were trained on any text at all. These untrained models have the exact same size and structure as the trained ones, but they are essentially random noise, with no knowledge of language. By comparing the trained models against their untrained twins, the researchers could see how much of the improvement came from actual learning and how much came simply from having a larger number of internal signals to work with. They used a method similar to a reservoir, where a large, untrained network acts as a source of raw material that can be shaped by a simple classifier. If the training was truly making the model more human-like, the trained version should have significantly outperformed the untrained version, especially as the models grew larger.

The results were surprising. When the researchers controlled for the size of the internal signals, the extra advantage of training disappeared. In fact, for models larger than a few billion parameters, the trained versions performed no better than the untrained ones. On several datasets, the extra benefit of training dropped to zero. This means that the massive improvements seen in earlier studies were not because the larger models had learned a deeper understanding of language, but simply because they had more internal dimensions to fit the data. The researchers found that as models grew beyond a certain point, the contribution of training to the model's fit over an untrained baseline dropped to zero, rather than increasing. The larger the model became, the less the training process added to the predictive power provided by the model's size alone.

The study also looked at different layers within the models, checking if the middle sections of the network, which some previous work suggested were more important, behaved differently. They found the same pattern: the contribution of training dropped to zero across all layers as the models grew. Even when they adjusted their statistical methods to account for the possibility that the models were simply hitting a limit on how well they could predict human data, the trend held. The researchers concluded that the success of these large models in predicting human behavior is largely due to a statistical effect where having more variables allows for a better fit, rather than a genuine reflection of human cognition.

This finding suggests a significant misalignment between the way these artificial systems are built and the way human brains work. The training process that makes these models so good at generating text does not necessarily make them better models of human reading or brain activity. In fact, the study argues that the untrained versions of these massive networks, which act as complex reservoirs of random connections, might be just as effective at predicting human data as the expensive, fully trained versions. The researchers propose that future studies should use these untrained models as a standard baseline to ensure that any improvements attributed to training are real and not just a byproduct of the model's size. While the idea that bigger is better has driven the field forward for years, this work suggests that in the realm of human language processing, there is a point where adding more size and training does not bring us closer to understanding the human mind, but instead takes us further away.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →