← Latest papers
💬 NLP

Probing for Reading Times

This study demonstrates that early-layer representations in language models effectively predict early-pass human reading times across five languages, outperforming scalar surprisal, while scalar surprisal remains superior for late-pass measures, highlighting a functional alignment between model depth and the temporal stages of human reading.

Original authors: Eleftheria Tsipidi, Samuel Kiegeland, Francesco Ignazio Re, Tianyang Xu, Mario Giulianelli, Karolina Stanczak, Ryan Cotterell

Published 2026-04-22
📖 5 min read🧠 Deep dive

Original authors: Eleftheria Tsipidi, Samuel Kiegeland, Francesco Ignazio Re, Tianyang Xu, Mario Giulianelli, Karolina Stanczak, Ryan Cotterell

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are watching a movie of someone reading a book. You can see their eyes stop (or "fixate") on certain words for a split second longer than others. In the world of reading research, these tiny pauses are like heartbeats of the brain. A longer pause usually means the brain is working harder to understand that specific word or phrase.

For a long time, scientists have tried to predict these pauses using a simple math trick called "Surprisal." Think of Surprisal as a "shock meter." If a sentence says, "The cat sat on the mat," the word "mat" has a low shock meter because it's expected. If the sentence says, "The cat sat on the... toaster," the shock meter spikes. The theory was: The higher the shock, the longer the pause.

But in this new paper, the researchers asked a bigger question: "Is the shock meter the whole story? Or does the computer's brain (the Language Model) have a more detailed map of how humans read?"

Here is the breakdown of their discovery, using some everyday analogies.

1. The "Layer Cake" of the AI Brain

Modern AI models (like the ones used in this study) aren't just one big brain; they are like a multi-layered cake.

  • The Bottom Layers (Early Layers): These are like the "appetizers." They handle the basics: recognizing letters, sounds, and simple word shapes.
  • The Middle Layers: These start mixing things up, looking at grammar and sentence structure.
  • The Top Layers (Late Layers): These are the "main course." They understand the deep meaning, the context, and the complex ideas.

The researchers peeled back the cake, layer by layer, to see which part of the AI's brain best predicted how long a human would pause while reading.

2. The Big Discovery: It Depends on When You Look

The results were fascinating because they showed that different layers of the AI predict different stages of human reading.

The "First Glance" (Early Layers)

When a human first looks at a word, their eyes make a quick decision: "Do I know this word? Does it look right?"

  • The Finding: The bottom layers of the AI were the best at predicting these quick, initial pauses.
  • The Analogy: Imagine you are walking into a grocery store. The bottom layers of the AI are like your eyes scanning the aisle to see if you recognize the brand of cereal. They are fast, structural, and focused on the "look" of the word. The AI's early layers capture this "first impression" better than the simple "shock meter" (Surprisal) ever could.

The "Deep Dive" (Top Layers & Surprisal)

Sometimes, a reader gets stuck. They might re-read a sentence to figure out the grammar or the deep meaning. This takes longer.

  • The Finding: For these long, deep pauses, the simple "shock meter" (Surprisal) actually worked best, often beating the complex AI layers.
  • The Analogy: This is like the moment you realize, "Wait, if the cat sat on the toaster, why isn't it on fire?" You have to stop and think about the logic. The simple math of "how unexpected is this?" is surprisingly good at capturing this moment of confusion. The deep, complex layers of the AI didn't add much extra value here; the simple surprise was enough.

3. The "Swiss Army Knife" vs. The "Specialist"

The researchers also found that the "best tool" changes depending on the language.

  • In English, the AI's early layers were great specialists for the first glance.
  • In languages like Greek, Hebrew, or Turkish, the simple "shock meter" often did just as well as the fancy AI layers.
  • The Takeaway: There is no single "magic bullet." Just like you wouldn't use a sledgehammer to crack a nut, you can't use one single AI measurement to predict every reading pause in every language.

4. The "Best of Both Worlds"

The most exciting part? When they combined the AI's early layers (the "first glance" specialist) with the Surprisal (the "shock meter"), the prediction got even better.

  • The Analogy: It's like having a team. One person is great at spotting the shape of the word (Early Layers), and another is great at calculating how weird the sentence is (Surprisal). When they work together, they can predict exactly when a human will pause, even better than either could alone.

Summary

This paper is a reminder that human reading is a two-step dance:

  1. Step 1: A quick, visual check (Is this a word I know?). The AI's early layers are the best at predicting this.
  2. Step 2: A slower, logical check (Does this make sense?). The simple Surprisal math is surprisingly good at predicting this.

The authors show that while AI models are getting smarter, they aren't just "black boxes." By looking inside their "layers," we can see a functional map that mirrors how our own brains process language, from the first split-second glance to the deep, thoughtful pause.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →