← Latest papers
💬 NLP

Modeling semantic association in self-paced reading with language model embeddings

This study demonstrates that while language model embeddings can quantify semantic association in self-paced reading and EEG data, the reliability of these effects on both neural (N400) and behavioral measures critically depends on methodological choices, with sentence-based embeddings outperforming word-based approaches in capturing semantic associations beyond word predictability.

Original authors: Sara Møller Østergaard, Kenneth Enevoldsen, Afra Alishahi, Bruno Nicenboim

Published 2026-06-08
📖 4 min read☕ Coffee break read

Original authors: Sara Møller Østergaard, Kenneth Enevoldsen, Afra Alishahi, Bruno Nicenboim

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are reading a story. Your brain is constantly trying to guess what word is coming next. If the story says, "The hiker's feet were cold and wet, so he needed a new pair of...", your brain instantly shouts, "Boots!" That's predictability.

But what if the story said, "...so he needed a new pair of sandals"? You wouldn't expect "sandals" (predictability is low), but your brain still understands the connection because sandals are also shoes for feet. That connection is semantic association.

This paper is like a detective story trying to figure out how to measure that "connection" between a word and its surroundings using computer models. The researchers wanted to see if these computer models could explain how fast we read and how our brains react (measured by electrical signals called the N400) when we encounter these words.

The Big Problem: Too Many Ways to Measure "Connection"

The researchers found that scientists have been using language models (AI that reads text) to measure this "semantic association" in about 10 different ways. It's like trying to measure the distance between two cities:

  • Some people use a straight line on a map.
  • Some people drive the actual roads.
  • Some people count the steps.
  • Some people fly a drone.

The paper asked: Does it matter which "ruler" we use?

The Experiment: The "Ruler" Test

The team took a large collection of Dutch stories that people had read while wearing EEG caps (which record brain waves) and while reading at their own pace. They then ran the text through 10 different "ruler" setups to calculate the semantic association.

These setups varied in two main ways:

  1. The Model: Did they use a model that looks at words one by one (like a dictionary definition), or a model that looks at the whole sentence as a single idea (like understanding the vibe of a paragraph)?
  2. The Context: Did they look at just the word before the target, the whole sentence, or the entire paragraph?

The Results: The "Vibe" Ruler Wins

Here is what they discovered, using simple analogies:

1. The "Dictionary" Ruler Failed
When they used models that looked at words individually (like a standard dictionary or a list of definitions), the results were messy.

  • The Brain Signal (N400): Sometimes it said words with a strong connection made the brain react less, and sometimes it said they reacted more. It was inconsistent.
  • Reading Speed: It didn't seem to predict how fast people read at all.
  • Analogy: Imagine trying to understand a joke by looking up every single word in a dictionary. You get the definitions, but you miss the humor. The "dictionary" approach missed the point of the story.

2. The "Vibe" Ruler Succeeded
When they used Sentence Embeddings (models trained to understand the "vibe" or overall meaning of a whole sentence), the results were clear and consistent.

  • The Brain Signal: These models showed that when a word was semantically linked to the story (even if unexpected), the brain reacted in a specific, predictable way.
  • Reading Speed: These models successfully predicted that people read words faster when they fit the "theme" of the story, even if they weren't the most obvious prediction.
  • Analogy: This is like understanding a joke because you get the context and the tone. The "vibe" model captured the theme of the story (e.g., "dragons" in a fairy tale) rather than just the individual words.

The "Dragon" Example

The paper gives a great example to show the difference.

  • The Text: A story about a "Dragon King." Later, the word "dragon" appears again.
  • The Dictionary Model: It sees the word "dragon" and thinks, "Okay, I know this word." But it doesn't realize how strongly this word connects to the whole story about the Dragon King.
  • The Vibe Model: It realizes, "Ah, this word 'dragon' is the heartbeat of this entire story!" It captures the deep thematic link that the dictionary model missed.

The Bottom Line

The main takeaway is that how you measure "meaning" changes what you find.

If you want to understand how the human brain processes the meaning of a story (not just guessing the next word), you can't just look at word-by-word definitions. You need a tool that understands the whole sentence and the overall theme.

The paper concludes that "Sentence Embeddings" (the "Vibe" ruler) are the best tool we currently have for studying how our brains connect words to the bigger picture of a story. The old way of just averaging word definitions doesn't seem to work as well for natural, flowing text.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →