← Latest papers
💬 NLP

On the Proper Treatment of Units in Surprisal Theory

This paper proposes a unified framework to disentangle the definition of linguistic units from evaluation regions in surprisal theory, arguing that tokenization should be treated as an implementation detail rather than a scientific primitive to ensure explicit and rigorous modeling of human processing effort.

Original authors: Samuel Kiegeland, Vésteinn Snæbjarnarson, Tim Vieira, Ryan Cotterell

Published 2026-05-01
📖 5 min read🧠 Deep dive

Original authors: Samuel Kiegeland, Vésteinn Snæbjarnarson, Tim Vieira, Ryan Cotterell

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Problem: The "Wrong Ruler"

Imagine you are a psychologist trying to measure how hard it is for a human brain to read a sentence. You have a theory called Surprisal Theory, which says: The harder a word is to predict, the more mental effort it takes to read.

To test this, you need a "ruler" to measure that effort. In the past, researchers used a ruler that matched the words perfectly (like measuring a table in inches). But today, we use powerful AI models (like GPT-2) to do the measuring. The problem? These AI models don't speak "human words." They speak in "tokens."

Think of a token like a Lego brick.

  • A human sees the word "don't".
  • The AI sees two bricks: "don" and "'t".
  • A human sees the word "words".
  • The AI sees one brick: "words".
  • A human sees the word "equal".
  • The AI sees one brick: "equal".

Because the AI's "bricks" (tokens) don't line up with our "words," researchers have been trying to glue the AI's bricks together to match human words. They've been using "ad hoc" (make-shift) glue. Sometimes they glue a space to the front of a word; sometimes to the back. Sometimes they ignore punctuation.

The Paper's Main Claim: This "gluing" is messy and inconsistent. It's like trying to measure a room with a ruler that changes its own length every time you look at it. The authors argue that the choice of what counts as a "unit" (a word, a letter, a phrase) is a scientific decision, not just a technical detail. We need to stop letting the AI's internal brick structure dictate our science.

The Solution: "Bring Your Own Units"

The authors propose a new framework where the researcher decides exactly what a "unit" is before looking at the AI.

  1. You choose the unit: Do you want to study whole words? Individual letters? Or maybe specific parts of speech?
  2. You translate the AI: Instead of forcing the AI to change, you build a "translator" (a mathematical tool called a transducer) that converts the AI's weird bricks into your chosen units.

The Analogy:
Imagine the AI is a chef who only speaks in "grams of flour, sugar, and eggs." You, the scientist, want to know how much "cake" is being made.

  • Old Way: You try to guess how many grams make a cake based on the chef's random measurements.
  • New Way: You build a machine that takes the chef's grams and automatically assembles them into perfect "cakes" (or "cupcakes," or "muffins") based on your recipe. The chef doesn't change; your machine just interprets the output correctly for your specific question.

The Experiment: Testing Different Rulers

To prove their point, the authors ran an experiment using the MECO dataset (a collection of eye-tracking data from people reading text). They tested four different "rulers" (unit inventories) to see which one best predicted how long people's eyes stayed on a word:

  1. Tokens: Using the AI's native bricks (e.g., "don" and "'t" as separate things).
  2. Characters: Measuring every single letter (e.g., d, o, n, ', t).
  3. Acontextual Words: Splitting text by simple rules (e.g., "cut at every space"). This is like cutting a string of beads at every gap, regardless of what the beads look like.
  4. Contextual Words: Using smart rules (like the Penn Treebank guidelines) that understand context. For example, knowing that a comma in "1,000" is part of the number, but a comma in "Hello, world" is a pause.

The Results: It Matters What You Measure

The study found that changing the ruler changes the results.

  • Whole Words (Contextual): When they used "smart" word units that respected punctuation and grammar, the AI's predictions matched human eye movements very well.
  • Tokens: The AI's native bricks worked okay, but not as well as the smart words.
  • Characters: Measuring letter-by-letter didn't work well for predicting reading time. Why? Because humans rarely stop their eyes on a single letter; they look at whole words.
  • The "Glue" Problem: They found that how you handle spaces (leading vs. trailing) changed the math significantly.

The Takeaway:
If you use a bad ruler, you might conclude that "Surprisal Theory is wrong" or "AI models are bad," when in reality, you just measured the wrong thing.

The Core Lesson

The paper concludes that tokenization is just a technical detail, like the brand of the camera lens. It shouldn't be the star of the show.

  • Old Mindset: "The AI uses tokens, so we must study tokens."
  • New Mindset: "I want to study how humans process words. I will take the AI's output and mathematically convert it into words so I can answer my scientific question correctly."

By treating the definition of a "unit" as a deliberate scientific choice rather than a default setting, researchers can get clearer, more accurate answers about how the human brain processes language.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →