← Latest papers
💬 NLP

PoTeC: A German Naturalistic Eye-tracking-while-reading Corpus

The Potsdam Textbook Corpus (PoTeC) is a novel, open-access German eye-tracking dataset featuring 75 participants (both domain experts and novices) reading 12 scientific texts within a fully-crossed factorial design, accompanied by comprehension assessments, linguistic annotations, and full preprocessing code to facilitate the study of expert versus non-expert reading strategies.

Original authors: Deborah N. Jakobi, Thomas Kern, David R. Reich, Patrick Haller, Lena A. Jäger

Published 2026-02-24
📖 4 min read☕ Coffee break read

Original authors: Deborah N. Jakobi, Thomas Kern, David R. Reich, Patrick Haller, Lena A. Jäger

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to understand how people read. For decades, scientists have done this by putting people in a lab and showing them tiny, carefully crafted sentences, almost like a chef testing a single spice in isolation. But in the real world, we don't read isolated words; we read entire paragraphs, textbooks, and news articles.

Enter PoTeC (Potsdam Textbook Corpus). Think of PoTeC as a massive, high-definition "flight recorder" for the human brain while it's reading. It's a new dataset that tracks exactly where people's eyes move when they read real German textbooks, not just made-up sentences.

Here is the breakdown of what makes this paper special, using some everyday analogies:

1. The "Expert vs. Novice" Gym Match

Most reading studies treat everyone the same. PoTeC is different because it's like a boxing match where you pair a Grandmaster against a Beginner, but with a twist: The same person plays both roles.

  • The Setup: They took 75 university students. Some were studying Physics, some Biology. Some were fresh undergraduates (novices), and some were PhD students (experts).
  • The Twist: Everyone read texts from both fields.
    • A Physics PhD student reading a Physics textbook is the Expert.
    • That same Physics PhD student reading a Biology textbook is the Novice.
  • Why it matters: This allows researchers to see how the same brain changes its reading strategy when it knows the material versus when it's struggling to understand it. It's like watching a professional chef cook a familiar dish versus trying to cook a recipe from a cuisine they've never seen before.

2. The "Black Box" of Eye Movements

The researchers used a super-precise camera (an eye-tracker) to record eye movements 1,000 times per second.

  • The Analogy: Imagine reading a book is like driving a car. Your eyes are the headlights. Sometimes you stare straight ahead (reading smoothly), sometimes you jerk your head back to check a sign you missed (re-reading), and sometimes you glance at the dashboard (skipping words).
  • The Data: PoTeC records every single "glance" and "stare." It tells us how long a person looked at a word, if they skipped it, or if they went back to re-read it.

3. The "Super-Annotation" Layer

Just having the eye movements isn't enough; you need to know what the words were. The team didn't just dump the text; they labeled it with a massive amount of metadata.

  • The Analogy: Imagine the textbook pages are covered in invisible ink. The researchers used different "magic pens" to reveal hidden layers:
    • The Grammar Pen: Marking every verb and noun.
    • The Difficulty Pen: Calculating how surprising a word is (e.g., "The cat sat on the..." vs. "The cat sat on the... toaster").
    • The Expert Pen: Marking which words are "jargon" that only a pro would know (like "homolog" in biology) versus words a normal person knows (like "protein").
  • This means researchers can ask: "Do experts skip the jargon because they know it, or do they stare at it longer because it's complex?"

4. The "Raw vs. Polished" Gem

One of the coolest features of PoTeC is that they didn't just give you the final, perfect data. They gave you the raw, messy data too.

  • The Analogy: Usually, a scientist gives you a finished cake. PoTeC gives you the cake plus the raw eggs, the flour, and the mixing bowl.
  • Why? Because sometimes the eye-tracker drifts (like a camera slowly tilting to the left). The researchers manually fixed this drift. By giving you both the "drifting" version and the "fixed" version, they are inviting other scientists to build AI tools that can automatically fix these errors in the future. It's an open invitation to improve the tools of the trade.

5. Why Should You Care?

This isn't just for linguists. This data is a goldmine for:

  • AI and Computers: Teaching computers to read like humans. If an AI can predict where a human's eyes will go, it understands the text better.
  • Education: figuring out exactly where a student gets stuck in a textbook so we can design better learning materials.
  • Biometrics: Could your unique eye movement pattern be used as a password? (Yes, apparently!)

The Bottom Line

PoTeC is a giant, open-source library of "how humans read." It's the first time we have a dataset that lets us compare experts and beginners side-by-side using real, difficult textbooks, with every single eye movement and linguistic detail meticulously recorded. It's like giving the scientific community a high-definition map of the brain's reading journey, complete with all the potholes, detours, and scenic routes.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →