← Latest papers
💬 NLP

Eye Tracking Based Cognitive Evaluation of Automatic Readability Assessment Methods

This paper introduces an eye-tracking-based cognitive evaluation framework that reveals traditional readability formulas, NLP methods, commercial systems, and frontier LLMs are poor predictors of real-time reading ease compared to psycholinguistic word properties, highlighting the need for new cognitively driven scoring approaches.

Original authors: Keren Gruteke Klein, Shachar Frenkel, Omer Shubi, Yevgeni Berzak

Published 2026-07-29
📖 4 min read☕ Coffee break read

Original authors: Keren Gruteke Klein, Shachar Frenkel, Omer Shubi, Yevgeni Berzak

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to figure out how easy a book is to read. For over a hundred years, scientists and teachers have tried to build "readability formulas"—mathematical recipes that look at a text and spit out a grade level, like "this is a 5th-grade text" or "this is a 12th-grade text." They usually do this by counting things like how long the sentences are or how many big words are in them. But here's the tricky part: just because a text has short sentences doesn't mean a human actually feels like it's easy to read while they are reading it. To really know if a text is easy, you need to see how a person's brain is working in real-time. This is where "eye tracking" comes in. It's like putting a tiny camera on a person's eye to see exactly where they look, how long they stare at a word, and if they have to look back (regress) because they got confused. It's the difference between guessing how fast a car is going by looking at the engine size versus actually driving it and feeling the speed.

This paper, titled "Eye Tracking Based Cognitive Evaluation of Automatic Readability Assessment Methods," is a big reality check for the world of text analysis. The researchers, led by Keren Gruteke Klein and her team at the Technion in Israel, decided to test the most popular ways we currently judge text difficulty. They gathered a massive group of 638 adult readers—some native English speakers and some who learned English as a second language—and had them read news articles on a computer while their eyes were tracked at a super-fast speed of 1,000 times per second. They used a clever trick: they took the same news stories and had human teachers rewrite them to be simpler. This allowed the researchers to compare the "hard" version and the "easy" version of the exact same content, isolating the writing style from the topic.

The team then asked a simple question: Which computer program is the best at predicting how much easier the "easy" version actually felt to the readers? They tested everything from old-school math formulas (like the Flesch Reading Ease score) to modern AI systems, commercial tools used in schools, and frontier Large Language Models (LLMs). Notably, the study evaluated models like GPT-4o and other cutting-edge systems (including versions labeled with future dates like 2025 in the paper) using specific prompts that asked the AI to assign a grade level or a difficulty score from 1 to 100. They compared these high-tech tools against some very simple, old-school measurements from psychology, like how often a word is used in the language or how surprising a word is to see next.

The results were a bit of a shock. Despite all the fancy technology and the incredible power of modern AI, the popular readability tools were terrible at predicting how easy a text actually felt to read in real-time. The old formulas, the commercial school tools, and even the cutting-edge AI models all failed to match the readers' eye movements. In fact, the best predictors turned out to be the simplest things: how long a word is, how common it is, and a measure called "surprisal" (which is basically a math way of saying "how unexpected is this word?"). The study suggests that while our current tools are great at guessing what grade level a text belongs to, they are missing the mark on the actual human experience of reading. The authors conclude that we need to stop relying on these old methods for high-stakes decisions and start building new systems that are actually driven by how our brains process words in the moment.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →