Repeated Sequences Reveal Gaps between Large Language Models and Natural Language
This paper introduces a novel evaluation framework based on repeated subsequences and Rényi entropies to demonstrate that, unlike natural language which exhibits stable long-range structural organization, state-of-the-art large language models display systematic and size-dependent deviations in their entropy growth patterns, revealing fundamental gaps in how they capture the deep statistical structure of text beyond surface fluency.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: Fluency vs. Structure
Imagine you are listening to two people tell a story.
- Person A speaks perfectly. Every sentence flows smoothly, the grammar is correct, and they never stutter.
- Person B also speaks perfectly, but they have a secret: they are actually just repeating the same few phrases over and over, or they are building a story by stitching together tiny, pre-made Lego bricks without ever creating a new shape.
For a long time, we thought Large Language Models (LLMs) like GPT were getting so good that they sounded exactly like humans. But this paper asks a deeper question: Do they actually understand how to build a long story from the inside out, or are they just really good at sounding smooth in the short term?
The author, Kumiko Tanaka-Ishii, developed a new way to "listen" to these stories not by checking if they make sense grammatically, but by looking at how they reuse their own words and patterns over long distances.
The Tool: The "Echo" Detector
To find the difference, the author uses a concept called Repeated Subsequences.
Think of a text as a long hallway.
- Natural Language (Humans): When a human writes a novel, they plant seeds early on. They introduce a character, a theme, or a specific phrase. As the story goes on, they come back to those seeds, water them, and grow new branches. The story "echoes" itself. It reuses old structures to build new meaning.
- LLMs (AI): The author found that AI models often act like a broken record or a hall of mirrors. They might repeat things, but they don't seem to have that same deep, structural "echo" where old parts of the story are actively reorganized to support the new parts.
The paper measures this by counting how many times small chunks of text (like 5-letter or 10-letter sequences) appear again later in the document.
The Two Types of Growth: The Garden vs. The Factory
The paper compares how "information" grows as a text gets longer. It identifies two main ways this can happen:
The Power-Law Garden (Natural Language):
Imagine a garden where you keep planting new, unique flowers, but you also keep reusing the same trellis and fence to support them. As the garden grows, the variety of flowers increases, but the structure (the fences and trellises) is reused over and over.
The paper finds that human writing behaves like this. It follows a "Log-Power" pattern. This means the text is constantly recycling its own structural DNA. It's efficient and deeply interconnected.The Factory Line (AI Models):
Imagine a factory that keeps churning out products. As the factory runs longer, it keeps making slightly different items, but it doesn't seem to have that same deep, recursive reuse of its own internal blueprint.
The paper found that AI models, especially older or smaller ones, behave differently. They show a "Power-Law" pattern that suggests they are constantly generating new structural complexity without the same level of deep, efficient reuse found in human writing.
The Experiment: The "Shuffle" Test
The author tested this on:
- Human-written novels (from Project Gutenberg).
- AI-generated stories (created by GPT-3.5, GPT-4, and newer models).
They matched the lengths so the comparison was fair. They then ran a mathematical "echo test" (using something called Rényi Entropy) to see how the text reused itself.
The Results:
- Humans: No matter how long the book was, the "echo" pattern stayed stable. Humans consistently reused their own structural patterns in a predictable, efficient way.
- AI: The AI's pattern changed depending on how "big" the model was.
- Smaller AI models were very chaotic in how they reused (or failed to reuse) structure.
- As the AI models got bigger (like GPT-5), they started to look more like humans, but they still showed a systematic difference. They were getting better at mimicking the "echo," but the underlying math of how they built their stories was still distinct from a human author.
The "Maximal Repetition" Trap
The paper also looked at the longest repeated phrase in a text (like finding the longest sentence that appears twice).
- Previous studies focused on these "extreme" repetitions.
- The author argues this is like trying to understand a forest by only looking at the tallest tree. It's unstable and misleading.
- Instead, looking at all the repeated chunks (the "distribution") gives a much clearer, more stable picture of the text's true structure.
The Bottom Line
This paper doesn't say AI is "bad" or that it can't write good stories. It says that AI and humans build stories differently.
- Humans write by constantly referencing and recombining their own past words and structures, creating a tight, efficient web of meaning that grows in a specific mathematical way (Log-Power).
- AI writes by predicting the next word, which creates a different kind of growth pattern (closer to Power-Law). Even the smartest AI models haven't quite mastered the human art of "structural recycling" over long distances.
The author concludes that we need new tools to measure AI. We can't just ask, "Did it pass the test?" or "Is it fluent?" We need to ask, "Does it build its world the way a human does?" And right now, the answer is: Not quite.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.