← Latest papers
💬 NLP

Complete asymptotic type-token relationship for growing complex systems with inverse power-law count rankings

This paper presents a deterministic, idealized model that derives a unified asymptotic expression for the type-token relationship in growing complex systems, demonstrating that Heaps' law emerges solely as a mathematical consequence of Zipf's law across all power-law exponents without relying on stochastic mechanisms.

Original authors: Pablo Rosillo-Rodes, Laurent Hébert-Dufresne, Peter Sheridan Dodds

Published 2026-03-31
📖 5 min read🧠 Deep dive

Original authors: Pablo Rosillo-Rodes, Laurent Hébert-Dufresne, Peter Sheridan Dodds

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Picture: Two Famous Rules of Growth

Imagine you are building a massive library. As you add more and more books (tokens), two famous rules usually describe what happens:

  1. Zipf's Law (The Popularity Contest): In any large collection of words, species, or ideas, a few things are super popular, and most things are rare. If you rank them from most common to least common, the popularity drops off like a slide. The #1 word appears twice as often as the #2, three times as often as the #3, and so on. This is the "Inverse Power-Law."
  2. Heaps' Law (The Vocabulary Growth): As you read more and more text (add more books to your library), you discover new words (types). But you don't discover them at a constant speed. At first, you find new words quickly. Later, you mostly just see words you've already read. The number of unique words grows slower than the total number of words.

The Question: For a long time, scientists wondered: Are these two rules just happening to exist together, or does one actually cause the other?

The Old Answer vs. The New Answer

A previous study (by Lü et al.) tried to connect these two rules. They built a mathematical bridge between them, but their bridge had some cracks. It worked well for some situations but fell apart when the "popularity slide" was very steep (when the most popular items were extremely dominant) or when it was very flat.

This new paper by Rosillo-Rodes, Hébert-Dufresne, and Dodds fixes the bridge.

They created a perfect, idealized model of a growing system. They didn't need to guess about random chance or complex biological mechanisms. They simply asked: "If we strictly follow the rule that popularity drops off like a power law, what must happen to the total number of unique items?"

The Analogy: The "Infinite Buffet"

Imagine a giant buffet where new dishes (types) are added every day.

  • The Rule: The most popular dish (say, Pizza) is served to 100 people. The second most popular (Burgers) is served to 50. The third (Salad) to 33, and so on. This is Zipf's Law.
  • The Goal: We want to know how many unique dishes are on the menu after 1,000 people have eaten (the total tokens).

The Old Math Mistake:
Previous scientists tried to estimate the total number of dishes by drawing a smooth curve under the "popularity slide."

  • The Problem: When the slide is very steep (like a cliff), the curve misses the first few steps. It's like trying to measure a staircase by looking at the shadow it casts; you miss the actual steps. This caused them to underestimate how many people it takes to see a certain number of dishes, especially when one or two dishes are overwhelmingly popular.

The New Solution:
The authors used a more precise mathematical tool (called the Euler-Maclaurin expansion).

  • The Analogy: Instead of looking at the shadow, they counted every single step of the staircase. They realized that for very steep slides, the "first step" (the most popular item) is so big that it dominates the whole system. Their new formula accounts for this "step" perfectly.

What They Found

They derived a single, unified formula that works for every type of system, whether the popularity drop is gentle or a steep cliff:

  1. When popularity drops slowly (Flat slide): The number of unique items grows almost linearly with the total size.
  2. When popularity drops at the "standard" rate (The famous Zipf's Law where α=1\alpha = 1): The growth follows a specific logarithmic pattern (it slows down, but in a predictable way involving a famous constant called the Euler-Mascheroni constant).
  3. When popularity drops very fast (Steep cliff): The system is dominated by just a few items. The number of unique items grows very slowly, like the root of the total size.

Why This Matters

The most exciting part of this paper is the conclusion: Heaps' Law is not a separate magic rule.

It is simply a side effect of Zipf's Law.

Think of it like this: If you have a crowd where a few people are shouting very loudly and everyone else is whispering (Zipf's Law), you will naturally find that as the crowd gets bigger, you hear fewer new voices (Heaps' Law). You don't need to invent a special "whispering mechanism" to explain it; the shouting pattern explains it all.

The Takeaway

  • Old View: Zipf's Law and Heaps' Law are two separate friends who happen to hang out together.
  • New View: Heaps' Law is just Zipf's Law wearing a different hat. If you know how the popularity ranks work, you can mathematically predict exactly how the vocabulary will grow, without needing to assume any random luck or complex biology.

The authors have provided the "master key" that unlocks the relationship between how things are ranked and how new things appear in any growing system, from the words in a novel to the species in a forest.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →