← Latest papers
💬 NLP

Entropy of Ukrainian

This paper presents the first study to estimate the entropy of the Ukrainian language by replicating Claude Shannon's 1951 character prediction experiment with 184 volunteers, yielding an upper bound of approximately 1.201 bits per character and comparing this result against current Large Language Models.

Original authors: Anton Lavreniuk, Mykyta Mudryi, Markiian Chaklosh

Published 2026-05-01
📖 5 min read🧠 Deep dive

Original authors: Anton Lavreniuk, Mykyta Mudryi, Markiian Chaklosh

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are playing a game of "Guess the Next Letter" with a friend. You show them a sentence that is cut off in the middle, like: "The cat sat on the..."

Your friend has to guess the next letter.

  • If they guess "m" (for mat), they are right. That was easy.
  • If they guess "z", they are wrong. They have to guess again.
  • If they guess "p" (for pillow), they are right.

This game measures something called Entropy. In the world of language, entropy isn't about heat or disorder; it's a measure of surprise.

  • High Entropy: The next letter is a total surprise. You have to guess many times before getting it right. The language feels chaotic and unpredictable.
  • Low Entropy: The next letter is obvious. You guess it immediately. The language feels predictable and repetitive.

The Experiment: Guessing Ukrainian

For decades, scientists have played this guessing game with English. In 1951, a man named Claude Shannon asked people to guess the next letter in English sentences. He found that English is quite predictable; you only need about 1.2 guesses on average to get the next letter right.

But nobody had ever played this game with Ukrainian until now.

Three researchers from Ukraine (Anton, Mykyta, and Markiian) decided to run this experiment for the first time. Here is how they did it, translated into simple terms:

1. The Players (The Volunteers)

Instead of hiring professional workers, they asked regular people on social media (mostly on Telegram) to play.

  • The Setup: They gave volunteers sentences from news articles.
  • The Twist: To keep people interested, they didn't start from the very beginning of the sentence. They showed the first 70 characters (about the length of a short sentence) and asked, "What comes next?"
  • The Game: The volunteer types a letter. If it's wrong, it turns red, and they try again. If it's right, they move to the next letter.
  • The Goal: They wanted to see how many "wrong guesses" it took before people got the right letter.

2. The Results: How Predictable is Ukrainian?

After collecting data from 184 volunteers who made over 17,000 guesses, the researchers crunched the numbers.

  • The Score: They found that the "surprise level" (entropy) of Ukrainian is approximately 1.201 bits per character.
  • What does that mean? It means Ukrainian is extremely similar to English in terms of predictability. Just like English, if you know the previous letters, the next letter in Ukrainian is usually quite obvious.
  • Redundancy: Because the language is so predictable, about 76% of the letters in a Ukrainian sentence are "redundant." You could actually remove a huge chunk of the text, and a native speaker could still guess the missing parts perfectly.

3. The "Cheating" Problem

The researchers had to be careful. Since this was an online game, some people might have:

  • Looked up the full sentence online.
  • Guessed randomly and got lucky.
  • Lost interest and stopped playing halfway through.

To fix this, they acted like a strict referee. They threw out the results of the people who played poorly (who might not have been trying) and the people who played too perfectly (who might have cheated). Even after being very strict and removing a lot of data, the result stayed around 1.20.

4. The AI Comparison: Humans vs. Robots

The researchers also asked modern AI models (Large Language Models) to play the same game.

  • The AI Advantage: The AI models were much better than the humans. Some of the best AI models could guess the next letter with a score of roughly 0.7.
  • The Catch: The researchers warn that the AI might be "cheating" in a different way. Because the AI was trained on massive amounts of news data (similar to the sentences used in the game), it might have memorized the specific sentences rather than truly understanding the language. It's like a student who memorized the answer key instead of learning the math.
  • The Conclusion: The AI is very good, but because it's so close to the "perfect" score, it's hard to tell if it's truly understanding Ukrainian or just overfitting to the specific news articles used.

The Big Takeaway

This paper is the first time we have a solid number for how predictable Ukrainian is.

  • Ukrainian is not chaotic. It is highly structured and predictable, just like English.
  • Humans are good at it. Regular people can guess the next letter with about 1.2 tries.
  • AI is better, but maybe too good. AI models can guess even faster, but we have to be careful not to confuse "memorization" with "intelligence."

The researchers admit their study wasn't perfect (they didn't have enough players, and the sentences were all from news articles), but it gives us a much better starting point than we had before. They have even shared their code and data so others can try to improve the game in the future.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →