Zipf's Law of Abbreviation in a Logographic Script: Coding-Theoretic Bounds on Chinese Character Stroke Counts
This study extends Zipf's law of abbreviation to logographic scripts by demonstrating that Chinese character stroke counts exhibit significant compression relative to optimal coding bounds, a pattern largely independent of script type that is further enhanced by historical simplification reforms and balanced against the structural necessity of two-dimensional stroke arrangement for character disambiguation.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Language is a system of trade-offs. On one side, there is the need to be understood clearly; on the other, the human desire to speak and write quickly. For over seventy years, linguists have observed a consistent pattern in how languages solve this problem: the words we use most often tend to be the shortest. This is known as the law of abbreviation. It appears in spoken sounds, in the length of written words, and even in the calls of animals. The idea is simple: if you say a word a million times a day, it pays to make it short. If you say it only once a year, it does not matter if it is long. But a deeper question has recently emerged. We know that languages follow this rule, but how close do they come to the absolute limit of efficiency? How much shorter could a language be if it were perfectly designed, and how much of that potential is left unused?
Until now, this question has been studied mostly in alphabetic languages like English or Spanish, where the cost of a word is measured by the number of letters it contains. A new study brings this investigation to a completely different kind of writing system: Chinese. In Chinese, the basic unit is not a letter but a character, and the cost of writing a character is measured by the number of brush or pen strokes required to draw it. Because Chinese characters are built from a fixed set of five basic stroke types, researchers can calculate the theoretical limit of efficiency with a precision that is impossible for alphabetic scripts. By analyzing the stroke counts of every character in the standard set against how often they are actually used, the study reveals that Chinese writing is remarkably efficient, yet it stops short of perfection for a very specific, structural reason.
The researchers began by gathering a massive amount of data. They took a database containing every single stroke order for the 20,902 standard Chinese characters and matched it with two huge collections of real-world text. One collection contained nearly 260 million characters from edited books and newspapers, while the other held over 190 million characters from social media posts. They counted the strokes for every character and calculated the average number of strokes needed to write the text people actually read. The results confirmed the expected pattern: the most common characters are indeed the simplest to write. The average character in the entire dictionary requires about 12.7 strokes to write, but the average character you encounter while reading a normal text requires only about 7.2 strokes. The most frequent hundred characters, which make up nearly 40 percent of all written text, average just six strokes each.
To understand how efficient this system really is, the team compared the real-world text to two theoretical extremes. The first was a random arrangement, where the most common words were assigned the longest, most complex characters by pure chance. The second was a perfect arrangement, where the most common words were assigned the shortest possible characters, and the rarest words got the longest ones. The actual Chinese writing system fell somewhere in between. It was far better than random, but it was not perfect. The study calculated an optimality score, which measures how close the system is to that perfect arrangement. For simplified Chinese, the score was 0.668, meaning the system is about two-thirds of the way toward the theoretical best. This figure is strikingly similar to the scores found for word lengths in dozens of other languages using alphabets and syllabaries, suggesting that there is a universal limit to how much any communication system can be compressed while remaining useful.
However, the study went further than just measuring efficiency. Because the Chinese writing system uses a closed set of five stroke types, the researchers could calculate the absolute mathematical limit of what is possible. They determined that if the system were a perfectly efficient code, the average character in running text would require only about 4.34 strokes. The fact that the real average is 7.22 strokes means the system is using about 1.66 times more strokes than the absolute minimum. In a typical alphabetic language, such a gap might be dismissed as simple inefficiency or a lack of planning. But the researchers found that this gap is not a mistake; it is a feature.
The reason the system cannot be compressed further lies in how the characters are built. If Chinese were a perfect one-dimensional code, like a string of letters where the order alone defines the word, then every character would have a unique sequence of strokes. But in Chinese, many different characters share the exact same sequence of strokes. For example, the strokes for "person," "enter," and "eight" are identical in sequence; they are distinguished only by the spatial arrangement of those strokes on the page. Similarly, the characters for "scholar" and "earth" use the same strokes in the same order but differ in the length of one horizontal line. Because the sequence of strokes is not unique, the system cannot be compressed to the mathematical limit without losing the ability to tell words apart. The "extra" strokes that seem redundant are actually necessary to create a two-dimensional structure that allows for visual components to be reused. This spatial arrangement makes the writing system transparent; a reader who knows the character for "water" can often guess the meaning of other characters that contain that same component, even if they have never seen them before.
The study also examined a major historical event to see if the system could be improved. In the mid-twentieth century, China underwent a simplification reform that changed the shape of thousands of characters, often reducing the number of strokes. The researchers treated this reform as a controlled experiment. They found that the reform successfully increased the efficiency of the writing system, raising the optimality score from 0.555 to 0.668. Crucially, the reform targeted the most frequent characters, shortening them by an average of two strokes, while leaving the rare, complex characters largely untouched. This is exactly what a perfect coding system would do: save the most effort on the forms used most often. The fact that the reform achieved this improvement suggests that the remaining gap between the current system and the theoretical limit is real and actionable, but it also shows that the vast majority of the compression had already happened naturally over centuries of use, long before any official reform.
Ultimately, the paper concludes that the Chinese writing system is not failing to reach perfection; it is solving a different problem. It sacrifices some compression to gain clarity and structure. The "inefficiency" of using more strokes than the mathematical minimum is the price paid for a system where meaning is built from recognizable parts arranged in space. The study confirms that while the law of abbreviation holds true, the limit of compression is not just a matter of counting strokes or letters. It is determined by the architecture of the script itself. The fact that Chinese, with its unique visual logic, lands in the same efficiency range as alphabetic languages suggests that all human communication systems are balancing the same fundamental pressures: the drive to be short and the need to be clear.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.