← Latest papers
💬 NLP

Three Buddhist Vocabularies: Computational Stylometry of the English Pali Canon across Sutta, Vinaya, and Abhidhamma

This paper presents a computational stylometric analysis of the English Pali Canon across the Sutta, Vinaya, and Abhidhamma Pitakas, revealing distinct lexical profiles for each corpus—such as the Abhidhamma's higher diversity and numeral density, the Theravada and Mulasarvastivada Vinayas' shared legal heritage, and significant translation variations over time—while confirming that all texts adhere to Zipf's law.

Original authors: Joy Bose

Published 2026-06-25
📖 5 min read🧠 Deep dive

Original authors: Joy Bose

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the Buddhist Canon (the Tipitaka) as a massive library containing three distinct wings, each written by different authors for different purposes. For a long time, researchers have studied the "Sutta" wing (the stories and sermons). This paper is like a new librarian who decides to walk into the other two wings—the Vinaya (the rulebook for monks) and the Abhidhamma (a technical encyclopedia of the mind)—and run a computer program to see how the language itself changes from room to room.

The researcher, Joy Bose, didn't read every word to understand the deep philosophy. Instead, they used a "word counter" and a "pattern detector" to see how the vocabulary behaves. Here is what they found, explained simply:

1. The Three "Flavors" of Language

Think of the three sections as three different types of parties:

  • The Sutta Wing (The Storytellers): This is the biggest room. The language here is like a campfire story. It repeats the same phrases often (like "the monk said" or "four things"). The computer found that the words here are very repetitive, creating a steady, rhythmic beat.
  • The Vinaya Wing (The Rulebook): This room is full of legal instructions. It's surprisingly similar to the Storytellers in how repetitive it is. Whether you read the modern translation (Brahmali) or the old one (Horner), the "vocabulary rhythm" is almost identical. It's like two different people reading the same phone book; the words might change slightly, but the structure is the same.
  • The Abhidhamma Wing (The Encyclopedia): This is the odd one out. It's a technical manual listing 52 types of feelings and 89 types of consciousness. Because it has to name so many specific, unique things, it uses a much wider variety of words. It's less repetitive and feels more like a modern science textbook than a story.

2. The "Zipf" Test: Who is the Star of the Show?

The researchers used a famous math rule called Zipf's Law. Imagine a rock concert where one singer sings 50% of the songs, the second singer sings 25%, and the rest split the remaining time. In most texts, a few common words (like "the," "and," "it") dominate the top spots.

  • In the Stories and Rules: The top 10 words are mostly grammar helpers (like "it" or "they").
  • In the Encyclopedia: The computer found something weird. The word "Consciousness" jumped into the top 10, pushing the grammar words down. It's as if, in a library of technical manuals, the word "Engine" appears so often it drowns out the word "The." This proves the Encyclopedia is obsessed with specific concepts, not just storytelling.

3. The "Vocabulary Overlap" (How much do they share?)

The researchers asked: If you take a handful of words from the Rulebook, how many of them would you also find in the Storybook?

  • The Result: About half of the words in the Rulebook are also in the Storybook. They share a common language because they are part of the same tradition.
  • The Time Travel Test: The researchers compared the Theravada Rulebook (used in Sri Lanka/Thailand) with the Mulasarvastivada Rulebook (used in Tibet) from 2,000 years ago. Even though they split up two millennia ago, they still share 50% of their vocabulary. It's like finding that a modern American English dictionary and a 19th-century British one still share half their words, proving they came from the same family tree.

4. The "Translation Shift" (How language changes over time)

The researchers took the exact same Rulebook text and compared two translations: one from 1938 and one from 2026.

  • The Shock: They only shared 24% of their vocabulary! That's a huge gap for the same text.
  • The Cause: It wasn't that the meaning changed, but the words changed.
    • Example 1: In 1938, a deep meditation state was called "musing" (which sounds a bit like daydreaming). In 2026, it's called "absorption" (which sounds much more focused).
    • Example 2: A serious offense was called "defeat" in 1938, but "expulsion" in 2026.
    • Example 3: The 2026 translator was more direct about anatomy and gender issues, while the 1938 translator used polite, vague language.

5. The "Number" Count

The Encyclopedia (Abhidhamma) is obsessed with counting. It uses number words (one, two, three, four) 3.26% of the time. The Storybook uses them only 2.09% of the time. This makes sense: if you are listing 89 types of consciousness, you have to say "one," "two," "three" a lot more than if you are just telling a story about a monk walking in the woods.

Summary

This paper is like a linguistic X-ray. It doesn't tell us what the Buddha taught in deep detail, but it shows us how the teaching was packaged.

  • Stories are rhythmic and repetitive.
  • Rules are procedural and shared across time.
  • Encyclopedias are dense, technical, and full of unique terms.
  • Translations change drastically over 88 years, swapping old-fashioned words for modern, precise ones.

The researcher concludes that while we can see these patterns clearly with computers, the deeper question of how these ideas evolved over time requires a more human, concept-level analysis, which is the next step in their work.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →