← Latest papers
💬 NLP

An Information-Theoretic Approach to Identifying Formulaic Clusters in Textual Data

This paper presents an information-theoretic algorithm utilizing weighted self-information distributions to identify formulaic clusters and stylistic layers in high-dimensional textual data, demonstrating its effectiveness in quantitatively analyzing compositional patterns within the multi-author Hebrew Bible.

Original authors: Gideon Yoffe, Yair Segev, Barak Sober

Published 2026-07-29
📖 7 min read🧠 Deep dive

Original authors: Gideon Yoffe, Yair Segev, Barak Sober

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery inside a massive, ancient library. This library isn't just full of random stories; it's a collection of texts that have been written, edited, and copied by hundreds of different people over thousands of years. Some parts were written by a single author with a unique voice, while others were stitched together from different sources, like a patchwork quilt. The challenge is: how do you figure out which patch belongs to which tailor without asking the tailors? You can't just look at the words because some writers might use the same words for different reasons. You need a way to measure the "rhythm" or the "predictability" of the writing.

This is where a branch of science called Information Theory comes in. Think of it as a way to measure how surprised you are when you read the next word in a sentence. If you are reading a story where anything can happen, you are constantly surprised; the information content is high, and the writing feels "loose" or "creative." But if you are reading a recipe or a legal contract, you know exactly what word comes next. "Add one cup of..." is almost always followed by "flour." That predictability means the information content is low because the pattern is rigid and repetitive. In this paper, the authors use this concept of "surprise" (or lack thereof) to find hidden patterns in ancient texts. They want to see if they can automatically spot the "recipe-like" sections of a book versus the "story-like" sections, just by measuring how predictable the words are.

The Detective's New Tool

The authors, Gideon Yoffe, Yair Segev, and Barak Sober, have built a new digital magnifying glass. They call it an unsupervised information-theoretic approach. Let's break that down: "Unsupervised" means the computer doesn't need a teacher to tell it what to look for; it figures it out on its own. "Information-theoretic" means it uses math to measure predictability.

Their main idea is simple but powerful: Formulaic text is predictable. When a writer uses a rigid structure, like a list of ancestors or a set of religious laws, they repeat the same phrases over and over. Because these phrases appear so often, they become very predictable. In the language of information theory, predictable things have low self-information (low surprise). Unpredictable, creative storytelling has high self-information (high surprise).

The team developed an algorithm that scans through a text, looking for chunks of writing that are unusually predictable. It doesn't just guess; it calculates a score for every part of the text based on how "surprised" a probability model would be by the words it sees. If a section is full of repetitive, formulaic phrases, the model isn't surprised at all, and the score stays low. If the section is a wild, creative narrative, the score goes up. By grouping the low-score sections together, the algorithm can isolate the "formulaic clusters"—the parts of the text that feel like they were written by a different hand or for a different purpose.

Testing the Tool on Ancient Texts

To see if their new tool actually works, the authors tested it on the Hebrew Bible, specifically the first three books: Genesis, Exodus, and Leviticus. These books are famous among scholars for being "composite," meaning they are believed to be made up of different sources woven together. One of the most debated questions is how to separate the Priestly source (P)—which is known for being very structured, legalistic, and repetitive—from the other, more narrative parts of the text.

The researchers didn't just look at the words; they looked at the structure of the words. They broke the text down into small chunks (like sentences or verses) and counted how often certain patterns appeared. They then ran their algorithm to see if it could naturally separate the text into two groups: one group that was highly repetitive (low surprise) and one that was more varied (high surprise).

Here is what they found:

  • In Genesis: They looked at the difference between the long, boring lists of family trees (genealogies) and the exciting stories of Abraham and Isaac. The algorithm successfully identified the genealogies as the highly predictable, formulaic cluster. This makes sense because family lists follow a strict, repetitive pattern, whereas stories are more fluid.
  • In Exodus: They tested the separation between the Priestly (P) sections (which contain laws about building the Tabernacle and religious rituals) and the non-Priestly narrative sections. The algorithm reliably separated these two layers, matching what traditional scholars have long believed. The Priestly parts showed up as the "low surprise" zone because they are full of repetitive instructions.
  • In Leviticus: This was the trickiest test. Leviticus is full of laws, but scholars debate whether it contains two different layers: the Priestly source (P) and a later "Holiness Code" (H). Both are repetitive, but they have subtle differences. The authors found that their method could distinguish between them, but only when they looked at the text through the right "lens." Depending on how they sized the chunks of text they analyzed, the algorithm sometimes highlighted the Priestly rules as the most repetitive, and other times it highlighted the Holiness Code. This suggests that both are formulaic, but they follow different kinds of patterns that show up at different scales.

How Sure Are They?

The authors are careful not to claim they have solved the mystery of the Bible once and for all. Instead, they show that their method suggests a way to quantify these differences.

They tested their algorithm on fake, computer-generated data first to make sure it could find patterns in a controlled environment. In those simulations, their method often did a better job than standard clustering tools (like k-means or Gaussian Mixture Models), especially when the data was sparse and high-dimensional (which is exactly what ancient texts are like).

When they applied it to the real Bible, the results were promising but nuanced. In the book of Exodus, their method agreed with expert scholars' classifications in 68% of the different parameter settings they tried, compared to only 30% for the standard k-means method. In Genesis, it agreed in 40% of cases versus 14.6% for k-means. In Leviticus, it agreed in 45.5% of cases versus 29.8% for k-means.

These numbers suggest that the information-theoretic approach is a strong tool for finding these hidden layers, but it's not a magic wand that works perfectly every time. The fact that the results change depending on how you slice the text (the size of the word groups and the window of verses) tells us that the "formulaic" nature of these texts is complex. It's not just one thing; it's a mix of patterns that appear at different scales.

Why This Matters

This paper doesn't just give us a new way to sort books; it offers a new way to listen to them. By using math to measure "surprise," the authors provide an objective way to see where the rhythm of a text changes. If a story suddenly becomes very predictable, it might be a sign that a different author picked up the pen, or that the text was copied from a template.

The authors conclude that this method is particularly useful for texts that have been edited and layered over time, like the Hebrew Bible. It allows researchers to see the "structural scaffolding" of a text without needing to guess which words are important beforehand. While the method doesn't prove exactly who wrote what, it provides a quantitative framework that supports many of the traditional theories about how these ancient books were put together. It turns the art of literary analysis into a science of patterns, showing us that even in the most complex stories, the fingerprints of structure and repetition can be found if you know how to measure the surprise.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →