← Latest papers
🧬 biology

Beyond cognacy

This paper demonstrates that multiple sequence alignment (MSA) derived from a pair-hidden Markov model offers a scalable, fully automated alternative to traditional expert-annotated cognate sets for constructing language phylogenies, yielding trees that are more consistent with established linguistic classifications and better at predicting typological variation.

Original authors: Gerhard Jäger

Published 2026-03-30
📖 5 min read🧠 Deep dive

Original authors: Gerhard Jäger

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine you are trying to draw a family tree for the entire human race, but instead of people, you are trying to map out how thousands of different languages are related to one another. This is the job of historical linguists.

For a long time, doing this was like trying to solve a massive jigsaw puzzle in the dark. Experts had to manually look at thousands of words, decide which ones were "cousins" (words that sound similar because they come from the same ancient ancestor), and then build the tree. This was slow, expensive, and limited to small groups of languages.

This paper is about a new way to solve that puzzle using computers, and it asks a simple question: Can we build these language family trees automatically, without needing a human expert to label every single word?

Here is the breakdown of the study, using some everyday analogies.

The Three Competitors

The author, Gerhard Jäger, set up a race between three different methods to build these language trees. Think of them as three different detectives trying to solve the same case.

  1. The Old School Detective (Expert Cognates):

    • How it works: This method relies on human experts who have already spent years labeling words as "cousins." It's like having a family reunion where everyone already knows who their relatives are.
    • The Problem: It's incredibly slow to get this data. You can only do it for small families, and it's hard to check if the experts made mistakes. It's like trying to map the whole world using only a few local maps.
  2. The Pattern Matcher (Automatic Clustering/PMI):

    • How it works: This computer method looks at words and tries to group them based on how often certain sounds appear together. It's like a computer looking at a bag of mixed-up socks and saying, "These red ones with stripes probably go together," without knowing what a sock actually is.
    • The Result: It's faster than the human method, but it's a bit fuzzy. It misses some of the deeper connections.
  3. The DNA Sequencer (Multiple Sequence Alignment / MSA):

    • How it works: This is the star of the show. Borrowing a technique from biology (where scientists align DNA strands to find mutations), this method lines up words from different languages like sentences in a book. It looks for patterns in how sounds change over time, even if the words aren't perfect matches.
    • The Analogy: Imagine you have three versions of a recipe for "soup." One says "add salt," one says "add a pinch of salt," and one says "add salt and pepper." Even though they aren't identical, a smart computer can see they are all trying to describe the same dish and how the instructions evolved. This method does that with thousands of languages at once.

The Big Test

The author didn't just guess which method was best; he put them to the test in three ways:

  1. The "Gold Standard" Check: He compared the computer-generated trees against the "official" family trees created by human experts (called Glottolog).

    • Result: The MSA method (the DNA sequencer) drew a map that looked most like the expert's map. It made the fewest mistakes.
  2. The "Typology" Check: He checked if the trees matched up with other facts about the languages, like grammar rules (e.g., "Does this language put the verb at the end of the sentence?").

    • Result: The MSA method was the best at predicting these grammar rules, suggesting its trees were the most accurate.
  3. The "Signal Strength" Check: He measured how much "noise" was in the data.

    • Result: The MSA method found the clearest signal. It was like listening to a radio station: the other methods had a lot of static, but the MSA method had a clear, strong broadcast.

The Verdict: Why This Matters

The study found that while the "Old School" expert method is still good for small, specific families, the MSA method is the future for looking at the big picture.

  • Scalability: You can't hire enough experts to label every language in the world. But you can write a computer program to do it.
  • Global Reach: The MSA method worked amazingly well when looking at languages from many different families mixed together. The other methods got confused and made mistakes when the data got too big.
  • No More Bottlenecks: This opens the door to creating a "Tree of Life" for all human languages, not just the ones we have already studied.

The Bottom Line

Think of this paper as the moment linguistics moved from hand-drawing maps to using GPS.

The old way (experts labeling words) is like walking through a forest with a compass and a notebook. It's accurate for the path you are on, but you can't see the whole forest. The new way (MSA) is like getting a satellite image. It sees the whole forest at once, spots the hidden trails, and gives you a much clearer picture of how everything is connected, all without needing a human to walk every single step.

This means we can finally start answering big questions about how human language evolved across the entire globe, something that was previously impossible due to the sheer amount of work required.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →