← Latest papers
💬 NLP

Methods, Data, and Conceptual Change: Reflections from Two Quantitative Diachronic Case Studies

This paper reflects on the interaction between quantitative methods and dataset properties in historical linguistics by comparing quad-based concept modelling of Early Modern English discourse with SynFlow analysis of scientific writing, ultimately arguing that methodological choices and data structure significantly shape the detection and interpretation of semantic change.

Original authors: Catherine Wong, Bach Phan-Tat, Susan Fitzmaurice

Published 2026-05-05
📖 4 min read☕ Coffee break read

Original authors: Catherine Wong, Bach Phan-Tat, Susan Fitzmaurice

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to understand how the meaning of words has changed over the last few hundred years. You have two massive libraries of old books, but they are very different from each other. One library is a chaotic, noisy bazaar filled with every kind of pamphlet, sermon, and story from the 1500s to the 1600s. The other is a quiet, orderly laboratory filled with precise scientific reports from the 1700s.

This paper argues that how you choose to look at these libraries changes what you see. It's not just about the words themselves; it's about the "lens" you use and the "room" you are looking into.

Here is a simple breakdown of the paper's two main experiments and what they teach us.

The Two Libraries (The Data)

  1. The Chaotic Bazaar (EEBO-TCP): This is a huge collection of Early Modern English texts. It's messy. People spelled words differently (like "shew" vs. "show"), the genres are all mixed up, and the grammar isn't always perfect. It's like trying to find a pattern in a pile of mixed-up puzzle pieces from different boxes.
  2. The Orderly Lab (Royal Society Corpus): This is a collection of scientific writing from the 1700s. It is much more uniform. The spelling is consistent, the sentences follow strict rules, and the topics are focused on science. It's like a neatly organized filing cabinet where every document looks similar.

The Two Lenses (The Methods)

The authors tried two different ways to find "conceptual change" (how ideas evolve) in these libraries.

Lens 1: The "Word Neighborhood" Map (Quad-based Modelling)

  • Used on: The Chaotic Bazaar.
  • How it works: Because the text is so messy, you can't rely on strict grammar rules. Instead, this method looks at groups of four words that frequently hang out together in a 100-word window.
  • The Analogy: Imagine you are trying to understand a person's personality by seeing who they hang out with at a party. You don't need to know exactly what they said to each other; you just need to know that "Liberty" always shows up with "King," "Law," and "Court," while "Freedom" always shows up with "God," "Soul," and "Prayer."
  • What they found: Even though "Liberty" and "Freedom" sound like synonyms, their "word neighborhoods" were totally different. The method showed that these words belonged to different conceptual worlds, even in the messy text.

Lens 2: The "Grammar Skeleton" Scan (SynFlow)

  • Used on: The Orderly Lab.
  • How it works: Because the scientific text is so structured, this method looks at the grammatical relationships between words. It asks: "Is this word acting as a subject? Is it being modified by an adjective?"
  • The Analogy: Imagine you are looking at a skeleton. You aren't just looking at the bones (the words); you are looking at how they are connected. You can see how the "concept of Air" changed by looking at what adjectives were glued to it.
  • What they found: In the 1700s, scientists started describing "Air" not just as a general thing, but as specific types of air (like "acid air" or "fixed air"). The method caught this shift by seeing that the words attached to "Air" became much more specific and scientific over time.

The Big Lesson: The Map is Not the Territory

The paper's main point is that you cannot separate the tool from the data.

  • If you try to use the "Grammar Skeleton" scan on the "Chaotic Bazaar," it will fail. The text is too messy; the grammar is too broken to build a reliable skeleton.
  • If you try to use the "Word Neighborhood" map on the "Orderly Lab," you might miss the fine details. You'd see that words hang out together, but you wouldn't see the precise grammatical shifts that define scientific progress.

The Takeaway:
Quantitative research isn't just about counting words. It's about realizing that your method shapes the reality you see.

  • One method sees change as a shift in who words hang out with (good for messy, diverse history).
  • The other sees change as a shift in how words are structured (good for clean, specialized history).

The authors conclude that we need to be honest about these limitations. We can't just say "we found a change in meaning." We have to say, "We found a change in meaning using this specific tool on this specific type of text." It's like saying, "I saw a fish swimming," rather than just "I saw a fish," because if you were looking through a telescope, you might have seen a bird instead. The tool dictates what you can see.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →