← Latest papers
💬 NLP

Losing My Composure: Predicting Compositionality Over Time

This paper introduces a novel dataset and task for predicting the diachronic trends of noun compound compositionality in German and English, revealing only a slight decrease in transparency over time and demonstrating that semantic models trained on narrow, temporally specific data slices outperform those trained on broad historical windows in capturing these gradual changes.

Original authors: Chris Jenkins, Emma Raimundo Schulz, Filip Miletić, Sabine Schulte im Walde

Published 2026-07-14
📖 5 min read🧠 Deep dive

Original authors: Chris Jenkins, Emma Raimundo Schulz, Filip Miletić, Sabine Schulte im Walde

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine language as a giant, bustling city where words are like buildings. Some buildings are simple shacks made of one brick; others are complex skyscrapers built by snapping two smaller buildings together, like a "coffee house" or a "freight train." For a long time, linguists had a hunch about these skyscrapers: they thought that as the city got older, these buildings would slowly lose their "blueprint." They believed that over centuries, a "coffee house" would stop looking like a house for coffee and start becoming a weird, mysterious unit that you just have to memorize, losing its connection to the words "coffee" and "house."

This paper is like a team of time-traveling detectives who decided to test that hunch. They didn't just look at two snapshots of the city (one from the past, one from today); instead, they built a high-speed camera that took pictures every single decade for hundreds of years. They focused on 23 German and 26 English compound words, like "bosom friend" or "Kaffeehaus," and asked human volunteers to rate, on a scale of 0 to 5, how much the meaning of the whole word still matched the meanings of its parts in each specific decade.

The Big Surprise: The Blueprint Stays Put
The detectives found that the old hunch was mostly wrong. They didn't find a massive, city-wide trend where these word-buildings were crumbling into mystery. In fact, the changes were so tiny they were almost invisible. If you measured the "transparency" of these words over time, the average change was a microscopic negative blip (about -0.016 for English and -0.001 for German). It's like checking the height of a skyscraper every ten years and finding it hasn't grown or shrunk by even a single grain of sand. The paper suggests that, contrary to popular belief, these compound words generally keep their "construction plans" intact; they don't necessarily become less logical or more idiomatic just because time passes.

The Time-Traveling Cameras
To make these measurements, the team built about 100 different "cameras" (computer models) to simulate how a word is understood in different eras. They tried two main types of cameras:

  1. The "Context-Aware" Cameras (like Modern BERT): These are fancy, high-tech lenses that look at the words surrounding a target to understand its meaning, kind of like how you understand the word "bank" differently if it's next to "river" or "money."
  2. The "Static" Cameras (like Word2Vec): These are simpler lenses that give a word one single, fixed definition based on its history, like a dictionary entry that never changes.

They also tested how much "historical data" to feed into these cameras. Some were fed a massive 50-year chunk of history (a "full" window), while others were fed just one decade at a time, or a sliding window that grew one decade at a time.

The Results: Less is More
Here is where the plot thickens. The team discovered that the "fancy" context-aware cameras didn't necessarily win the race. In fact, the simpler, static cameras performed just as well, and sometimes even better, at predicting how compositionality changed over time.

More importantly, the paper argues that the size of the history you feed the camera matters a lot. When they trained the models on a massive 50-year block of text (the "full" schedule), the models got confused and missed the subtle, decade-by-decade shifts. However, when they trained the models on narrow, single-decade slices (the "decades" schedule), the models aligned much better with the human ratings. It's as if trying to understand the mood of a city by reading a 50-year-old newspaper archive is less effective than reading the daily paper from exactly the day you are interested in. The paper suggests that for tracking these tiny, gradual changes, you need to zoom in on the specific decade, not blur it together with the past.

The Verdict
So, what's the final word? The paper doesn't claim to have solved the mystery of language forever. Instead, it suggests that the idea of compounds inevitably becoming "less compositional" over time is likely an exaggeration. The changes are there, but they are small and messy, not a straight line down.

The study also warns us that while fancy AI models are cool, they aren't always the best tool for historical detective work. Sometimes, a simpler model trained on a very specific slice of time (just one decade) is more accurate than a super-complex model trained on half a century of data. The authors measured these results using statistical scores (like R² values around 0.36 for the best English models), which shows a decent but not perfect match between the computer's guesses and human ratings.

In the end, the paper is a reminder that language is a journey, but for these specific word-buildings, the journey hasn't been a dramatic collapse into confusion. They've mostly stayed true to their original blueprints, decade after decade, proving that sometimes, the simplest way to look at the past is to look at it one year at a time.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →