DHPLT: large-scale multilingual diachronic corpora and word representations for semantic change modelling
This paper introduces DHPLT, an open resource comprising large-scale, multilingual diachronic corpora across 41 languages with pre-computed embeddings and lexical substitutions, designed to address the scarcity of data for semantic change modeling beyond high-resource languages.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you want to study how human language evolves, like watching a river change its course over decades. You need a massive collection of old letters, newspapers, and diaries to see how words like "mouse" (the animal vs. the computer device) or "cloud" (the sky vs. data storage) shifted their meanings.
The problem? For most of the world's languages, those "old letters" don't exist in a neat, organized library. They are scattered, messy, or simply non-existent.
Enter DHPLT. Think of DHPLT as a massive, time-traveling digital library built by researchers to fix this problem. Here is how it works, broken down simply:
1. The Source: The "Internet Time Machine"
Instead of waiting for librarians to organize history, the researchers went straight to the source: the Internet. They used a giant project called HPLT (High-Performance Language Technologies) that has already scraped billions of web pages.
- The Analogy: Imagine the internet is a giant, chaotic ocean. HPLT is a net that scoops up clean, filtered water.
- The Trick: To turn this ocean into a timeline, they didn't look at when the text was written (which is often hidden or missing). Instead, they looked at when the web crawler downloaded the page.
- If a bot downloaded a page in 2015, that page must have existed by 2015. It's like finding a time capsule: if you found it in 2015, the contents couldn't have been made after 2015.
2. The Collection: 41 Languages, 3 Time Capsules
The researchers didn't just grab random pages. They built three specific "time windows" for 41 different languages (from English and Spanish to Tamil and Georgian):
- Window 1 (2011–2015): The "Early Internet" era.
- Window 2 (2020–2021): The "Pandemic" era (when the world changed rapidly).
- Window 3 (2024–Present): The "AI Boom" era.
Each window contains 1 million documents per language. That's a lot of reading material!
3. The "Target Words": The Stars of the Show
You can't read a million books to find one word. So, the researchers picked a "Wanted List" of about 18,000 interesting words for each language (nouns, verbs, and adjectives).
They didn't just give you the text; they did the heavy lifting for you. They pre-calculated three different types of "meaning maps" for these words:
- Static Maps (Word2Vec): Like a dictionary definition that doesn't change based on context. It's the "average" meaning of a word in that era.
- Contextual Maps (Token Embeddings): Like a high-definition photo. It shows how the word was used in a specific sentence. Did "AI" mean a robot in a video game, or a chatbot? This map knows the difference.
- Substitutes: If you had to replace the word with another one in that sentence, what would you use? This helps track how the "vibe" of a word changed.
4. Why This Matters: Seeing the Shift
The paper shows that this tool works beautifully. They tested it with the word "AI":
- In 2011–2015: The word was hanging out with "video games," "robots," and "cars."
- In 2020–2021: It started hanging out with "chatbots" and "machine learning."
- In 2024: It's now surrounded by "ChatGPT," "generative," and "LLMs."
They saw the exact same shift in Spanish ("IA") and Russian ("ИИ"). It's like watching a movie in fast-forward, seeing exactly when a word's personality changed.
The Catch (Limitations)
The researchers are honest about the flaws. Because they used "download dates" instead of "writing dates," a document downloaded in 2024 might actually have been written in 2010.
- The Analogy: It's like finding a 2010 newspaper in a 2024 trash can. You know it exists in 2024, but it's not new.
- The Result: The timelines aren't perfectly sharp, but they are "good enough" to see the big picture of how language evolves.
The Bottom Line
DHPLT is a gift to the scientific community. Before this, studying how language changes was like trying to build a house with only a few bricks from a dozen different countries. Now, they have a warehouse full of bricks from 41 different cultures, sorted by time, with blueprints already drawn up. It allows researchers to finally ask: "How does the world speak differently today compared to ten years ago?" and actually get an answer.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.