← Latest papers
💻 computer science

GLeMM: A large-scale multilingual dataset for morphological research

The paper introduces GLeMM, a large-scale, fully automated, and multilingual derivational morphology dataset covering seven European languages, designed to enable data-driven research and computational testing of form-meaning relations in word formation.

Original authors: Nabil Hathout, Basilio Calderone, Fiammetta Namer, Franck Sajous

Published 2026-09-25
📖 6 min read🧠 Deep dive

Original authors: Nabil Hathout, Basilio Calderone, Fiammetta Namer, Franck Sajous

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Language is a living system where words constantly evolve, splitting and merging to create new meanings. When we take a word like "happy" and turn it into "happiness," or "teach" into "teacher," we are engaging in a process called derivation. This is how human beings expand their vocabulary, creating complex ideas from simple roots. For decades, linguists have tried to map these connections, asking why some words change in predictable ways while others seem to follow their own unique rules. They have wondered if the way a language builds new words is the same across different cultures, or if every language has its own secret logic. Until recently, answering these questions meant relying on the intuition of experts and small, hand-picked lists of examples. It was a bit like trying to understand the weather by watching a single cloud; you might get a sense of the pattern, but you miss the storm.

A team of researchers has now built a massive digital library to change how we study this process. They created a resource called GLeMM, which stands for a large-scale multilingual collection of word-formation data. Instead of relying on small lists, they gathered information on nearly 1.2 million pairs of related words across seven European languages: English, French, German, Italian, Polish, Russian, and Spanish. The goal was to move beyond guesswork and let the sheer volume of data reveal the true structure of how languages grow. By automating the process of finding these connections, the researchers could see patterns that were previously invisible, offering a clearer picture of how form and meaning interact when we create new words.

The researchers did not sit down to manually write down every word connection. Instead, they turned to Wiktionary, the free, online dictionary that anyone can edit. They realized that the definitions people write for words often contain clues about how those words are related. For instance, if the definition of "spryness" says it is "the property of being spry," the word "spry" is right there in the explanation. The team developed a computer method to scan millions of these definitions, looking for pairs of words where one is defined using the other. They then checked if the two words looked similar enough to be related, such as sharing a common root. By cross-referencing these findings with other lists of related words, they filtered out random matches and kept only the pairs that showed a consistent, repeating pattern. This allowed them to build a vast network of word families for each of the seven languages, capturing everything from simple suffixes like "-ness" to more complex changes involving prefixes and multiple steps.

What they found challenges some long-held ideas about how language works. For a long time, many linguists believed that word families were like complete webs, where every word in a family was directly connected to every other word. If you had a word for "travel," a word for "traveler," and a word for "traveling," the theory suggested they were all equally linked in a perfect triangle. However, the data from GLeMM suggests this is rarely the case. In the massive dataset, most word families are not perfect webs. Instead, they look more like a central hub with spokes, where new words branch out from a main root but do not necessarily connect back to each other. The researchers found that only a tiny fraction of word groups form these perfect triangles. In fact, for every 100 pairs of words that could theoretically form a triangle, only a handful actually do. This suggests that the way we build words is more directional and hierarchical than previously thought, with a clear flow from a base word to its derivatives, rather than a chaotic web of mutual connections.

The study also shed light on how different languages handle the same concepts. While all seven languages in the study create new words, they do it with different tools and frequencies. For example, the researchers looked at how languages express the idea of "doing something again." In French, they found a specific pattern where a prefix is added to a verb to reverse an action, creating a whole new family of words that mirrors the original. In English, similar patterns exist but are less uniform. The dataset showed that while the underlying logic of word formation is shared, the specific rules vary significantly. One language might prefer to combine two existing words to make a new one, while another prefers to add a small ending to a single word. The researchers also discovered that some word formations are "backwards." Sometimes, a shorter word is actually derived from a longer one, even though it looks like the shorter one should be the root. The data helped identify these reverse patterns by showing that they occur much less frequently than the standard direction, allowing the researchers to spot them as exceptions to the rule.

Despite the massive scale of the project, the researchers are careful to note that their resource is not perfect. Because the data was gathered automatically from a public dictionary, it contains some errors and inconsistencies. Some word pairs that look related might not be, and some definitions might be slightly off. However, the sheer size of the collection means that these errors are rare enough that they do not distort the big picture. The researchers estimate that the accuracy of the word pairs is very high, with over 90 percent of the connections in the English and French sections being correct. The resource is designed to be a starting point for further investigation, a place where other scientists can test their own theories against a vast amount of real-world data. It is not a final answer to the mystery of language, but it is a powerful new tool that allows us to see the landscape of word formation with a clarity that was impossible before.

The creation of GLeMM represents a shift in how we study language, moving from small, curated examples to massive, data-driven exploration. By treating the dictionary as a living corpus of millions of relationships, the researchers have opened a window into the hidden architecture of human speech. They have shown that while languages are diverse, they follow certain structural principles that can be measured and observed. The findings suggest that the way we build words is not a random collection of habits, but a structured system with clear patterns, exceptions, and a distinct direction. As the researchers plan to expand this work to include more languages and refine their methods, this resource promises to deepen our understanding of the human mind's ability to create meaning from sound. It reminds us that every new word we speak is part of a vast, interconnected history, and now, for the first time, we have a map large enough to see the whole territory.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →