Semantic Space of Parts of Speech
This paper challenges the traditional view of parts of speech as crisp categories by using word2vec embeddings and neural networks to map words into a three-dimensional semantic space, revealing the inherent fuzziness and boundary relationships between parts of speech across five languages.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Language is often taught as a system of neat boxes. We learn that a word is either a noun, a verb, or an adjective, and that it belongs to only one of these categories at a time. This crisp categorization helps us build dictionaries and teach grammar, but it does not always match the messy reality of how words actually work. Some words feel like they belong in two places at once, or they sit right on the border between categories, refusing to be pinned down. Linguists have long suspected that these boundaries are fuzzy, but proving it has been difficult because traditional methods rely on human judgment and rigid rules. Now, a team of researchers from Charles University in Prague and the University of Potsdam in Germany has used a different approach to map this uncertainty. They did not ask people to decide where words belong; instead, they let the words speak for themselves by analyzing how they behave in millions of sentences.
The researchers started with a simple idea: words that appear in similar situations tend to have similar meanings. This concept, known as distributional semantics, suggests that if you look at the company a word keeps, you can understand what it is. For example, a word that often appears next to "the" and "big" is likely a noun, while one that follows "will" or "can" is likely a verb. Computers can turn these patterns into mathematical maps, where every word is a point in a vast space. The closer two points are, the more similar the words are in how they are used. The team took this a step further by asking a computer to learn the rules of parts of speech on its own. They fed the computer a massive collection of text in five different languages: French, Czech, Finnish, Russian, and English. The computer was given the words and their standard labels, and it was asked to find the hidden patterns that connect them.
To make sense of the complex data, the researchers built a special kind of computer program, a neural network, designed to squeeze all the information down into just three dimensions. Imagine a room filled with thousands of floating dots, each representing a word. In the computer's mind, these dots exist in a space with hundreds of directions, which is impossible for a human to see. The researchers forced the computer to flatten this space into a shape we can actually visualize: a three-dimensional map. On this map, the computer placed the words based on how similar they are in meaning and usage. The result was not a list of rules, but a landscape. In this landscape, the most common types of words formed distinct clusters, or "tentacles," that stretched out from a central point.
When the researchers looked at these maps, they saw that the traditional categories of nouns, verbs, and adjectives were indeed the strongest groups. In every language they studied, these three groups formed clear, separate shapes. However, the space between them was not empty. The researchers found that many words did not sit strictly inside one tentacle but rather in the space where the shapes touched or overlapped. For instance, in French, words that can be both verbs and adjectives, like "forgotten," were found right on the border between the verb and adjective groups. In English, words that can be both nouns and verbs, such as "laugh" or "care," were located where the noun and verb tentacles blended together. This confirmed that the fuzziness linguists suspected is real and measurable. The words that cause the most confusion for traditional grammar are the ones that naturally live in these transition zones.
The study also revealed surprising details about specific languages. In Russian, the researchers noticed that short forms of adjectives, which are used to describe a state like "the man is young," clustered closer to verbs than to other adjectives. This makes sense because, in Russian grammar, these short forms often function as the main action of a sentence. In Finnish, a language with a very different structure from the others, the map showed that adverbs did not form their own distinct group but were scattered among the other categories. This suggests that in Finnish, the line between an adverb and other parts of speech is much less clear than in English or French. Even in languages where the rules seem strict, the data showed that words often drift depending on how they are used.
One of the most striking findings was that numerals, or numbers, often formed their own separate group, even though traditional grammar sometimes treats them as a type of noun or adjective. In the maps for French, Czech, and Russian, the numbers stood apart from the main clusters, suggesting that they have a unique semantic identity that sets them apart from other words. This happened even though the computer was not told to look for numbers; it discovered this pattern purely by analyzing how the words were used in real text. The researchers also found that the way words are grouped can change depending on the specific dataset, showing that the boundaries are not fixed but fluid.
The researchers did not set out to create a new grammar book or to prove that one language is better than another. Their goal was to see what happens when we let the data decide the categories. They found that while the main categories of nouns, verbs, and adjectives are strong and stable, the edges are soft. Words often belong to multiple categories at once, or they shift depending on the context. The study suggests that the rigid boxes we use to teach grammar are useful tools, but they do not capture the full picture of how language works. The true nature of a word is not a single label, but a position in a vast, shifting space where it can be close to many different ideas.
By visualizing these relationships, the researchers provided a new way to see the structure of language. Instead of arguing about whether a word is a noun or a verb, we can now see exactly where it sits in relation to both. This approach does not replace the old rules, but it adds a layer of depth that shows the complexity of human communication. The maps show that language is not a collection of isolated islands, but a continuous landscape where words flow into one another. The study confirms that the fuzziness of language is not a mistake or a flaw, but a fundamental feature of how we use words to describe the world.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.