Constructing Parallel Multidimensional Chromatic Lexicons for Corpus-Assisted Analysis of Russian and English Texts
This paper addresses the scarcity of tools for corpus-assisted analysis of color terms by developing multidimensional Russian and English chromatic lexicons and demonstrating their utility through a comparative pilot study of Andrei Bely and Emily Dickinson's poetry, which revealed significant cross-linguistic differences in the frequency and characteristics of color usage.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine language as a giant, bustling city where every word is a building. Most of the time, we just walk past the buildings without thinking about their colors. But what if we could paint the entire city? What if we could measure exactly how often the "Red" buildings appear compared to the "Blue" ones, or how "bright" the neighborhood feels? This is the world of corpus linguistics, a field where researchers use massive digital libraries of text to spot patterns that the human eye might miss. At the heart of this study is the idea that color isn't just a visual thing; it's a language tool. Just as a painter chooses specific shades to make a picture feel warm or cold, writers choose specific color words to make a reader feel happy, sad, or excited. But here's the tricky part: different languages have different "color palettes." Russian and English, for instance, don't just have different words for the same colors; they have different amounts of words and different ways of building them. This paper asks a big question: How can we build a fair, side-by-side map of these two different color worlds so we can compare them without getting lost?
The researcher, Larisa Nikitina, decided to build two custom "color dictionaries" (lexicons) to solve this puzzle. Think of these not as simple lists, but as high-tech, multi-layered filters. Instead of just listing "red" or "blue," these dictionaries sort every color word into three specific categories: Hue (the actual color, like red or green), Saturation (how intense or muddy the color is, like a bright fire-engine red versus a dusty rose), and Temperature (whether the color feels warm like a sunset or cool like ice). They also included "visual descriptors," which are words that describe how light or dark something looks, even if they aren't a specific color name.
To make these dictionaries work, the team had to be very careful. They couldn't just use Google Translate, because that would be like trying to match a Russian snowflake to an English snowflake by only looking at the word "snow." They had to manually check every single entry. For example, they realized that in English, the word "citrine" can mean both a gemstone and a yellowish color, but in Russian, the equivalent word only means the gemstone. If they hadn't checked this, their computer might have counted a ruby ring as a "red" color word in Russian, which would be wrong. They also had to deal with Russian's tricky grammar, where one word stem can spawn dozens of variations, making sure they didn't accidentally count "egg yolk" just because it shares a root with the word for "yellow."
Once the dictionaries were built—one with 224 entries for Russian and one with 141 for English—the team put them to the test. They treated the dictionaries like a pair of high-powered binoculars and scanned the poetry of two very different writers: Andrei Bely, a Russian Symbolist poet, and Emily Dickinson, an American poet from the 19th century. They wanted to see what these two poets' "color fingerprints" looked like.
The results were striking. When they looked strictly at words that definitely referred to color, they found that Bely used color terms 3.4 times more often than Dickinson did. In the specific sample of poems they analyzed, Bely's text had 304 confirmed color mentions, while Dickinson's had 125. It wasn't just that Bely used more colors; he used them differently. He leaned heavily on "low saturation" colors (muted, greyed-out tones) and used visual descriptors like "bright" or "dark" far more frequently than Dickinson. Dickinson, on the other hand, had a higher relative frequency of the color purple, which aligns with what other scholars have noticed about her work, though it wasn't the most common color in this specific sample.
The researchers also ran a "sensitivity test" to see if their results would change if they included "maybe" colors—words like "gold" or "silver" that could be a color or just a material. Even when they added these ambiguous words, the main finding held true: Bely's poetry was still significantly more colorful than Dickinson's. However, one interesting shift happened with the color yellow. In the strict count, yellow wasn't significantly different between the two poets. But when they included words like "gold," yellow suddenly became much more frequent in Bely's work. This showed the researchers that how you classify a word can change the story you tell about a specific color, even if the overall picture stays the same.
Ultimately, this paper didn't just tell us which poet liked purple more. It proved that you can build a fair, scientific way to compare how different languages handle color. By creating these aligned, multidimensional dictionaries, the researchers gave linguists a new tool to explore how writers use the full spectrum of language to paint pictures in our minds, showing us that while the colors might be universal, the way we speak them is wonderfully unique.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.