PERCEPT: A Corpus for POS Tagging and Analysis of Persian-English Code-Mixing
This paper introduces PERCEPT, the first large-scale Persian-English code-mixed corpus annotated with Universal Dependencies POS tags, along with an LLM-assisted annotation framework and a comprehensive linguistic analysis revealing platform-specific patterns in code-mixing across social media.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the internet as a giant, chaotic playground where people from all over the world are shouting, whispering, and chatting all at once. In this playground, many people don't just speak one language; they mix them together like ingredients in a smoothie. This is called "code-mixing." It's when someone might start a sentence in their native tongue and suddenly throw in an English word because it feels just right, or because it's the only word they know for a new gadget. For computers, this is a nightmare. If a computer is trying to understand a sentence, it expects words to follow strict rules. But when languages get mixed up, the computer gets confused, like a chef trying to follow a recipe that suddenly switches from metric cups to imperial ounces in the middle of baking a cake. To teach computers how to handle this messy, fun reality, scientists need a special dictionary and a set of rules—a "corpus"—that shows them exactly how these mixed words work.
Enter PERCEPT, a new project that acts like a giant, organized library for Persian speakers who love to mix English into their posts. The researchers, a team of linguists and computer scientists from Iran, noticed that while we have maps for other mixed-language pairs, the Persian-English mix was a dark, unexplored forest. They wanted to build a flashlight to see what was happening there. So, they collected 6,800 posts from three popular digital hangouts: X (formerly Twitter), Instagram, and Digikala (a massive online shopping site). They didn't just collect the words; they used a super-smart AI assistant named Gemini to act as a detective, tagging every single English word (or English word written in Persian letters) with its "job" in the sentence. In grammar, this job is called a "Part-of-Speech" (POS). Is the word a noun (a thing)? A verb (an action)? An adjective (a description)? The team also asked the AI to guess what the post was talking about, like "shopping," "politics," or "funny memes."
Here is the exciting part: the team didn't just build the library; they went inside and started reading the books to see what patterns they could find. They discovered that when Persian speakers mix in English, they mostly use it for nouns (things). It's like they are borrowing the "objects" from English but keeping the "actions" and "descriptions" in Persian. For example, they might say, "I bought a new laptop," where "laptop" is the borrowed noun, but the rest of the sentence stays Persian. However, the "flavor" of the mix changes depending on where you are. On the shopping site, Digikala, people mixed in a lot of brand names and product models (which are special names, or "proper nouns"). On X, they mixed in more verbs (actions), and on Instagram, they leaned toward adjectives (descriptions).
The researchers also looked at where these mixed words hide in a sentence. They found a surprising consistency: no matter which platform you use, mixed words tend to appear at the beginning or the middle of a sentence, rarely at the very end. It's as if the speakers drop the English word in early to set the scene, then finish the thought in Persian. But the most interesting discovery was about "triggering." On Digikala, if someone used one English word, they were very likely to use another one right after. It's like once you start talking about a specific brand or product, the English words just keep flowing. This didn't happen as much on Instagram or X.
Finally, they checked which topics caused the most mixing. Unsurprisingly, the topics of "Industry & Commerce" and "Brand & Business" had the highest amount of code-mixing. When people talk about shopping, tech, or business, they simply can't help but sprinkle in English terms. The team confirmed their AI's work by having human experts double-check a sample, and the humans agreed with the AI almost all the time, proving that their new library is reliable.
In short, this paper doesn't just give us a new dataset; it gives us a clear picture of how Persian speakers naturally blend languages. It suggests that while the types of words they borrow change based on the platform, the way they mix them follows a steady rhythm. This map is now open for anyone to use, helping to build better translation tools, smarter chatbots, and a deeper understanding of how we communicate in a world where languages are constantly dancing together.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.