The Annotation Bottleneck in Persian Text NLP: Persian as an Annotation-Scarce Language
This paper argues that Persian should be characterized as an "annotation-scarce" language rather than a globally low-resource one, revealing that its NLP challenges stem from uneven task coverage, incompatible schemes, and access barriers rather than a simple lack of labeled data volume.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the world of computers that understand human language, there is a persistent belief that some languages are simply harder to teach than others. This difficulty is often blamed on a lack of digital text, as if a language needs a massive library of unlabelled books to function. However, the real challenge for artificial intelligence is not just having words, but having those words carefully marked up with instructions. Imagine a library where every book is written in a foreign language; if you want a computer to learn from them, you cannot just give it the raw text. You must provide a guide that explains what every sentence means, who is speaking, and what the emotions are. This process of adding instructions is called annotation. For many languages, this guide exists in abundance. For others, it is missing, making it nearly impossible to train reliable systems even if the raw text is plentiful.
This is the specific puzzle researchers in Tehran have been investigating regarding Persian, also known as Farsi. For years, the standard description of Persian in the field of computer science has been that it is a "low-resource" language, a label suggesting it is fundamentally short on data. But a new analysis suggests this label is too blunt an instrument. The researchers argue that Persian is not lacking in raw material; in fact, it has a massive digital footprint, appearing on nearly one percent of the world's known websites and in vast collections of news, blogs, and literature. The true problem, they find, is not a lack of words, but a lack of usable guides. The available instructions are scattered, inconsistent, or missing entirely for specific types of tasks, creating a bottleneck where the computer knows the language but cannot reliably understand the nuances of a medical report, a casual conversation, or a complex logical argument.
To reach this conclusion, the team at the University of Tehran did not simply count datasets. Instead, they built a detailed map of the Persian digital landscape, cataloging thirty-four distinct text resources available as of mid-2026. They looked at everything from ancient poetry collections to modern social media feeds, and from news archives to speech recordings. They also stepped outside the academic bubble to measure how much Persian actually exists on the open web, using independent tools to scan millions of pages. Their goal was to see if the amount of labelled data matched the sheer volume of text available. They found that the relationship is far from even. In some areas, like identifying names in news articles or parsing the grammar of standard sentences, the amount of labelled data is surprisingly high, even when compared to English. In these specific islands, Persian is well-served.
However, the picture changes dramatically when the researchers looked at other tasks. When they examined resources for understanding logical arguments or inferring meaning from context, the amount of labelled data dropped significantly, falling well below what would be expected given the language's presence on the web. The study also highlighted that the data that does exist often suffers from friction. Different projects use different rules for marking up text, making it difficult to combine them. Some resources are locked behind unclear licenses, while others lack the necessary documentation to be used correctly. This means that even when a dataset exists, it might be unusable for a new researcher trying to build a system for a different purpose, such as analyzing legal contracts or understanding regional dialects.
The researchers extended their investigation to spoken language to see if the pattern held true for audio. They found a similar story: while there are now thousands of hours of recorded Persian speech available, the depth of the instructions attached to them varies wildly. Some recordings have detailed phonetic breakdowns, while others have only basic transcripts. This confirms that the issue is not a total absence of data, but an uneven distribution of the specific, high-quality supervision needed for different jobs. The study explicitly rejects the idea that Persian is globally deficient in labelled volume. Instead, it proposes a more precise diagnosis: Persian is an "annotation-scarce" language. This term describes a situation where the scarcity is not in the total number of words, but in the availability of consistent, accessible, and well-documented instructions across the full range of tasks a computer might need to perform.
The implications of this finding are significant for how the field moves forward. If the problem were simply a lack of text, the solution would be to scrape more websites. But since the problem is the lack of usable guides, the solution requires a shift in strategy. The researchers suggest that the community needs to stop treating every new dataset as a standalone victory and start focusing on making existing resources easier to find, combine, and trust. They call for a public registry that tracks not just the size of a dataset, but its license, its specific rules, and how it overlaps with other resources. They also emphasize the need for better documentation and the protection of the people who create these annotations, ensuring that the data reflects the diversity of the language rather than just a single standard dialect.
Ultimately, this paper serves as a correction to a long-held assumption. It clarifies that Persian is not a language struggling to find its voice in the digital age; it is a language with a loud and visible presence that is currently hampered by a fragmented infrastructure of instructions. The path forward is not to simply collect more raw data, but to organize the data that already exists, fill the specific gaps where instructions are missing, and ensure that the guides we build are robust enough to handle the complexity of human communication. By shifting the focus from quantity to quality and usability, the field can move past the bottleneck and build systems that truly understand the richness of the Persian language.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.