← Latest papers
💬 NLP

CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data

This paper introduces CommonLID, a community-driven, human-annotated benchmark covering 109 languages that reveals existing language identification models significantly overestimate their accuracy on noisy web data, thereby providing a crucial resource for developing more representative multilingual corpora.

Original authors: Pedro Ortiz Suarez, Laurie Burchell, Catherine Arnett, Rafael Mosquera-Gómez, Sara Hincapie-Monsalve, Thom Vaughan, Damian Stewart, Malte Ostendorff, Idris Abdulmumin, Vukosi Marivate, Shamsuddeen Has
Published 2026-06-10
📖 4 min read☕ Coffee break read

Original authors: Pedro Ortiz Suarez, Laurie Burchell, Catherine Arnett, Rafael Mosquera-Gómez, Sara Hincapie-Monsalve, Thom Vaughan, Damian Stewart, Malte Ostendorff, Idris Abdulmumin, Vukosi Marivate, Shamsuddeen Hassan Muhammad, Atnafu Lambebo Tonja, Hend Al-Khalifa, Nadia Ghezaiel Hammouda, Verrah Otiende, Tack Hwa Wong, Jakhongir Saydaliev, Melika Nobakhtian, Muhammad Ravi Shulthan Habibi, Chalamalasetti Kranti, Carol Muchemi, Khang Nguyen, Faisal Muhammad Adam, Luis Frentzen Salim, Reem Alqifari, Cynthia Amol, Joseph Marvin Imperial, Ilker Kesen, Ahmad Mustafid, Pavel Stepachev, Leshem Choshen, David Anugraha, Hamada Nayel, Seid Muhie Yimam, Vallerie Alexandra Putra, My Chiffon Nguyen, Azmine Toushik Wasi, Gouthami Vadithya, Rob van der Goot, Lanwenn ar C'horr, Karan Dua, Andrew Yates, Mithil Bangera, Yeshil Bangera, Hitesh Laxmichand Patel, Shu Okabe, Fenal Ashokbhai Ilasariya, Dmitry Gaynullin, Genta Indra Winata, Yiyuan Li, Juan Pablo Martínez, Amit Agarwal, Ikhlasul Akmal Hanif, Raia Abu Ahmad, Esther Adenuga, Filbert Aurelian Tjiaranata, Weerayut Buaphet, Michael Anugraha, Sowmya Vajjala, Benjamin Rice, Azril Hafizi Amirudin, Jesujoba O. Alabi, Srikant Panda, Yassine Toughrai, Bruhan Kyomuhendo, Daniel Ruffinelli, Akshata A, Manuel Goulão, Ej Zhou, Ingrid Gabriela Franco Ramirez, Cristina Aggazzotti, Konstantin Dobler, Jun Kevin, Quentin Pagès, Nicholas Andrews, Nuhu Ibrahim, Mattes Ruckdeschel, Amr Keleg, Mike Zhang, Casper Muziri, Saron Samuel, Sotaro Takeshita, Kun Kerdthaisong, Luca Foppiano, Rasul Dent, Tommaso Green, Ahmad Mustapha Wali, Kamohelo Makaaka, Vicky Feliren, Inshirah Idris, Hande Celikkanat, Abdulhamid Abubakar, Jean Maillard, Benoît Sagot, Thibault Clérice, Kenton Murray, Sarah Luger

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the internet as a massive, chaotic library containing books in thousands of different languages. If you want to build a "super-reader" (a large language model) that can understand all these books, you first need a librarian who can quickly sort the books into the correct language sections. This librarian is called a Language Identification (LID) system.

The paper you shared, CommonLID, argues that the current librarians are doing a terrible job, especially with the messy, noisy books found on the open web. Here is a breakdown of their findings using simple analogies:

1. The Problem: Librarians Who Only Know "Clean" Books

For a long time, researchers thought the job of sorting languages was "solved." They tested their librarians using very clean, perfect books like government documents or translated news (think of these as encyclopedias). On these clean books, the librarians scored 95% or higher.

However, the real internet is more like a bustling flea market. It has slang, typos, mixed languages, and random noise. The paper found that when you take those same "expert" librarians and put them in the flea market, they get confused. They start mislabeling languages, especially for languages that don't have many books available (low-resource languages).

The Analogy: It's like testing a chef on a perfect, quiet kitchen with fresh ingredients, and then expecting them to cook a gourmet meal in a windy tent with expired spices. The chef fails, not because they are bad, but because the test didn't match the real-world chaos.

2. The Solution: A New, Real-World Test (CommonLID)

To fix this, the authors created CommonLID. Think of this as a new, rigorous exam designed specifically for the "flea market" conditions of the web.

  • Who took the exam? Over 80 native speakers from around the world (the community) helped create it.
  • What was on the test? They took real, messy text from the web (Common Crawl) and had native speakers manually sort it line-by-line.
  • How big is it? It covers 109 different languages, many of which have been ignored by previous tests. It's like adding a whole new wing to the library that was previously locked.

3. The Results: The "Experts" Are Overconfident

The authors took eight of the most popular "librarian" tools (like CLD2, GlotLID, fasttext) and put them through this new CommonLID exam. They also tested them on the old, clean exams to compare.

The Shocking Findings:

  • The Old Scores Were Fake: The high scores on the clean exams were misleading. They made the tools look much smarter than they actually are.
  • The Real-World Struggle: On the new CommonLID test, the best tools only got about 60-70% correct. That's a failing grade in many school systems.
  • The "Bible" Bias: Many of these tools were trained mostly on religious texts (like the Bible). When they see a religious text, they are great. But when they see a tweet, a forum post, or a blog, they get lost. It's like a student who memorized the dictionary but can't hold a conversation.

4. The Human Element: Giving Credit to the Helpers

A unique part of this paper is how they built the dataset. Instead of hiring cheap crowd-workers, they invited researchers and native speakers to help.

  • The Rule: If you annotated (sorted) at least 100 documents, you became a co-author of the paper.
  • Why? This ensures that the people who speak the languages are the ones credited for improving the technology, rather than just being invisible data entry workers.

5. The Big Picture: We Need Better Tools

The paper concludes that we cannot build fair, global AI systems if our "librarians" keep failing to sort the books correctly.

  • Current State: No single tool works well for all languages and all types of text (web vs. clean).
  • The Takeaway: We need to stop trusting the old, clean tests. We need to use new, messy, real-world tests like CommonLID to find tools that actually work in the real world.

In short: The paper says, "Stop pretending the language sorting problem is solved. We built a new, harder test using real people and real web data, and it proved that our current tools are failing the languages that need help the most."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →