Qaamuuska-NLP: A Structured 46K-Entry Somali Lexicon Extracted from a Print Dictionary
This paper introduces Qaamuuska-NLP, a structured 46,314-entry machine-readable Somali lexicon extracted via a deterministic rule-based pipeline from the 2012 monolingual dictionary *Qaamuuska Af-Soomaaliga*, which preserves detailed grammatical and lexical information while retaining original entry text for verification.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Language is more than just a stream of words; it is a structured system where every word carries a specific job, a history of how it changes form, and a network of relationships with other words. For computers to understand a language, they need a map of this structure. This map is called a lexicon, a digital dictionary that does not just list words and their meanings, but also records grammatical details like whether a noun is masculine or feminine, or how a verb changes when it describes an action done to someone versus an action done alone. For many of the world's languages, these maps are already vast and detailed, allowing computers to translate, search, and analyze text with great accuracy. However, for Somali, a language spoken by millions across the Horn of Africa and the diaspora, these digital maps have been incomplete. While Somali has a rich tradition of printed dictionaries, the information inside them has remained locked in a format designed for human eyes, making it difficult for machines to use systematically.
A team of researchers at Ogaal Labs has now opened that lock. They have created a new digital resource called Qaamuuska-NLP, a structured collection of 46,314 entries extracted directly from a major printed Somali dictionary published in 2012. The researchers did not simply scan the pages of the book and turn the images into text, a process that often introduces errors. Instead, they worked with the digital text layer that already existed inside the PDF file of the dictionary. They built a set of precise, rule-based instructions to read the book's layout, identify where one word entry ended and the next began, and pull out the specific grammatical clues hidden in the formatting. The result is a clean, machine-readable database that preserves the original dictionary's organization while adding a layer of structure that computers can instantly query. This work does not replace the original book, but rather translates its deep linguistic knowledge into a new language that software can speak.
The source of this new resource is Qaamuuska Af-Soomaaliga, a comprehensive monolingual dictionary edited by Annarita Puglielli and Cabdalla Cumar Mansuur. Unlike simple word lists, this book contains a wealth of grammatical information that is essential for understanding how Somali works. It tells the reader the gender of nouns, the class of verbs, and whether a verb requires a direct object. It also includes synonyms, cross-references to other words, and labels that indicate if a word belongs to a specific field like medicine or law. The challenge for the researchers was that this information was not stored in neat columns of data; it was embedded in the visual design of the printed page. The dictionary uses two columns of text, specific abbreviations, and small superscript numbers to distinguish between words that look the same but have different meanings. To build their digital version, the researchers wrote a computer program that could navigate this layout without needing to guess or learn from examples. The program reads the left column of a page, then the right column, and scans for the repeating patterns that signal the start of a new word entry.
Once the program identified a word entry, it broke the text down into its component parts. It separated the main word from its definition, its grammatical tags, and its references to other words. The team found that the dictionary contained 46,314 distinct records. Of these, about 25,000 provided actual definitions, while the remaining 21,000 were cross-references that pointed the reader to another entry. The extraction process was highly successful at capturing the grammatical details that make the dictionary useful. For the 34,726 noun entries, the system successfully identified the gender for nearly all of them, marking them as masculine, feminine, or capable of being either. For the 11,445 verb entries, the system captured the conjugation class and whether the verb was transitive or intransitive for almost every single one. The researchers also preserved the original text of every entry alongside the new structured data. This ensures that anyone using the database can always check the original source to verify the information, keeping the connection between the digital map and the physical book intact.
The researchers tested the quality of their work by looking at how well the grammatical labels in the dictionary matched the actual shape of the words. They asked a simple question: if you only look at the end of a Somali word, can you predict its grammatical category? They used a straightforward method that looked at the last few letters of the words to guess their gender or verb class. The results showed that the dictionary's labels were not random. For verb classes, the method was correct 93.5 percent of the time, far better than simply guessing the most common class. For noun gender, the method was correct 86.4 percent of the time. These numbers suggest that the grammatical information in the dictionary is deeply tied to the physical form of the words, confirming that the extracted data reflects real linguistic patterns rather than arbitrary choices. However, the researchers noted that this test was a diagnostic tool to check the internal consistency of the resource, not a full test of a computer's ability to analyze Somali text in the real world.
The new resource fills a specific gap in the landscape of Somali language technology. While other projects have combined multiple dictionaries or focused on generating lists of word forms, this work focuses on preserving the structure of a single, authoritative source. It captures the editorial decisions and the specific way the original authors organized the language. The database includes 11,415 synonym links and 21,032 cross-references, creating a network that shows how words relate to one another within the dictionary's own logic. It also identifies 1,945 entries that belong to specialized fields such as medicine, physics, and law, providing a starting point for building technical vocabulary in Somali. The researchers are careful to note that this resource is a representation of that specific 2012 dictionary, not a complete census of every word used in Somali today. The coverage and the specific definitions reflect the choices made by the original editors.
There are limitations to this approach, which the authors openly acknowledge. Because the extraction rules were designed specifically for the layout of this one book, they cannot be directly applied to a different dictionary without modification. The team also noted that the original PDF contained some inconsistencies in how apostrophes were typed, which could affect exact matches between words, though they kept the original text to allow for future corrections. Furthermore, while the extraction code and the statistical summaries of the data are being prepared for public release, the full list of 46,314 words and their definitions cannot be shared freely yet. This is because the original dictionary is copyrighted, and the researchers are still working to secure the necessary permissions to redistribute the complete content. Until that permission is granted, the full dataset remains a private research artifact, though the methods used to create it are available for others to study and adapt.
The value of this work lies in its ability to turn a static book into a dynamic tool. By converting the dictionary into a structured format, the researchers have made it possible for computers to search for words by their grammatical properties, to trace the network of synonyms and cross-references, and to analyze the distribution of specialized terms. This does not solve every problem in Somali language processing, but it provides a solid foundation of high-quality, curated data. It demonstrates that valuable linguistic knowledge often sits dormant in printed books, waiting for a method to unlock it. The Qaamuuska-NLP project shows that with a careful, rule-based approach, it is possible to recover this knowledge without losing the nuance and structure of the original source, offering a new resource for researchers, educators, and developers working to bring Somali into the digital age.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.