← Latest papers
💬 NLP

Trilingual Topic Modeling of Sri Lankan Parliamentary Debates

This paper presents an end-to-end framework utilizing LLM-based text extraction and multilingual embedding clustering to successfully perform unsupervised topic modeling on Sri Lanka's trilingual parliamentary debates, overcoming challenges like complex layouts and code-mixing to identify thematic structures that align with major national events.

Original authors: Himath Dhanapala, Haren Daishika, Himandhi Kuruppu, Sithija Seneviratne, Ashini Kavindya, Patalee Narasinghe, Sandeepa Weerasekara, Nisansa de Silva, Sandareka Wickramanayake

Published 2026-08-24
📖 4 min read☕ Coffee break read

Original authors: Himath Dhanapala, Haren Daishika, Himandhi Kuruppu, Sithija Seneviratne, Ashini Kavindya, Patalee Narasinghe, Sandeepa Weerasekara, Nisansa de Silva, Sandareka Wickramanayake

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the vast landscape of modern computing, a field known as natural language processing strives to teach machines how to read, understand, and organize human speech. For decades, this work has focused heavily on languages like English, where the rules of grammar are relatively straightforward and the digital tools are plentiful. However, the world is filled with languages that behave very differently. Some, like Sinhala and Tamil, spoken widely in Sri Lanka, are agglutinative, meaning they build complex words by stringing together many small pieces of meaning, making it difficult for standard computer programs to recognize that two different-looking words actually share the same root. Furthermore, in many parts of the world, people do not speak in just one language at a time; they mix them freely within a single sentence, a phenomenon known as code-mixing. When these linguistic complexities meet the messy reality of official government records, which often arrive as poorly formatted digital documents, the result is a mountain of text that standard software simply cannot climb. Understanding what is being said in these records is crucial, as they hold the keys to a nation's political priorities, economic struggles, and social concerns, yet they have remained largely locked away from systematic analysis.

A team of researchers from the University of Moratuwa in Sri Lanka has now built a new way to unlock this treasure trove of information. They turned their attention to the Hansards, the official transcripts of the Sri Lankan Parliament, which contain thousands of speeches delivered in Sinhala, Tamil, and English, often mixed together within a single address. These documents are notoriously difficult to process because they exist as scanned PDFs with confusing two-column layouts and inconsistent formatting that breaks traditional reading software. To solve this, the researchers first employed a powerful artificial intelligence tool capable of reading these messy documents and extracting the actual spoken words, ignoring the visual noise of the page. They then fed this cleaned text into a system designed to understand the meaning of words across different languages, rather than just counting how often they appear. By mapping the speeches into a digital space where similar ideas sit close together, regardless of the language used, they were able to group thousands of individual speeches into coherent themes.

The result of this effort is a clear map of the nation's political conversation over nearly a decade, from 2017 to 2026. The researchers analyzed 19,553 speeches and successfully organized them into 30 major topics, ranging from energy security and economic crises to national security and healthcare. What makes this achievement significant is that the computer did this without being told what to look for; it discovered the themes on its own. The system found that speeches about the economy surged dramatically during the 2022 financial crisis and the subsequent period of civil unrest, while discussions about national security spiked following the 2019 Easter Sunday attacks. Crucially, the model treated the three languages as a single, unified stream of conversation. A speech in Tamil about a specific policy issue was grouped together with a speech in Sinhala about the same issue, proving that the underlying meaning mattered more than the language used to express it.

The researchers also tested whether older, more traditional methods of analyzing text could handle this task. They found that standard approaches, which rely on counting word combinations, failed completely. These older methods could not bridge the gap between the different languages or handle the complex word structures of Sinhala and Tamil, resulting in fragmented and meaningless groups. In contrast, the new approach, which uses deep learning to understand context, successfully identified the structure of the debates. The team also explored a hybrid method that combined the deep understanding of meaning with a focus on specific keywords to see if it could make the groups even sharper. While this method did make the boundaries between topics clearer, it also meant that some speeches were left out of the groups entirely, suggesting a trade-off between precision and coverage.

Ultimately, this work demonstrates that it is possible to make sense of complex, multilingual political discourse without human intervention. By treating the diverse languages of Sri Lanka as a single, rich dataset, the researchers were able to reveal how the nation's attention shifts in response to real-world events. The findings show that the topics discussed in parliament are not random; they rise and fall in direct response to the country's challenges, from the management of public debt to the care of the sick. This new framework provides a solid foundation for future studies, allowing researchers and the public to track the evolution of political priorities with a clarity that was previously impossible, turning a chaotic archive of speeches into a structured story of a nation's journey.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →