← Latest papers
💬 NLP

Sri Lanka Document Datasets: A Large-Scale, Multilingual Resource for Law, News, and Policy

This paper introduces a large-scale, multilingual open dataset collection comprising over 278,000 documents in Sinhala, Tamil, and English, covering Sri Lankan law, news, and policy to support research in computational linguistics and socio-political studies.

Original authors: Nuwan I. Senaratna

Published 2026-07-03
📖 4 min read☕ Coffee break read

Original authors: Nuwan I. Senaratna

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine Sri Lanka's public records—laws, news, court decisions, and government reports—as a massive, chaotic library scattered across hundreds of different buildings. Some books are locked in glass cases (PDFs), others are just scribbled on loose paper (websites), and many are written in three different languages: Sinhala, Tamil, and English. For a regular citizen, a journalist, or a researcher, trying to find a specific fact in this mess is like looking for a needle in a haystack that keeps moving.

This paper introduces a project called the Sri Lanka Document Datasets, which acts like a super-organized librarian and a digital moving truck combined. Here is what they have built:

1. The Big Collection (The "Digital Warehouse")

The team has gathered 278,621 documents (about 80.7 GB of data) into one neat, machine-readable package. Think of this as building a single, giant warehouse where every piece of important information from Sri Lanka is sorted, labeled, and ready to be used.

The collection includes:

  • The "Law Library": Parliamentary debates (Hansards), Supreme Court rulings, and police press releases.
  • The "Rulebook": Actual laws (Acts), draft laws (Bills), and urgent government notices (Gazettes).
  • The "Newsstand": Thousands of news articles and government updates.
  • The "Weather & Economy Report": Tourism stats, fish prices, flood warnings, and bank reports.

All of this is available in three languages, making it a true multilingual treasure chest.

2. The Magic Machine (The "Collection Pipeline")

How did they get all this data? They didn't hire a team of people to manually copy-paste documents. Instead, they built an automated robot system.

  • The Scout: This robot visits official government websites every day. It's polite; it checks the "Do Not Enter" signs (robots.txt) and waits its turn so it doesn't crash the websites.
  • The Sorter: When the robot finds a new document, it doesn't just grab it; it checks if it's already there. If it's new, it downloads it, reads it, and cleans it up.
  • The Translator & Cleaner: The system takes messy PDFs (which are like pictures of text) and turns them into clean, searchable text. It fixes broken lines and organizes the data into a standard format (JSON) so computers can easily understand it.
  • The Daily Update: This robot runs automatically, several times a day. If a new law is passed or a new flood warning is issued, the system catches it and adds it to the collection immediately.

3. The Open Door Policy (Licensing and Access)

Usually, big data sets are locked behind paywalls or require special permission to use. This project is different. It's like a public park rather than a private club.

  • Free to Enter: Anyone can download the data for free.
  • Free to Change: You can take the data, fix mistakes, or build your own tools on top of it, as long as you give credit to the original creators.
  • Always Available: The data is mirrored on two popular platforms (GitHub and Hugging Face), ensuring that even if one site goes down, the library is still open.

4. What's Next? (Future Plans)

The paper outlines three main goals for the future:

  1. Expand the Library: Add more types of documents from other government departments and historical archives.
  2. Polish the Language: Improve how the system reads complex sentences in Sinhala and Tamil, ensuring the "robot" understands the nuances of the languages perfectly.
  3. Read the Handwritten/Scanned Stuff: Currently, the system struggles with blurry or scanned documents. They plan to teach the robot to use "eyes" (OCR technology) to read text from low-quality images, turning even the messiest old papers into clean digital text.

Why Does This Matter?

The paper argues that by turning fragmented, hard-to-read records into a clean, open, and searchable database, they are giving a superpower to researchers, journalists, and citizens. It allows them to track how laws change, monitor government decisions, and study the country's history without getting lost in the paperwork. It's about making the "digital record" of Sri Lanka accessible to everyone, not just those with the time or tools to dig through the archives.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →