← Latest papers
💻 computer science

The Saudi Legal Corpus for AI: An Auditable, Official-Source-Grounded, Article-Level Corpus of Saudi Arabian Legislation

This paper introduces the Saudi Legal Corpus for AI, a comprehensive, officially sourced, and auditable article-level dataset of 290 Saudi legislative instruments featuring unique provenance tracking, structured analytical layers like citation graphs and defined-term glossaries, and a retrieval benchmark demonstrating the superior performance of metadata-enhanced search over standard baselines for Arabic legal NLP.

Original authors: Abdullah Almohammedi

Published 2026-09-01
📖 6 min read🧠 Deep dive

Original authors: Abdullah Almohammedi

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the modern world, computers are increasingly asked to answer questions about the law. To do this accurately, they rely on a technique called retrieval-augmented generation. This process works like a librarian who, before answering a question, first searches a vast library of trusted documents to find the exact passage that holds the answer. If the computer can find the right text, it can summarize it for the user without making things up. However, this system only works if the library exists in a format the computer can read and understand. For many languages and legal systems, especially in the Arab world, this library has been missing. The laws of Saudi Arabia, a major global economy, are published in official gazettes and on government websites, but they exist as human-readable pages and documents, not as structured data that software can easily search through. Without a machine-readable version of these laws, artificial intelligence cannot reliably assist with legal questions in the Kingdom, leaving a significant gap in both technology and access to justice.

To bridge this gap, a researcher named Abdullah Almohammedi has built the Saudi Legal Corpus for AI, a massive, structured digital collection of the Kingdom's laws. This is not a simple list of links or a collection of scanned PDFs. Instead, it is a carefully organized database containing 290 distinct legislative instruments, ranging from major statutory laws to detailed implementing regulations and procedural rules. The collection covers twelve different areas of law, including commercial business, criminal justice, labor rights, and taxation. The core of this work is the transformation of 290 laws into 15,689 individual article-level records. Each record is a single, self-contained piece of text, such as one specific article from a law, totaling approximately 1.2 million words in Arabic. The entire collection is designed to be auditable, meaning every single piece of text can be traced back to its exact official source, ensuring that the computer is reading the real law and not a summary or a guess.

The construction of this corpus follows a strict set of rules to ensure accuracy and trustworthiness. The most important rule is that the official Arabic text is the only version that counts. While the collection includes English and Chinese translations, these are clearly marked as non-authoritative references, useful for understanding but not legally binding. Every single article in the database carries a label that explains exactly where it came from and how it was verified. The researcher used a four-tier system to grade the reliability of each source, distinguishing between laws that were cross-checked against multiple official documents and those that came from a single source. This level of detail allows users to filter the data based on how much evidence they need. For instance, a system designed for high-stakes legal advice might only use the most rigorously verified sources, while a research project might accept a broader range. The database also includes a "freshness manifest," a list that flags 16 specific tracks where the official government portal might not yet have the latest amendments, ensuring that users are aware of potential outdated information rather than being misled by it.

Beyond just storing the text, the researcher added several layers of analysis that have never existed for Saudi law before. The corpus includes a map of how laws reference one another, showing which articles cite other articles. This citation graph contains over 3,600 references, revealing which laws act as central hubs in the legal system, such as the Labor Law and the Companies Law, which are frequently cited by other regulations. There is also a hand-classified map showing which laws have replaced or repealed older ones, helping to track the history of legal changes. Additionally, the project created a glossary of nearly 2,000 terms that are specifically defined within the laws themselves, along with over 3,300 definitions. This helps clarify how the same word might mean different things in different contexts. To make the data ready for modern AI tools, the text was also broken down into smaller, manageable chunks, allowing computers to process specific sections of a law without getting lost in thousands of pages of text.

To test how well this new resource works, the researcher created a benchmark using 519 specific questions about Saudi law, each paired with the correct article that contains the answer. These questions were written by reading the actual laws and then formulating queries that a person might ask. When the researchers tested a standard search method against this new, metadata-rich system, the results were telling. The new system correctly identified the right article 93.3 percent of the time, compared to 90.4 percent for the standard method. The difference was most noticeable when the questions asked for definitions, such as "what is a sale contract?" In these cases, the new system was significantly more accurate, finding the correct definition 90.9 percent of the time versus 74.5 percent for the standard method. This gap demonstrates that the extra work put into organizing the data and adding detailed labels pays off, making it much easier for computers to find the right information even when the search terms do not perfectly match the text.

The release of this corpus is a significant step forward for legal technology in the region, but it comes with clear boundaries and limitations. The collection is not exhaustive; it covers the principal laws but does not include every single ministerial decision or municipal rule. The researcher is transparent about this, documenting the scope of the collection and the specific risks associated with each piece of data. The work is explicitly not an official government publication, and the English and Chinese layers are not official translations. The only binding text remains the original Arabic as published in the official gazette. By providing this structured, verified, and auditable resource, the project offers a foundation for building reliable legal assistants and conducting deeper studies on the structure of Saudi law. It provides a template for how other countries without machine-readable official gazettes can organize their own legal systems for the digital age, ensuring that the law remains accessible, understandable, and grounded in fact.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →