Automatic Construction of a Legal Citation Graph from 100 Million Ukrainian Court Decisions: Large-Scale Extraction, Topological Analysis, and Ontology-Driven Clustering
This paper presents the automatic construction of a massive legal citation graph from over 100 million Ukrainian court decisions, revealing that unsupervised analysis of 502 million citation edges effectively encodes legal domain boundaries, predicts legislative importance with near-perfect accuracy, and detects regime changes, thereby enabling the creation of an ontology-driven workflow for LLM-assisted legal analysis.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the entire legal system of Ukraine as a massive, bustling library containing over 100 million books (court decisions). For years, these books sat on shelves, mostly unread by computers, because the way they referenced other laws was messy, inconsistent, and written in a complex human language.
This paper describes a project where the authors built a super-fast robot librarian to read every single one of those books and map out exactly how they connect to one another.
Here is the breakdown of what they did and what they found, using simple analogies:
1. The Robot Librarian (The Extraction)
The authors created a specialized tool (a "regex pipeline") that acts like a high-speed scanner.
- The Task: It had to find references to laws inside 100.7 million court decisions. That's 1.1 terabytes of text—enough to fill a small library.
- The Speed: It did this in just 5 hours on a standard computer server. To put that in perspective, a human reading one page a minute would take centuries to do this.
- The Result: The robot found 502 million connections (citations). It was incredibly accurate, getting a perfect score on a test sample. It successfully identified six different types of references, from specific articles in the Civil Code to rulings by the Supreme Court.
2. The Map of Connections (The Graph)
Once the robot finished reading, the authors didn't just have a pile of notes; they built a giant spiderweb (a graph).
- How it works: Imagine every court decision is a dot, and every law it mentions is another dot. The robot drew a line between them.
- The Shape: They found that this web follows a "Power Law." In plain English, this means a tiny handful of laws are super-popular (cited millions of times), while most laws are cited very rarely.
- Analogy: Think of it like a social media network. A few celebrities (like the main articles on theft or civil procedure) have millions of followers, while most regular people have only a few. In Ukraine, the "celebrity" laws are the procedural rules that judges use in almost every single case.
3. Discovering Hidden Neighborhoods (Ontology & Clustering)
The most surprising part of the paper is what happened when they looked at how laws are cited together.
- The Discovery: They used a computer algorithm to group laws that are frequently cited in the same cases. They didn't tell the computer what "Civil Law" or "Criminal Law" was. They just let the data speak.
- The Result: The computer naturally grouped the laws into four distinct neighborhoods: Civil, Criminal, Administrative, and Commercial.
- Analogy: Imagine walking into a room full of people and asking them to sort themselves into groups without telling them the rules. If you ask them to group by "who they talk to," the accountants will naturally cluster together, and the artists will cluster together. The computer did this with laws, proving that the structure of the court system itself reveals the categories of law, even without human experts labeling them.
4. Reading the History (Time Travel)
Because they had data from 2007 to 2026, they could watch the legal system change in real-time.
- Reforms: When a new law code was introduced (like in 2012 or 2017), the map showed a sudden "earthquake." Old laws stopped being cited, and new ones exploded in popularity.
- The 2022 Invasion: The paper notes a specific "spike" in 2022. The variety of laws being cited suddenly became much more chaotic and diverse (entropy increased). New laws about the war appeared on the map for the first time, showing how the legal system instantly adapted to the crisis.
5. Predicting the Future
The authors tested if they could guess which laws would become important in the future.
- The Test: They used the history of citations to predict the top 1,000 most important laws for the years 2020–2026.
- The Score: They were 99.8% accurate.
- The Lesson: The best way to know if a law is important isn't to read the text; it's to see how often judges cite it. The "popularity" of a law is a perfect predictor of its future importance.
Why This Matters (According to the Paper)
The authors explain that this map is now being used to help Artificial Intelligence (AI) understand Ukrainian law.
- Currently, AI models often hallucinate or make things up because they don't have a solid "map" of how laws relate to each other.
- This project provides that map. It gives the AI a "domain layer"—a structured, data-driven guide that tells the AI, "These laws belong together, and this is how they are used in real life." This helps the AI give more accurate, grounded legal advice.
In summary: The paper is about building a massive, automated map of Ukraine's legal system. By letting the data draw the lines between laws, they discovered that the court system naturally organizes itself into clear categories, that a few laws hold the whole system together, and that this map can teach AI how to think like a lawyer.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.