← Latest papers
💻 computer science

Freeing the Law with LOCUS: A Local Ordinance Corpus for the United States

This paper introduces LOCUS, a comprehensive, machine-readable corpus of nearly all publicly available U.S. municipal and county ordinances derived via OCR from fragmented sources, accompanied by harmonized access layers and specialized AI models to enable large-scale legal research and analysis of local laws.

Original authors: Denis Peskoff, Joe Barrow, Christopher Vu, Diag Davenport

Published 2026-06-19
📖 5 min read🧠 Deep dive

Original authors: Denis Peskoff, Joe Barrow, Christopher Vu, Diag Davenport

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the law as a massive, multi-layered library. For years, legal AI (computer programs that understand law) has been very good at reading the "big books" on the top shelves: federal laws, state statutes, and court cases. These are the rules that apply to everyone in the country or a whole state.

But there is a whole basement level of the library that has been locked away and filled with dust: local laws. These are the rules your city or county makes about things like zoning (where you can build a house), noise levels, business licenses, and animal control.

The paper introduces LOCUS (Local Ordinance Corpus for the United States), which is essentially a massive project to unlock that basement, clean up the books, and put them on a digital shelf where computers can finally read them.

Here is a breakdown of what they did, using simple analogies:

1. The Problem: A Library with No Catalog

Before LOCUS, finding local laws was like trying to find a specific recipe in a library where every town had its own cookbook, but:

  • The books were in different languages (formats).
  • Some were handwritten, some were typed, and some were just photocopies of photocopies (scanned PDFs).
  • There was no master list telling you which town had which book.
  • You couldn't just ask a librarian (a search engine) for "all the rules about parking in America" because the books were scattered across thousands of different websites designed for humans to browse, not for computers to scan.

2. The Solution: The "LOCUS" Project

The authors built a giant pipeline to fix this. Think of it as a team of robots with super-powered eyes and brains:

  • The Collectors: They sent robots to visit thousands of city and county websites to download the law books.
  • The Translators (OCR): Since many laws were just pictures of text (scanned PDFs), they used a special AI (called LightOnOCR) to "read" the pictures and turn them into digital text, like converting a handwritten letter into a Word document.
  • The Organizers: They cleaned up the text, removing page numbers and headers, and chopped the massive books into individual "laws" (like separating a single recipe from a whole cookbook).
  • The Labelers: They used AI to tag every single law with labels like "This is about buildings," "This is about business," or "This is about noise."

3. The Result: A "County-Harmonized" Map

The final product is a dataset covering 9,239 cities and counties. However, to make it useful for research, they created a simplified version called the "County-Harmonized Access Layer."

The Analogy: Imagine you want to study the weather across the US. You don't need a separate report for every single street corner; you need one reliable report for every county.

  • For every US county, LOCUS picks the "biggest" local law book available (either the county's own book or the book of the biggest city inside that county).
  • This doesn't mean it solves every legal question (like "Can I build a shed in this specific backyard?"), but it gives researchers a solid, comparable foundation to study local laws across the whole country.

4. Giving the Laws a "Personality Score"

One of the coolest parts of the paper is that they didn't just organize the text; they gave the laws "scores" on four specific personality traits. They trained AI to read a law and guess how it feels on a scale:

  1. Opacity (Confusing vs. Clear): Is this law written in plain English, or is it a confusing maze of jargon?
  2. Paternalism (Protecting You vs. Protecting Others): Is the law trying to stop you from hurting yourself (like a curfew for kids), or is it trying to stop you from hurting others (like a noise ordinance)?
  3. Enforcement Discretion (Strict vs. Flexible): Does the law give police officers a lot of room to decide what to do, or is it a strict "yes/no" rule?
  4. Problem Salience (Urgent vs. Minor): Does the law treat the issue as a huge emergency, or a minor annoyance?

Example: A law saying "No one under 16 can go to a festival without an adult" gets a high Paternalism score (it's protecting the kid) and a low Opacity score (it's very clear).

5. Why This Matters

The authors found that local laws aren't just a random mess. They have a hidden structure.

  • Cities tend to focus on "nuisance" (noise, pets, public order).
  • Counties tend to focus on "zoning" (land use, building permits).
  • Different parts of the country have different "flavors" of laws (e.g., Florida laws tend to be very confusing/opaque, while other states are clearer).

The Bottom Line

The paper claims that by "freeing" these local laws from their dusty, fragmented formats and turning them into a clean, searchable dataset, they have built the infrastructure for the next generation of legal AI.

They aren't claiming that AI can now act as a judge or give legal advice. Instead, they are saying: "We have finally built the map and the library. Now, researchers can start asking big questions about how local laws work, how they differ across the country, and how they affect people's daily lives."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →