ImmigrationReason: A Structured Dataset of U.S. Immigration Appeals for Legal Reasoning Research
This paper introduces ImmigrationReason, a large-scale structured dataset of 12,375 U.S. Citizenship and Immigration Services administrative appeals that addresses the gap in legal NLP resources for administrative adjudication by providing detailed, expert-verified records of legal frameworks, evidence sufficiency, and adjudicator errors to enable advanced research in outcome prediction, error analysis, and regulatory agent design.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The law is often imagined as a grand theater of courtroom drama, where judges in robes weigh evidence and deliver final verdicts on the fate of nations. Yet, the vast majority of legal decisions never reach a courtroom. They happen in the quiet, bureaucratic corridors of administrative agencies, where government officials review millions of applications for rights, benefits, and status every year. In the United States, one of the most significant of these processes involves employment-based immigration, where foreign nationals petition to live and work in the country. These cases are not decided by a single judge in a trial, but by a complex system of officers who apply specific rules to individual stories, often leading to denials that can be appealed to a higher administrative body. Understanding how these decisions are made, and where they might go wrong, is crucial not just for the people involved, but for anyone interested in how artificial intelligence might one day assist in such high-stakes regulatory work.
A team of researchers at Stanford University has now opened a window into this hidden world of administrative justice with a new resource called ImmigrationReason. They have gathered and organized a massive collection of 12,375 real-world decisions from the Administrative Appeals Office, the body that reviews denied immigration petitions. Spanning two decades from 2005 to 2026, this dataset is unique because it does not just contain the raw text of the decisions; it breaks them down into a structured map of the legal reasoning used in each case. The researchers have labeled every specific piece of evidence, tracked whether the original officer was right or wrong, and even captured the exact words the higher authority used to criticize the initial decision. This level of detail transforms a pile of legal documents into a structured dataset that can be studied by computers, offering a rare glimpse into the mechanics of how the government decides who gets to stay.
The work began by collecting every publicly available decision from the immigration appeals office for two specific types of employment visas. The team faced a significant hurdle: many of the older documents from before 2017 were simply scanned images of paper files, which are notoriously difficult for computers to read accurately. Standard software often garbled the text, turning legal citations into nonsense or missing entire footnotes. To solve this, the researchers used a powerful new type of artificial intelligence to transcribe these documents, creating clean, readable text that preserved the original structure and citations. They then built a sophisticated system to read these cleaned documents and extract specific information, such as the type of visa being sought, the legal arguments made, and the final outcome. To ensure the computer was not making mistakes, they ran the extraction process three times using different methods and had a third, highly advanced AI model act as a judge to resolve any disagreements between the first two attempts. This rigorous process ensured that the final dataset was as accurate and reliable as possible.
What makes this collection so valuable is the depth of its annotations. In most legal datasets, the focus is on broad categories, like whether a case was won or lost. Here, the researchers went much deeper. They analyzed each decision to see how the appeals office evaluated every single requirement of the law. For example, in a case requiring proof of "extraordinary ability," the law lists ten different ways a person might qualify. The dataset tracks exactly how the appeals office ruled on each of those ten points for every single case. It also records whether the higher authority agreed with the original officer, partially disagreed, or completely overturned the decision. Perhaps most importantly, the dataset includes nearly 9,000 direct quotes where the appeals office explicitly criticized the original officer for making a legal error, such as ignoring evidence or applying the wrong rule. This provides a unique library of examples showing exactly how legal reasoning can fail and how it should be corrected.
The data also reveals a natural experiment in how the law changes over time. In December 2016, the rules for one type of visa were completely rewritten, shifting from one set of standards to another. Because the dataset covers the years before and after this change, it allows researchers to see exactly how the legal landscape shifted. They can observe how the success rates for applicants changed, how the arguments used by lawyers evolved, and how the government's interpretation of the law adapted to the new framework. This creates a clear boundary in the data that helps scientists test whether artificial intelligence can understand that laws are not static, but living systems that change over time.
Beyond the structure of the law, the dataset uncovers interesting patterns in how different government offices operate. The researchers found that the rate at which the appeals office overturned the original decisions varied significantly depending on which regional office issued the initial denial. Some offices had their decisions overturned more than half the time, while others were overturned far less frequently. This suggests that the experience or interpretation of the law can differ depending on where an application is processed. By having this information organized in a structured way, researchers can now study these institutional differences in a way that was previously impossible with raw text alone.
The creation of this resource opens the door for a new kind of research in legal artificial intelligence. Because the dataset includes the specific evidence presented, the legal rules applied, and the final judgment, it can be used to train computers to understand the logic of legal reasoning. It can help build systems that can predict the outcome of a case based on the facts, or more importantly, systems that can detect when a legal decision contains an error. Since the government is already beginning to use artificial intelligence in its decision-making processes, having a tool that can identify mistakes in human reasoning is a critical step toward ensuring that these new technologies are safe and fair. The researchers have made the entire dataset, along with the code used to build it, available to the public, inviting scientists and legal experts to explore the intricate machinery of administrative justice and to build tools that can help make it work better for everyone.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.