A Short Survey of Viewing Large Language Models in Legal Aspect
This survey explores the integration of large language models into the legal field by examining their applications in tasks like judgment prediction and document analysis, while also addressing associated challenges such as privacy, bias, and explainability, and outlining future directions for specialized legal AI.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a world where the law is not just a set of rules written in stone, but a living, breathing library that changes every day. In this library, a new kind of librarian has arrived: a machine that can read millions of documents, understand complex arguments, and write answers to legal questions in seconds. This machine is a large language model, a type of artificial intelligence trained on vast amounts of text. For years, researchers have asked if these machines can solve specific legal puzzles, like predicting a court outcome or finding a relevant case. But the real question is no longer just about solving a puzzle; it is about whether we can trust the machine's answer when a person's freedom, money, or rights depend on it. The law is unique because it relies on authority. A legal answer is only valid if it points to the right source, from the right place, at the right time, and if that source actually supports the conclusion. If a machine gives a confident answer based on an old rule or a law from a different country, it is not just wrong; it is dangerous.
A researcher from China has taken a fresh look at how these machines are being used in the legal field. They reviewed hundreds of studies published between 2022 and 2025 to understand what these systems can actually do and, more importantly, where they might fail. Instead of just looking at how well a model answers a single question, they proposed a new way to think about the entire system. They call this an "authority-grounded" system. In simple terms, this means the machine must not only generate text but also prove its work. It must show exactly which documents it used, confirm that those documents are still valid, and explain how they support its answer. The researcher found that current studies often measure only one part of this process, such as how fast the machine retrieves information or how well it mimics human writing, without checking if the information is legally sound.
The researcher analyzed 327 technical studies to build a clear map of the field. They discovered that the tools used to make these machines smarter are not enough on their own. For example, a system might be excellent at finding relevant documents, but if it cannot tell the difference between a binding law and a mere suggestion, or if it cannot check the date of a law to see if it has been changed, the system is not ready for real-world use. The study highlights that accuracy in a test does not equal reliability in practice. A machine might get the right answer on a quiz but fail to explain why, or it might cite a source that looks correct but is actually outdated or from a different legal jurisdiction. The researcher argues that we need to stop treating these machines as simple answer generators and start treating them as components of a larger, more careful workflow.
To fix these gaps, the author suggests a new framework that connects three main ideas: what the machine is asked to do, how it is built to do it, and what proof we have that it is safe to use. They found that many existing systems lack the ability to "abstain," or admit when they do not know the answer, which is a critical skill in law. They also noted that while some systems use special tools to check their work, these tools often fail to catch subtle errors, such as a law that has been overruled by a newer court decision. The researcher proposes that future systems must include strict checks for the source of every piece of information, the time it was valid, and the specific legal area it applies to. They also emphasize that humans must remain in the loop, ready to review the machine's work and take responsibility for the final decision.
The paper does not claim that these machines are ready to replace lawyers or judges. Instead, it suggests that we are at a turning point where we must define exactly what it means for an artificial intelligence to be trustworthy in a legal setting. The researcher points out that we cannot simply rely on a high score from a test to prove a system is safe. We need a new kind of report card that shows not just the answer, but the entire path the machine took to get there, including the sources it checked and the rules it followed. This approach would allow courts, lawyers, and the public to see exactly where the system might be weak and where human oversight is needed. The study concludes that while the technology is advancing quickly, the rules for using it safely have not caught up. The path forward requires building systems that are transparent, accountable, and designed to handle the complex, changing nature of the law, ensuring that when a machine speaks, it speaks with the weight of verified truth behind it.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.