Testing Frontier LLMs on Indian Statutory Interpretation: The Indian Construction Canon Benchmarking (ICCB) -Pilot Benchmark
This paper introduces the ICCB-Pilot benchmark to evaluate frontier LLMs on Indian statutory interpretation, revealing that while models can often predict correct legal outcomes, they fail to correctly identify and apply the underlying canons of construction, suggesting their reasoning relies on surface pattern-matching rather than genuine jurisprudential understanding.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the world of law, reading a statute is rarely as simple as reading a dictionary definition. When a judge faces a vague or ambiguous law, they do not merely guess the answer; they follow a specific set of mental tools known as "canons of construction." These are established rules of thumb that guide how a text should be understood. For instance, one rule might say that if a law is meant to punish someone, it should be read very strictly, word for word. Another rule might say that if a law is designed to help the poor, it should be read broadly to include as many people as possible. These tools are the difference between a judge simply recalling a past case and a judge actually reasoning through a new problem. As artificial intelligence systems begin to enter courtrooms and legal offices, a critical question arises: do these machines truly understand how to use these tools, or are they just very good at guessing the right answer based on patterns they have seen before?
A team of researchers from the Manipal Academy of Higher Education in India set out to test this question using the latest generation of artificial intelligence models. They created a specialized test, which they call a benchmark, focused entirely on how these machines interpret Indian laws. Instead of asking the computers to simply predict the outcome of a case, the researchers asked them to explain their thinking. The test presented the models with forty different legal scenarios drawn from real Indian court judgments. In each scenario, the model had to identify which specific rule of interpretation should be used, explain that rule correctly, and then apply it to the facts to reach a conclusion. The researchers evaluated the models on five distinct aspects: whether they picked the right rule, whether they described that rule accurately according to Indian legal tradition, whether they applied it correctly to the facts, whether they reached the right final answer, and how clear their overall reasoning was.
The results revealed a striking gap between what the machines could do and how they did it. The most powerful artificial intelligence models tested were surprisingly good at getting the final answer right. In many cases, they arrived at the same conclusion a human judge would have reached. However, when the researchers looked at how the models got there, the picture changed. The models frequently failed to correctly identify the specific legal rule they were using. They often reached the correct outcome without knowing the right tool to use, suggesting they were matching the situation to a familiar pattern rather than following a logical path. It is as if a student could solve a complex math problem correctly but could not explain which formula they used or why it worked. The study found that when the researchers explicitly told the models which rule to use, the models became much better at naming that rule, but their ability to explain it or apply it did not improve significantly. This suggests that the models possess the knowledge of these legal rules but struggle to access it on their own without a direct prompt.
The researchers also observed that the models behaved differently depending on how the question was asked. When left to figure out the problem on their own, the models often hesitated or refused to answer, particularly the open-source model tested. However, when the researchers gave them a hint about the type of rule to use, the models were more willing to engage. Despite this improvement in willingness, the core issue remained: the models were still better at guessing the right outcome than at constructing the correct legal argument. One model, in particular, consistently reached the correct outcome more often than it correctly identified the legal principle, yet it still made significant errors in its final conclusions. This dissociation between a correct answer and a correct reasoning process is a significant finding. It implies that if these systems are deployed in real legal settings based only on their ability to predict outcomes, they might provide correct answers for the wrong reasons, which could be dangerous in a system where the reasoning itself is just as important as the result.
The study concludes that while these artificial intelligence systems are powerful tools, they have not yet mastered the specific kind of reasoning required for legal interpretation. They appear to rely heavily on pattern recognition rather than the step-by-step application of legal rules that human judges use. The researchers emphasize that this does not mean the technology is useless, but it does mean that we cannot trust these systems to reason like lawyers just because they get the right answer. The study serves as a pilot, a small-scale test designed to prove that this specific type of evaluation is possible and necessary. The team plans to expand this work in the future, testing more models and a wider variety of laws to see if these patterns hold true across the board. For now, the findings suggest a clear warning: in the complex world of law, getting the right answer is not enough; the machine must also be able to show its work.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.