Figurative Justice: Detecting metaphors in Hindi judgements with qualitative assessment and transformers
This paper introduces the Hindi Legal Metaphor Corpus (HiLeMe) and a transformer-based detection model to address the lack of metaphor analysis in low-resource Hindi legal texts, aiming to decode judicial decision-making and expand automated legal discourse tools to other Indian languages.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Language in a courtroom is rarely just a neutral tool for listing facts. While laws are often imagined as a rigid set of rules written in plain ink, the people who write and interpret them—judges, lawyers, and legislators—frequently rely on figurative language to make sense of abstract ideas. They speak of "walls" separating church and state, "scales" balancing justice, or "roots" of a contract. These are not merely decorative flourishes; they are powerful cognitive tools that shape how legal concepts are understood, argued, and decided. When a judge describes a situation as a "war" or a right as an "object," they are doing more than using a colorful phrase; they are framing the reality of the case, potentially influencing the severity of a punishment or the interpretation of a human right. This is especially critical in languages like Hindi, where the nuances of such metaphors can shift the entire meaning of a legal argument, yet until now, no one had built a way for computers to spot them automatically in Indian court documents.
A team of researchers from Estonia and India has taken the first step toward solving this problem by creating a specialized dataset and a computer model designed to find metaphors in Hindi legal judgments. Their work, titled "Figurative Justice," addresses a significant gap in the field of artificial intelligence. While computers have become quite good at reading and understanding English, Spanish, and other major languages, they struggle with "low-resource" languages like Hindi, particularly when those languages are used in the complex, high-stakes environment of the law. The researchers started by gathering thousands of sentences from real district court judgments in the Indian state of Uttar Pradesh. These documents were written entirely in Hindi and contained the raw material of legal reasoning. To make sense of this text, the team enlisted six legal experts to act as human teachers for the computer. These experts read through the sentences and marked every instance where a word was being used in a figurative sense rather than its literal meaning, following a strict, established method for identifying metaphors. This process resulted in a new collection of data, which the researchers named the Hindi Legal Metaphor Corpus, containing over 7,000 sentences and nearly 162,000 words.
With this carefully labeled data in hand, the researchers trained a sophisticated computer model to recognize the difference between literal and figurative language. They used a type of artificial intelligence known as a transformer, which is a system capable of understanding the context of words within a whole sentence. The model was fed the Hindi sentences and taught to predict whether a specific part of the text contained a metaphor. The results showed that the computer could distinguish between literal and figurative language with an overall accuracy of about 72 percent. This is a promising start, proving that it is possible to teach a machine to see the hidden layers of meaning in Hindi legal texts. However, the study also revealed the difficulty of the task. When the model specifically looked for metaphors, it was correct about half the time it made a guess, and it missed more than two-thirds of the actual metaphors present in the text. This suggests that while the computer can spot obvious examples, it still struggles with the subtle, nuanced, and highly contextual metaphors that often appear in legal arguments.
The researchers found that the challenge lies partly in the nature of legal language itself, where metaphors are often deeply embedded in the structure of the argument rather than standing out as obvious comparisons. In their analysis of the mistakes the model made, they saw that the computer sometimes confused standard legal terms with metaphors, and at other times failed to recognize a metaphor because the context was too complex. The study does not claim to have solved the problem of detecting metaphors in Hindi law, but rather to have built the first reliable map of the terrain. By creating this dataset and demonstrating that a computer model can achieve a baseline level of success, the team has provided a foundation for future work. Their findings suggest that to truly understand judicial decisions in Hindi, and to ensure that the figurative language used by judges does not unintentionally bias outcomes, we need tools that can read between the lines. This research opens the door to applying similar techniques to the other twenty-one scheduled languages of India, aiming to make the hidden power of legal metaphors visible and open to critical review.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.