Prompting from the bench: Large-scale pretraining is not sufficient to prepare LLMs for ordinary meaning analysis
This paper argues that large-scale pretraining alone is insufficient to prepare large language models for reliable legal "ordinary meaning" analysis, as empirical evidence reveals their conclusions lack robustness against minor prompt variations and show only moderate correlation with human judgments, rendering them unsuitable for authoritative use in high-stakes judicial contexts.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a judge in a courtroom trying to figure out what a confusing word in a law actually means to the average person. For example, does a contract term like "landscaping" cover building a trampoline? Traditionally, judges might ask themselves, "What would my neighbor think?" or look up words in a dictionary.
Recently, some legal experts and judges have started asking Large Language Models (LLMs)—the same kind of AI that powers chatbots—to answer this question. The idea is: "If an AI has read almost everything ever written on the internet, it must know exactly how ordinary people use words, right?"
This paper is a big, careful experiment to test if that idea is true. The researchers acted like "tough critics" sitting on the bench, putting these AI models through a rigorous stress test. Here is what they found, explained simply:
1. The "Broken Clock" Problem
The researchers asked 15 different AI models (ranging from small ones to the massive GPT-4) to decide if 138 different insurance scenarios were covered or not.
- The Finding: Some of the AI models were like a broken clock that is stuck on "No." No matter what the story was about, they gave the exact same answer every single time.
- The Metaphor: Imagine asking a weather forecaster, "Is it raining?" If they say "Yes" for a sunny day, a cloudy day, and a snowstorm, they aren't actually looking at the sky. They are just repeating a script. Several of the AI models did exactly this, failing to actually "read" the specific details of the legal case.
2. The "Magic Trick" of Prompting
The researchers then tried a magic trick: they asked the exact same question, but just changed the way they asked it.
- Scenario A: "Is John covered? Yes or No?"
- Scenario B: "Is John not covered? Yes or No?"
- Scenario C: "Do you agree that John is covered?"
- The Finding: The AI's answer would flip-flop wildly. In one version of the question, the AI might be 90% sure John is covered. In a slightly different version (like adding a "not" or flipping the order of options), the AI might become 90% sure he is not covered.
- The Metaphor: It's like a suggestion box that changes its mind based on the color of the paper you write on. If you ask a model a question one way, it says "Yes." If you ask it the same question but put the "No" option first, it says "No." This means a lawyer could potentially "shop around" for the prompt that gives them the answer they want, rather than getting the truth.
3. The "Translator" Who Doesn't Speak the Language
The researchers compared the AI's answers to what real humans (ordinary English speakers) actually thought.
- The Finding: The AI and the humans were only moderately in agreement. Even the best AI models (like GPT-4) only got it right about 83% of the time if you knew exactly how to interpret their probability scores. But in a real courtroom, you don't get to peek at the scores; you just get the answer.
- The Metaphor: Imagine hiring a translator to tell you what a crowd of people is thinking. If the translator gets it right 83% of the time, that sounds good. But in a legal case, getting it wrong 1 in 6 times is a disaster. It's like a GPS that gets you to the right city 83% of the time but drops you in a swamp the other 17%.
The Bottom Line
The paper argues that just because an AI has read a lot of data (large-scale pretraining) doesn't mean it understands how ordinary people think.
The authors conclude that these AI models are currently too fragile and unreliable to be used as the final word in legal cases. They are not "authoritative" sources of ordinary meaning yet. If a judge relies on them without checking their work, they might be trusting a broken clock or a suggestion box that changes its mind based on the font size.
In short: The AI is a fluent talker, but it's not a reliable thinker when it comes to deciding what words mean to regular people.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.