LLMs Should Not Yet Be Credited with Decision Explanation
This position paper argues that LLMs should not yet be credited with explaining human decisions because current evidence primarily supports prediction and rationalization rather than genuine explanation, advocating instead for a calibrated standard that distinguishes between these capabilities to prevent premature redefinitions of explanatory progress.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to figure out why a person made a specific choice, like buying a red car instead of a blue one.
Now, imagine you have a very smart, well-read robot (an LLM) that can look at the situation and tell you two things:
- The Prediction: "I bet they bought the red car."
- The Story: "They bought the red car because they love the color red and it matches their personality."
The paper by Wenshuo Wang argues that we are currently making a big mistake. We are treating the robot's story as if it is a scientific explanation of the human's mind. The author says: Stop giving the robot credit for "explaining" human decisions until it has much stronger proof.
Here is the breakdown using simple analogies:
1. The Three Levels of "Credit"
The paper says there are three different things a robot can do, and they require different levels of proof. Think of it like a video game with three difficulty levels:
Level 1: The Fortune Teller (Decision Prediction)
- What it does: The robot guesses the outcome correctly. "They picked the red car."
- The Proof: It just needs to be right often.
- The Paper's View: This is great! We should give the robot credit for being a good guesser. But guessing right doesn't mean you know why they picked it. Maybe the robot just noticed that most people pick red on Tuesdays, not because of the person's love for red.
Level 2: The Smooth Talker (Rationale Generation)
- What it does: The robot writes a convincing story. "They picked the red car because it's stylish and matches their shoes."
- The Proof: Humans read the story and say, "That sounds reasonable and makes sense."
- The Paper's View: This is also useful! The robot is good at writing. But a smooth story isn't the same as the truth. A lawyer can write a perfect defense for a client who is actually guilty. The story sounds good, but it doesn't prove what was actually happening in the client's mind.
Level 3: The Mind Reader (Decision Explanation)
- What it does: The robot proves it actually tracked the specific mental gears turning inside the human's head that led to the choice.
- The Proof: This is the hard part. You can't just ask the robot to write a story. You have to test it.
- The Paper's View: We are not there yet. Currently, the robot is just a Level 1 Fortune Teller and a Level 2 Smooth Talker. We are mistakenly calling it a Level 3 Mind Reader.
2. The Problem: "The Persuasive Narrator"
The paper warns that we are falling for a trick. Because the robot is so good at writing (Level 2) and so good at guessing (Level 1), we assume it must also understand the cause (Level 3).
The Analogy:
Imagine a magician who pulls a rabbit out of a hat.
- Prediction: The magician says, "I will pull a rabbit out." (He is right).
- Rationale: He says, "I pulled the rabbit out because I used a special magic spell." (This sounds plausible and magical).
- Explanation: To really explain it, you'd need to show the secret compartment in the hat or the mechanism he used.
Right now, LLMs are just the magician telling us the spell. They haven't shown us the secret compartment. They are just "rationalizing" (making up a good reason) for the outcome they already predicted.
3. The Solution: The "Bridge Standard"
The author proposes a new rule for when we can finally say, "Okay, this robot does explain human decisions." The robot needs to build a bridge between its story and the actual human mind.
To cross this bridge, the robot's claim must pass four tests:
- Name the Target: Don't just say "It explains the choice." Say exactly what it is tracking. Is it tracking "fear of losing money"? Is it tracking "attention to bright colors"? You have to pick a specific thing to measure.
- Beat the "Fake" Alternatives: You have to prove the robot isn't just a "Smart Guess + Smooth Talker." You need to show that if you trick the robot (e.g., hide the answer first), it can't just make up a story that happens to be right. It must prove it found the real reason, not just a convenient excuse.
- Use the Right Tools: If you claim the robot tracks "attention," you shouldn't just ask it to write an essay. You should use tools like eye-tracking or reaction times to see if the robot's story matches what the human's eyes actually did.
- Keep it Small: Don't say "This robot explains all human decisions." Say "This robot explains this specific type of choice for this specific group of people."
4. Why This Matters
The paper isn't saying robots are useless. It's saying we need to be honest about what they are good at.
- Current State: We are treating robots like Psychic Storytellers. We think they are revealing the deep secrets of the human mind because they write convincing stories.
- Proposed State: We should treat them as Powerful Tools that are great at guessing and generating ideas, but we shouldn't call them "Explaners" until they pass the hard tests.
The Bottom Line:
If we keep giving robots credit for "explaining" things they haven't truly proven, we might stop looking for the real answers. By calibrating our credit (giving them credit for what they actually do, not what we wish they do), we can turn them from persuasive storytellers into reliable instruments for discovering the truth about human behavior.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.