Measuring Judgment Quality in Natural-Language Explanations: Evidence from Forecasting Tournaments
This paper introduces Explanation Quality Markers (EQMs), a scalable, LLM-based method for scoring sixty reasoning patterns in natural-language explanations that effectively predicts forecasting accuracy and outperforms traditional text-analysis techniques.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a manager trying to hire the best financial advisor. You have a stack of 55,000 applications. Each application has a specific prediction (e.g., "There is a 20% chance the stock will rise") and a written paragraph explaining why they think that.
Your problem? You can't read all 55,000 paragraphs to see who is actually good at guessing the future. You need a way to quickly scan the text and say, "This explanation looks smart," or "This one looks like a guess."
This paper introduces a new tool called Explanation Quality Markers (EQMs) to solve that problem. Here is how it works, using simple analogies.
The Old Way: Counting Words and Checking Dictionaries
Before this study, researchers tried to judge these explanations using "pre-LLM" tools. Think of these like a spell-checker or a word counter.
- They would count how many words were in the paragraph.
- They would look for specific "smart-sounding" words (like "therefore" or "however").
- They would check if the text was complex.
The Result: These tools were like trying to judge a chef's cooking skills by counting how many spices they used. It didn't really tell you if the food tasted good. The paper found these old methods had almost no connection to whether the person was actually accurate.
The New Way: The "Super-Reader" AI
The authors used a Large Language Model (an advanced AI, like GPT-4o) to act as a Super-Reader. Instead of just counting words, this AI was trained to look for 60 specific "reasoning patterns" that experts know are good or bad.
Think of the AI as a detective looking for clues in a suspect's story.
- Good Clues (Green Flags): Does the person look at past data? Do they admit they might be wrong? Do they break a big problem into small, manageable pieces?
- Bad Clues (Red Flags): Is the person just guessing? Are they overconfident ("It will definitely happen!")? Are they ignoring evidence that contradicts their view?
The AI reads the explanation and gives it a score based on how many "Green Flags" and "Red Flags" it finds.
The Big Discovery: The "Trash Detector"
The researchers tested this AI against the actual results of the predictions (did the stock go up or down?). Here is what they found:
- The AI is a Great "Trash Detector": The AI is incredibly good at spotting bad explanations. If an explanation is full of "Red Flags" (like guessing or overconfidence), the AI flags it, and that person is almost always wrong.
- The AI is a Weak "Star Finder": The AI is less good at spotting the absolute best experts. Just because an explanation looks perfect doesn't guarantee the person is a genius.
- Analogy: Imagine a metal detector at a beach. It is amazing at finding the rusty nails and broken glass (the bad stuff). But it isn't very good at distinguishing between a shiny piece of copper and a real gold nugget (the very best stuff).
How It Compares to Humans
The researchers also asked actual humans to read these explanations and rate them.
- Humans are easily fooled by length: Humans tended to give higher scores to longer explanations. They thought, "Wow, this person wrote a lot, so they must be smart."
- The Reality: The length of the text had almost nothing to do with whether the prediction was correct.
- The AI vs. Humans: The AI was much better at predicting who was accurate than the humans were. The humans were distracted by the "fluff" (length and fancy words), while the AI focused on the actual logic.
Why This Matters
This tool is useful because it requires no track record.
- Usually, to know if a forecaster is good, you have to wait months or years to see their past results.
- With EQMs, you can look at a single written explanation and get a good idea of whether that person is likely to be accurate or not.
Summary
- The Problem: We have too many expert explanations to read, and we don't know which ones are trustworthy.
- The Solution: An AI that scans for 60 specific reasoning habits (like checking for bias or using data).
- The Result: The AI is much better than old computer methods and better than human readers at spotting bad reasoning. It helps decision-makers filter out the "noise" and the "guessers," even if it can't perfectly identify the "geniuses."
Note: The paper strictly tested this on geopolitical and economic forecasting tournaments. It does not claim this method works for medical diagnoses, legal judgments, or other fields yet; it only proves it works for predicting future events in these specific tournaments.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.