Precision Is Not Faithfulness: Coverage-Aware Evaluation of Grounded Generation with a Complete Oracle
This paper introduces a coverage-aware evaluation framework for grounded generation that exposes the limitations of existing precision-only faithfulness metrics by demonstrating, through complete-oracle benchmarks in Formula 1 and weather domains, that current models achieve high faithfulness by under-reporting relevant facts, and proposes a unified score and verifier-guided method to improve both precision and recall.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Problem: "Silence is Golden" (But Only Half True)
Imagine you are a teacher grading a student's essay. You have a strict rule: "Every sentence you write must be 100% true."
If the student writes a 500-word essay but gets one fact wrong, they fail. But here is the trick: If the student writes only one sentence that is true, and says absolutely nothing else, they get a perfect score of 100%.
This is the problem the paper identifies with current AI evaluation tools. These tools measure Precision (how many of the things the AI did say were true). They reward the AI for being super cautious and saying very little. It's like a student who refuses to answer any questions unless they are 100% sure, resulting in a blank test paper that somehow gets an "A" because nothing on it is wrong.
The paper argues that Faithfulness (being a good, honest AI) shouldn't just mean "don't lie." It should also mean "don't leave out the important stuff."
The Solution: The "Complete Answer Sheet"
To prove this, the researchers needed a way to measure not just what the AI got right, but also what it missed. This is called Recall.
Usually, this is impossible. If you ask an AI to write a story about a random day in Paris, how do you know if it missed a detail? There is no "answer key" for a creative story.
The Researchers' Trick: Formula 1 Racing
The researchers used Formula 1 (F1) racing data as their "answer key."
- The Analogy: Think of an F1 race like a math problem with a single, undeniable solution. We know exactly which lap a driver pitted, what tires they used, and who they passed. There is no guessing.
- The "Complete Oracle": Because the data is so precise, the researchers could create a "Complete Answer Sheet" for every race decision. They knew exactly every single fact that should have been mentioned in a perfect explanation.
This allowed them to grade the AI on two things:
- Precision: Did the facts it did say match the answer key? (Is it lying?)
- Recall (Coverage): Did it mention all the facts that were on the answer key? (Is it being helpful?)
The Shocking Discovery: The "Smart" AI Was Actually Lazy
The researchers tested the world's most advanced AI models (like Grok, GPT-5, and Gemini) on explaining F1 race strategies.
- The Result: The model with the highest "Precision" score (the one that never lied) was actually the worst at being helpful.
- Why? It was playing it safe. It would say, "The driver pitted on lap 12." (True! 100% score). But it would skip the fact that the driver switched to hard tires, or that they lost 5 seconds, or that they defended against a rival.
- The Flip: When the researchers added a penalty for missing facts (Recall), the rankings changed completely. The models that were slightly less "perfect" but much more talkative and detailed became the winners.
The Metaphor:
Imagine two doctors giving you a diagnosis.
- Doctor A says: "You are alive." (100% accurate, 0% helpful).
- Doctor B says: "You have a broken leg, a sprained ankle, and a concussion. Here is the treatment for each." (95% accurate, 100% helpful).
Current AI tools would give Doctor A a perfect score and Doctor B a lower score because Doctor B took a tiny risk of being slightly wrong on one detail. This paper says: Stop rewarding Doctor A.
The Weather Test: It's Not Just About Racing
To make sure this wasn't just a weird quirk of F1 data, they tested the same idea on Weather Forecasts.
- They gave the AI real weather data (temperature, wind, rain chance) and asked it to write a forecast.
- The Result: The same thing happened. The most "precise" AI was the one that said the least. The most "faithful" (accurate + complete) AI was the one that covered all the facts, even if it made a tiny mistake here or there.
The Fix: A "Verifier" That Guides the AI
The paper also offers a solution. Instead of just asking the AI to "write a story," they built a system where:
- The AI writes a draft.
- A "Verifier" (a strict checker) looks at the draft against the data.
- The Verifier tells the AI: "You got this fact right, but you forgot to mention the wind speed, and you lied about the rain."
- The AI fixes it and tries again.
This "Verifier-Guided" method helped the AI become both accurate and complete without needing a human to rewrite the text.
Summary of Key Takeaways
- Precision Faithfulness: Just because an AI doesn't lie doesn't mean it's doing a good job. If it says nothing, it can't lie, but it's useless.
- The "Abstention" Trap: Current AI benchmarks reward models for staying silent and avoiding risks.
- The Complete Oracle: You can only measure if an AI is "missing" facts if you have a perfect, complete list of what should be there (like F1 data or weather records).
- The Ranking Flip: When you measure both accuracy and completeness, the "best" AI changes. The models that are brave enough to say more (and cover more facts) are actually the best, even if they aren't perfect.
- Language Doesn't Matter: This problem happens in English, Spanish, and Portuguese equally.
In short: We need to stop grading AI only on whether they tell the truth, and start grading them on whether they tell the whole truth.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.