← Latest papers
💬 NLP

Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness

This paper introduces GAMUT, a novel benchmark and two-level meta-rubric framework designed to evaluate the factual completeness of long-form open-ended generations by converting structured content requirements into machine-gradable checklists, revealing that current frontier models struggle to achieve comprehensive coverage across diverse, evidence-backed domains.

Original authors: Xilun Chen, Zhaleh Feizollahi, Ross Goodwin, Seungwhan Moon, Scott Yih, Pinar Donmez, Babak Damavandi, Luna Dong

Published 2026-07-22
📖 7 min read🧠 Deep dive

Original authors: Xilun Chen, Zhaleh Feizollahi, Ross Goodwin, Seungwhan Moon, Scott Yih, Pinar Donmez, Babak Damavandi, Luna Dong

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Quest for the Complete Answer

Imagine you are teaching a robot to be a helpful assistant. You don't just want it to avoid lying; you want it to be genuinely useful. In the world of artificial intelligence, this is the difference between precision and recall. Think of precision as a strict teacher who only gives you a red "X" if you write something wrong. If you say, "The sky is green," the teacher marks it wrong. But if you say, "The sky is blue," the teacher gives you a gold star, even if you forgot to mention that it's also blue with white clouds, or that it changes color at sunset. That's precision: checking if what you did say is true.

Now, think of recall as a different kind of teacher who asks, "Did you tell me everything you know about the sky?" If you only said "The sky is blue," this teacher might say, "That's true, but you missed the clouds and the sunset. You didn't give me the full picture." For a long time, AI researchers have been great at the first teacher (precision), making sure robots don't hallucinate fake facts. But they've struggled with the second teacher (recall). They haven't had a good way to measure if a robot's answer is complete, especially when the answer isn't just a list of facts but a complex story with steps, relationships, and open-ended details. This paper steps into that gap, asking: "How do we grade an AI not just on whether it's right, but on whether it's thorough?"

Enter GAMUT: The "Did You Miss Anything?" Test

The researchers at Meta AI have built a new benchmark called GAMUT (Grounded Assessment of Multimodal Factuality). Think of GAMUT as a super-challenging trivia game designed to test if an AI can be a "deep researcher" rather than just a fact-reciter.

Here's the setup: Imagine you are wearing smart glasses. You look at a weird, flaky pastry in a bakery, or a strange car engine, or a specific type of plant. You ask your glasses, "What is this, and how is it made?" The AI has to look at the picture, go out to the internet, dig up information from multiple sources, and write a long, detailed answer.

The problem with previous tests was that they treated answers like a simple checklist. They would ask, "Did you mention rice?" Yes. "Did you mention tomatoes?" Yes. "Pass!" But real life isn't a flat list. Sometimes you need to know the order of steps (you can't fry the rice before you wash it). Sometimes you need to know that there are many valid ingredients, and you just need to name enough of them to be helpful, not every single one in existence. Sometimes you need to explain why something happens, connecting two facts together. A simple "yes/no" checklist fails to capture these nuances.

The Two-Level Magic Trick
To solve this, the authors invented a clever two-step grading system, which they call a Two-Level Meta-Rubric.

  1. Level 1: The Master Plan (The Meta-Rubric). First, a human expert (aided by a smart AI) creates a detailed "Master Plan" for what a perfect answer looks like. This plan isn't just a list; it's a structured map. It says, "Okay, for this question, you must identify the dish (Critical). You must list the core ingredients (Critical). You should mention at least two of these optional spices (Flexible). And you must describe the cooking steps in the exact right order (Process)." This level captures the messy, complex reality of a good answer.
  2. Level 2: The Checklist (The Binary Rubric). Here's the trick. The computer then mechanically translates that complex Master Plan into a long, flat list of simple "Pass/Fail" questions that a robot judge can easily grade. It turns "You need to cover enough optional spices" into a specific check: "Did the answer mention at least 2 of these 6 specific spices?" It turns "Do the steps in order" into a check: "Did the answer say 'fry' before 'simmer'?"

This way, the humans get to design a rich, nuanced test, but the grading is done by a robot checking simple boxes, which is much more reliable and consistent than asking a robot to "guess" how good an answer feels.

The Game: 1,813 Questions and Real Images

The team built a massive dataset called GAMUT containing 1,813 questions. These aren't made-up questions; they are based on real photos taken by people wearing smart glasses. The photos show everyday things: a specific type of shoe, a local landmark, a weird vegetable, or a car part. The questions are designed to be "everyday deep research"—things a curious person would actually ask while looking at something, but which require digging into the internet to answer fully.

They tested 14 different AI models (both the big, expensive ones from tech giants and the open-source ones) on this benchmark. The results were eye-opening.

The Findings:

  • It's Hard. Even the smartest model, Gemini 3.1 Pro, only got a score of 58.7%. That's barely a passing grade in high school! This means the best AI in the world is still missing a huge chunk of the information a complete answer should have.
  • Missing vs. Wrong. The biggest problem wasn't that the AIs were lying (contradicting facts). The biggest problem was omission. The models were skipping important details. They were like students who wrote a short essay when the teacher asked for a book. They got the main facts right (precision), but they didn't give the full story (completeness).
  • The "Tax" of Vision. The researchers also tested the models without the pictures (just text). When they did this, the scores went up by about 10 to 20 points for everyone. This shows that part of the difficulty was just "seeing" the object correctly. But even after removing the vision part, the models still had gaps in their knowledge. The ranking of the models stayed the same, proving that the test is really measuring how much the AI knows, not just how well it sees.
  • Robustness. They tested the results with three different "judge" AIs. The rankings stayed the same no matter who was doing the grading, which means the test is fair and reliable.

Why This Matters

This paper doesn't just say "AI is getting better." It says, "We have a new way to measure if AI is actually helpful." By moving beyond simple "is this fact true?" checks to "did you cover the whole topic?" checks, GAMUT reveals that even our most advanced models are still missing the forest for the trees. They can tell you a fact is true, but they often forget to tell you the whole story.

The authors suggest that for AI to truly be a useful assistant in our daily lives—helping us cook, fix things, or learn about the world—it needs to master factual completeness. GAMUT gives us the ruler to measure that progress, and right now, the rulers show we have a long way to go. The best models are only about 60% of the way there, leaving plenty of room for the next generation of AI to learn how to be truly thorough.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →