← Latest papers
🤖 machine learning

Coverage, Not Averages: Semantic Stratification for Trustworthy Retrieval Evaluation

This paper proposes "semantic stratification," a framework that formalizes retrieval evaluation as a statistical estimation problem by organizing documents into entity-based clusters and generating targeted queries to ensure comprehensive coverage and transparent identification of failure modes, thereby overcoming the inherent biases of current heuristic-based evaluation methods.

Original authors: Andrew Klearman, Radu Revutchi, Rohin Garg, Rishav Chakravarti, Samuel Marc Denton, Yuan Xue

Published 2026-04-23
📖 5 min read🧠 Deep dive

Original authors: Andrew Klearman, Radu Revutchi, Rohin Garg, Rishav Chakravarti, Samuel Marc Denton, Yuan Xue

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Problem: The "Average" Trap

Imagine you are a restaurant critic. You want to know if a new restaurant is good.

The Old Way (Current Benchmarks):
You go there, order five dishes: three are amazing steaks, one is a terrible salad, and one is a burnt soup. You calculate the "average" taste.

  • Result: The average score looks pretty good because the steaks were so delicious. You tell the world, "This restaurant is a 4-star gem!"
  • The Reality: The restaurant is actually terrible at making salads and soup. If you go there specifically for a salad, you'll be disappointed. The "average" hid the failure.

The Paper's Argument:
This is exactly what is happening with RAG (Retrieval-Augmented Generation) systems. These are AI tools that search a database to answer questions. Currently, we test them using a small list of questions (queries).

  • The questions are mostly about popular, easy topics (the "steaks").
  • They rarely ask about niche, difficult, or complex topics (the "salads and soups").
  • Because the AI gets the easy questions right, the average score looks high. But in the real world, when a user asks a hard question, the AI fails miserably. The average score is lying to us.

The Solution: "Semantic Stratification" (The Map)

The authors propose a new way to test these AI systems. Instead of just grabbing a random handful of questions, they want to build a complete map of the knowledge base first.

Think of the document collection (the database the AI searches) as a giant, unorganized library.

  1. The Old Way: You throw darts at the library to pick books to test. You might hit the "History" section 50 times and the "Quantum Physics" section zero times.
  2. The New Way (Stratification):
    • Step 1: Map the Library. The authors use AI to read every book and group them into "neighborhoods" based on what they are about (e.g., "Medical Procedures," "Financial Advice," "Climate Science").
    • Step 2: Check the Gaps. They look at the map and say, "Hey, we have 500 questions about History, but we have zero questions about Quantum Physics, even though there are 200 books on that topic!"
    • Step 3: Fill the Gaps. They use AI to generate new questions specifically for those empty neighborhoods.

This ensures that the test covers every neighborhood in the library, not just the popular ones.


How It Works: The "Neighborhood" Analogy

The paper breaks the library down into two types of "neighborhoods" to make sure the test is fair:

1. Semantic Neighborhoods (What the topic is)

  • Analogy: "The Medical District" vs. "The Finance District."
  • Goal: Make sure we have questions for every district, not just the ones with the most tourists.

2. Structural Neighborhoods (How hard the question is)

  • Analogy: Some questions are like asking, "Where is the nearest coffee shop?" (Easy, direct). Others are like asking, "How does the coffee shop's supply chain affect the price of beans in 2025?" (Hard, requires connecting many dots).
  • Goal: The authors found that some AI systems are great at simple questions but terrible at complex ones. By testing both, we see the full picture.

The Results: What Did They Find?

When they applied this new "Map-Based" testing to real-world data (like medical documents), they found shocking things that the old "Average" method missed:

  • The "Blind Spots": There were entire sections of the medical library (like "Statistical Methods" or "Clinical Trial Designs") containing thousands of documents that the old tests never touched. The AI was never tested on these, so we didn't know it was failing there.
  • The "Fake Winners": One AI system looked like the winner because it was great at the popular topics. But when they tested the "blind spots," that same AI failed completely. Another system was mediocre on popular topics but excellent on the hard ones.
  • The "Broken Dictionary": They found cases where the AI failed not because it was dumb, but because the library itself was messy. For example, a document might be about a specific disease, but the question used a different name for it. The AI couldn't connect the dots because the "dictionary" didn't match.

Why Does This Matter?

For the Average Person:
If you use an AI to help you with your health, your taxes, or your job, you want to know: "Will this AI work for my specific problem?"
The old testing method says, "Yes, it's 90% accurate!" (based on easy questions).
The new method says, "It's 90% accurate on general topics, but if you ask about [Specific Niche Topic], it will fail 100% of the time."

The Takeaway:
Don't trust the Average. Trust the Coverage.
Just because a car has a great top speed (average score) doesn't mean it can handle off-road terrain (the hard, niche questions). This paper gives us a way to test the "off-road" capabilities of AI so we don't get stuck in the mud when we need it most.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →