← Latest papers
🤖 AI

The Chronicles of RiDiC: Generating Datasets with Controlled Popularity Distribution for Long-form Factuality Evaluation

This paper introduces RiDiC, a configurable multilingual dataset of 3,000 entities across varying popularity tiers designed to evaluate and reveal hallucinations in long-form factuality generation by frontier LLMs.

Original authors: Pavel Braslavski, Dmitrii Iarosh, Nikita Sushko, Andrey Sakhovskiy, Vasily Konovalov, Elena Tutubalina, Alexander Panchenko

Published 2026-04-02
📖 4 min read☕ Coffee break read

Original authors: Pavel Braslavski, Dmitrii Iarosh, Nikita Sushko, Andrey Sakhovskiy, Vasily Konovalov, Elena Tutubalina, Alexander Panchenko

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a super-smart librarian (an AI) who can answer any question you ask. If you ask, "Who is the President of the United States?" or "What is the capital of France?", the librarian answers instantly and correctly. These are the "famous" questions, the ones everyone knows.

But what happens if you ask, "Tell me everything about the Little River of Gloom in a remote valley in Peru," or "Describe the specific features of a 1998 car model that only 50 people ever bought"?

This is where the story of RIDIC begins.

The Problem: The Librarian's "Famous Person" Bias

The researchers behind this paper noticed something tricky about AI. When an AI talks about famous things (like the Nile River or the Toyota Camry), it's usually accurate. But when it talks about obscure, "long-tail" things (like a tiny river or a rare car), it starts to make things up. It hallucinates. It sounds confident, but it's lying.

Most tests for AI only check if it knows the famous stuff. It's like testing a chef only on how well they can make a grilled cheese sandwich. You never know if they can actually cook a complex, rare dish until you ask them to.

The Solution: Building a "Popularity Menu"

The authors created a special tool (a pipeline) to build a new testing ground called RIDIC. Think of this as a menu designed specifically to test the AI's knowledge across the entire spectrum of popularity.

They didn't just pick random things. They carefully selected 3,000 items from three specific categories:

  1. Rivers (Nature's highways)
  2. Natural Disasters (Earth's tantrums)
  3. Car Models (Human engineering)

For each category, they picked items in three "tiers" of fame:

  • The Head (The Stars): Famous things everyone knows (e.g., The Amazon River, Hurricane Katrina, the Honda Civic).
  • The Torso (The Middle Class): Moderately known things (e.g., a specific river in Michigan, a flood in Spain, a specific Cadillac model).
  • The Tail (The Ghosts): Extremely obscure things that very few people know about (e.g., a tiny creek in Russia, a minor landslide in 1977, a rare French car).

They did this in two languages: English and Chinese, to see if the AI struggled more in one language than the other.

The Experiment: Putting the AI to the Test

Once they built this menu, they asked three different AI "librarians" (Llama, Qwen, and GPT-5) to write a detailed story about each item on the list.

Then, they acted as strict fact-checkers. They compared the AI's stories against the "source of truth" (Wikipedia and other databases). They broke the stories down into tiny facts and checked each one: Is this true? Is this made up?

The Findings: What They Discovered

The results were like a report card for the AI, revealing some surprising grades:

  1. The "Long-Tail" Trap: As the items got less famous, the AI got worse. It's like a student who aces the math test but fails the history test because they only studied the chapters everyone else studied. The AI struggled significantly with the "Tail" items, often inventing facts to fill the silence.
  2. The Language Gap: The AI was much more accurate in English than in Chinese. It's as if the librarian has a massive, well-organized English library but only a few scattered, dusty pamphlets in Chinese. When the AI tried to write about obscure Chinese topics, it hallucinated even more.
  3. The "Big Brain" Advantage: The biggest AI (GPT-5) was the most accurate, but even it wasn't perfect. The smaller AIs (Llama and Qwen) struggled much more with the rare items, showing that having a bigger "brain" (more parameters) helps with remembering obscure facts.
  4. The Domain Difference: The AI was surprisingly good at talking about cars and disasters but terrible at talking about rivers. It seems the AI just knows more about cars and storms than about waterways!

Why This Matters

This paper is like giving the AI industry a new kind of mirror. Instead of just seeing how the AI looks when it's wearing its "famous" hat, we can now see how it looks when it's dealing with the obscure, the rare, and the difficult.

By releasing this dataset and the tools to create it, the authors are saying: "Don't just test your AI on the easy stuff. If you want to trust it with real-world problems, you need to know if it can handle the obscure stuff without making things up."

In short, RIDIC is a spotlight that shines on the dark corners of an AI's knowledge, ensuring that when the lights go down and the famous topics fade away, the AI doesn't start making up stories in the dark.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →