LLMpedia: A Transparent Framework to Materialize an LLM's Encyclopedic Knowledge at Scale
LLMpedia reveals that large language models' true encyclopedic knowledge is significantly lower than standard benchmarks suggest by generating a fully open, retrieval-free encyclopedia of one million articles that exposes substantial factuality gaps and subject biases across model families.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a brilliant, encyclopedic friend who has read almost every book in the world. You ask them a specific question, like "Who won the 1998 World Cup?" and they answer perfectly. You ask another, "What is the capital of France?" and they get it right again. Based on these tests, you might conclude your friend knows everything with near-perfect accuracy.
But what happens when you ask them to write a full, 2,000-word biography of a random historical figure, or explain a complex topic they've never been directly asked about before? Do they still know everything, or do they start making things up?
This is the story of LLMpedia, a new project that decided to stop asking simple trivia questions and instead asked AI models to write their own encyclopedia from scratch.
The Problem: The "Quiz Show" Trap
For years, we've tested AI models (like the ones powering chatbots) using "benchmarks." Think of these like standardized tests or quiz shows. The researchers pick a list of questions (e.g., "What is the chemical symbol for gold?") and see how many the AI gets right.
The paper argues that these tests are misleading. They suffer from "Availability Bias." It's like testing a chef only on their ability to make a perfect omelet. If they ace the omelet test, you assume they are a master chef. But if you ask them to cook a full, complex banquet from memory without a recipe, they might burn the soup or serve you a shoe.
The benchmarks only test what the researchers thought to ask. They miss the vast, uncharted territory of what the AI actually knows but has never been asked to explain.
The Experiment: Building a Library from Memory
To fix this, the researchers built LLMpedia. Instead of asking the AI specific questions, they gave it a seed (a starting topic, like "Vannevar Bush," a real-life inventor) and said: "Write an encyclopedia article about this person. Then, look at the links inside your article, pick a new person mentioned there, and write an article about them. Keep doing this."
They did this without letting the AI look up answers on the internet. They forced the AI to rely entirely on its internal memory (its "parametric knowledge").
- The Scale: They generated nearly 1 million articles across three different AI models.
- The Method: It's like a game of "telephone" where the AI passes the baton from one topic to another, expanding outward like a tree growing branches.
The Shocking Results
When they checked the facts in these million articles, the picture looked very different from the "90%+ accuracy" seen in quiz shows.
- The Drop in Accuracy: For the top model (gpt-5-mini), the accuracy on topics that did have Wikipedia pages dropped from the expected 90%+ to just 74.7%. That's a massive gap.
- The "Frontier" Problem: When the AI wrote about topics that didn't exist on Wikipedia (the "frontier"), the accuracy plummeted to 63.2%. The AI was confidently making up facts about things it didn't actually know.
- The "Capture Trap": The researchers compared their open, transparent AI encyclopedia to Grokipedia (a famous AI encyclopedia created by xAI).
- Grokipedia sounded very similar to Wikipedia (like a student copying the teacher's homework). It was very similar in wording but had a high rate of factual errors.
- LLMpedia sounded different and unique (like a student writing their own essay). It was less similar to Wikipedia but much more factually accurate.
- The Analogy: Grokipedia is like a student who memorizes the textbook so well they can recite it, but if you ask a slightly different question, they get confused. LLMpedia is like a student who actually understands the concepts, even if they explain them in their own words.
Why This Matters
The paper reveals a "Capture Trap." If an AI is too similar to Wikipedia, it might just be copying Wikipedia's phrasing rather than using its own knowledge. This hides the fact that the AI might be hallucinating (making things up) or getting confused about who is who.
By making their process completely open (showing every prompt, every draft, and every fact-check), the LLMpedia team proved that:
- Benchmarks lie: High scores on quizzes don't mean the AI is reliable in real-world, long-form writing.
- Memory is fragile: AI models are great at answering specific questions but struggle to weave a consistent, factual narrative over a long text.
- Transparency is key: We need to see how AI generates information, not just what it says, to trust it.
The Takeaway
Think of LLMpedia as a "stress test" for AI memory. It showed us that while our AI friends are great at trivia, they are still prone to making up stories when asked to write a full book from scratch. The paper calls for a new way of testing AI: not just by asking questions, but by watching them create knowledge and seeing if they can keep their facts straight when no one is holding a quiz sheet over them.
In short: Don't trust the AI just because it aced the test. Watch it write a story, and you might find it's still learning how to tell the truth.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.