Beyond Questions: Evaluating What Large Language Models (Actually) Know
This paper introduces "Beyond Questions" (BeQu), a new open knowledge evaluation paradigm that assesses large language models on the knowledge they naturally express through open-ended elicitation rather than predefined questions, revealing significant insights into model capabilities across various factors like scale and reasoning effort.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to figure out what a student actually knows about a famous historical figure, like Martin Luther King Jr.
The Old Way (The "Quiz Show" Approach)
For years, researchers have tested AI models like a strict quiz show host. They ask specific, narrow questions: "What was his birth date?" or "Who was his wife?"
- The Problem: This is like judging a chef only by asking them to make a grilled cheese sandwich. If the chef is a master pastry chef but terrible at sandwiches, the quiz says they are a bad cook. The test only measures what the quiz designer thought to ask. If the designer never asked about the chef's favorite color, the AI's knowledge of that fact remains invisible. This is called availability bias.
The New Way (The "Show and Tell" Approach)
This paper introduces a new method called Open Knowledge Evaluation. Instead of a quiz, imagine asking the AI: "Tell me everything you know about Martin Luther King Jr."
- The Goal: We don't just check if the AI gets the right answer to a specific question. We look at the entire story it tells. We check:
- Precision: Did it make up facts (hallucinations)?
- Recall: Did it remember everything it knows, or did it leave out important details?
The Tool: BeQu (Beyond Questions)
To test this, the authors built a massive playground called BeQu.
- They picked 10,000 random things (people, places, events) from Wikipedia.
- They built a "Reference Library" for each one, gathering facts from Wikipedia and the web.
- They asked 20 different AI models to "tell everything they know" about these 10,000 things.
- Then, they used a super-smart AI judge to compare the stories the models told against the Reference Library to see what was true, what was false, and what was missing.
What They Found (The Results)
Big Models Still Win (But Open Ones are Catching Up):
Think of the AI models as students. The "commercial" students (paid, big models like Claude Opus) are still getting the highest grades. However, some "open-source" students (free models) are scoring remarkably close, proving they are very competitive.Thinking Hard Doesn't Help Much:
You might think, "If I tell the AI to 'think harder' or 'reason more' before answering, it will know more." The paper found this is mostly a myth for this specific task. Whether the AI spent 1 second or 10 seconds "thinking," the amount of facts it pulled out didn't change much. The knowledge was already there; it just needed to be spoken.The "Strict Teacher" vs. The "Creative Writer":
The researchers tried forcing the AI to answer in a strict format (like a spreadsheet with specific columns).- Result: The AI became very accurate (high precision) but stopped sharing many facts (low recall). It was like a student who only answers exactly what the teacher asks, refusing to add any extra context.
- Conversely, when allowed to be free and open, the AI shared more facts, even if it occasionally made a small mistake.
Asking for More Gets You More (Until it Doesn't):
When the researchers asked the AI to list "5 facts" vs. "100 facts," the AI with the larger request usually shared more knowledge. However, if they pushed the AI too hard to list hundreds of facts, it started making things up (hallucinating) just to fill the quota.The "Fake Person" Test:
They asked the AI about things that don't exist (like the "International Airport of Andorra"). The best models simply said, "I don't know anything about that." The weaker models started inventing fake details, showing they were guessing rather than knowing.
The Bottom Line
This paper argues that we need to stop treating AI like a multiple-choice test machine. To truly understand what an AI knows, we need to let it talk freely and see how much of the real world it can describe, how accurately it describes it, and what it chooses to leave out. The new "BeQu" benchmark is a tool to measure this "Show and Tell" ability, revealing that while AI is getting smarter, it still has a lot of hidden knowledge it isn't always choosing to share.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.