← Latest papers
💬 NLP

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy

This paper reveals that large language models often exhibit significant instability in their answers when faced with meaning-preserving paraphrases, demonstrating that standard accuracy metrics can mask reliability issues and that self-paraphrasing strategies can effectively recover latent knowledge to improve performance.

Original authors: Kazem Faghih, Yize Cheng, Shoumik Saha, Mobina Pournemat, Armin Gerami, Soheil Feizi

Published 2026-07-28
📖 5 min read🧠 Deep dive

Original authors: Kazem Faghih, Yize Cheng, Shoumik Saha, Mobina Pournemat, Armin Gerami, Soheil Feizi

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Illusion of the Perfect Robot

Imagine you are talking to a very smart robot friend who has read almost every book in the library. You ask it a simple question, and it gives you a brilliant, correct answer. You feel confident, right? But what if you asked the exact same question again, just using slightly different words? Would it still give you the same answer? This is the big question behind a new study from computer scientists at the University of Maryland. They are looking at "Large Language Models" (LLMs), which are the super-smart AI brains behind tools like chatbots. These models are famous for being accurate on tests, but this study asks a tricky question: Are they actually reliable, or are they just good at guessing based on how a question is phrased? Think of it like a student who memorizes the exact wording of a practice test. If you change the question just a little bit, does the student still know the answer, or do they get confused? The researchers wanted to find out if these AI models truly understand the world or if they are just playing a game of "word matching."

The Same Question, The Wrong Answer

The researchers decided to play a game with 13 different AI models, ranging from smaller, open-source ones to the massive, powerful models used by big tech companies. They took questions about real-world facts (like "How many chess games did Anderssen lose in 1864?") and math problems, and then they asked the models the exact same questions but rewritten in dozens of different ways. They made sure the new questions meant the exact same thing, just like saying "How are you?" versus "What's up?" or "How's it going?"

Here is the surprising twist they found: The AI models were often inconsistent. Even though the questions meant the same thing, the models would sometimes get the answer right the first time, and then get it wrong when the question was rephrased. In fact, for many questions, the models flipped between being right and wrong depending on the wording. The study found that mismatch rates—where the model gives a different answer to the same question—could reach more than 23%. That means if you asked a model the same question 100 times in different ways, it might give you a different (and wrong) answer nearly a quarter of the time, even though it knew the right answer all along!

The "Hidden Knowledge" Gap

The most fascinating part of the story is what this says about what the AI actually "knows." The researchers discovered that the models often did have the correct answer hidden inside them. If you asked the same question in enough different ways, the model would eventually get it right at least once. This suggests the knowledge is there, but the model is unreliable at retrieving it.

To visualize this, imagine a student taking a test.

  • Standard Accuracy: The student gets 80% right on the first try.
  • Reliable Capability: The student gets 80% right every single time you ask the question, no matter how you phrase it.
  • Latent Capability: The student gets the right answer at least once if you ask them 10 different ways.

The study found a huge gap between "Reliable Capability" and "Latent Capability." The models often knew the answer (Latent) but failed to show it consistently (Reliable). In some cases, the models could answer an additional 30% of questions correctly if you just asked them in a different way, but they failed to do so on the standard test. This means the standard "accuracy" scores we see on leaderboards might be hiding a lot of instability. The models aren't necessarily stupid; they are just fragile. They rely on specific word patterns rather than a deep, rock-solid understanding.

The "Flip" and the "Fix"

The researchers also looked at something called "flip rates." This is when a model gets a question right the first time, but then gets it wrong when you rephrase it. Or, even more confusingly, it gets a question wrong the first time, but then gets it right when you rephrase it. This "flip-flopping" behavior shows that the model's brain is wobbly. It's like a compass that points North when you hold it one way, but points East when you tilt it slightly.

However, there was a glimmer of hope. The researchers tried a simple trick called Self-Paraphrasing. Instead of just asking the model one question, they asked it to first rewrite the question in its own words, and then answer based on those new versions. It's like asking a student to "re-read the question in your own head before answering." This simple step helped the models bridge the gap, improving their performance by helping them access that hidden, latent knowledge.

What This Means for the Future

The main takeaway from this paper is that standard accuracy scores can be misleading. Just because a model gets a high score on a test doesn't mean it's truly reliable in the real world. In real life, people ask questions in all sorts of weird ways, and if an AI can't handle that, it's not ready for prime time. The study suggests that we need to stop just looking at "how many questions did it get right?" and start asking "how consistent is it?"

The authors suggest that for AI to be truly trustworthy, it needs to be tested on its ability to stay consistent across different phrasings. Until then, we might be overestimating how smart these robots really are. They might be brilliant at memorizing the script, but they still need to learn how to improvise without losing their place.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →