← Latest papers
💬 NLP

Epistemic Familiarity is Associated With Belief Stability in Large Language Models

This paper introduces the P-StaT framework to demonstrate that large language models exhibit greater belief stability when processing familiar fictional statements compared to unfamiliar synthetic ones, revealing that epistemic familiarity is a key factor in mitigating semantic-induced belief instability.

Original authors: Samantha Dies, Courtney Maynard, Germans Savcisens, Tina Eliassi-Rad

Published 2026-07-22
📖 5 min read🧠 Deep dive

Original authors: Samantha Dies, Courtney Maynard, Germans Savcisens, Tina Eliassi-Rad

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a super-smart robot how to tell the difference between a fact, a lie, and a made-up story. You might think that once the robot learns the rules of "truth," it would stick to them no matter how you ask the question. But in the world of Large Language Models (LLMs)—the powerful AI brains behind chatbots and search tools—things are a bit more wobbly. These models are like giant libraries that have read almost everything on the internet, but they don't "know" things the way humans do; they predict what words should come next based on patterns. A key question for scientists is: if you slightly change the context or the way a question is framed, does the robot's belief in a fact stay steady, or does it wobble and collapse? This is the study of "epistemic stability." If a robot believes a city exists today, but changes its mind just because you mentioned a fictional story nearby, that's a problem. We care about this because these models are increasingly used to give us information, and we need to know if their beliefs are rock-solid or if they are easily shaken by a little bit of confusion.

Enter a new study by Samantha Dies and her team, who decided to shake the robot's world to see what would fall out. They built a framework called P-StaT (which stands for Perturbation Stability of Truth). Think of P-StaT as a "belief stress test." The researchers took 21 different AI models and gave them a bunch of statements to judge. Some statements were real facts (like "Surat is in India"), some were obvious lies, and the tricky part was the "Neither" statements—claims that aren't true or false in the real world.

The team split these "Neither" statements into two camps. The first camp was Familiar Fictional statements. These are made-up things that sound like they belong in a storybook or a movie, like "The city of Bikini Bottom is in the Pacific Ocean." Even though Bikini Bottom isn't real, the AI has probably seen it in its training data because it's a popular cartoon. The second camp was Unfamiliar Synthetic statements. These are completely new, weird combinations of words the AI has likely never seen before, like "The city of Eustapor is in Oklanian." These are designed to be total strangers to the model.

The researchers then played a game of "what if." They told the AI, "Okay, pretend these made-up things are actually true," and watched to see if the AI would suddenly start doubting its real facts. It's like telling a person, "Imagine that dragons are real," and then asking, "So, is the sky blue?" If the person suddenly says, "I'm not sure, maybe the sky is green because of dragons," their belief system is unstable.

The results were fascinating. When the AI was exposed to the Unfamiliar Synthetic statements (the totally new, weird ones), it got very shaky. In the behavioral tests (where the AI was just chatting back), the models retracted their beliefs about real facts more than 50% of the time. That's a coin flip! If you flipped a coin 100 times, you'd expect about 50 heads; here, the AI lost its confidence in real facts more often than not just because it was confused by something totally new.

However, when the AI was exposed to the Familiar Fictional statements (the cartoon cities and movie monsters), it stayed much calmer. It knew that Bikini Bottom was a cartoon, so it didn't let that idea mess up its knowledge about real cities. The study suggests that the AI's instability isn't just about the words being weird; it's about how unfamiliar the context feels. The models seem to have a "comfort zone" of things they've seen before, and when you push them into the unknown, their grip on the truth slips.

The researchers also looked at what kind of facts the AI was most likely to drop when it got confused. They found that the AI didn't just randomly forget things. It tended to drop beliefs about things that were already a bit fuzzy to begin with—like technical medical terms, obscure geography, or statements that had a little bit of ambiguity. It's as if the AI's belief system is a house of cards: when you blow a gentle breeze of familiar fiction, the house stands. But when you blow a hurricane of unfamiliar nonsense, the cards that were already slightly wobbly (the tricky, technical ones) are the first to fall.

In short, this paper suggests that epistemic familiarity—how well an AI knows a topic—is a huge factor in how stable its beliefs are. If an AI is asked to reason about things it has never seen before, it becomes much more likely to hallucinate or change its mind about things it actually knows. The authors aren't saying the AI is broken, but they are showing us that "stability" is a new way to measure how reliable an AI is, alongside the usual "accuracy" tests. It turns out, for these digital brains, knowing the rules of the game is just as important as knowing the facts of the game.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →