← Latest papers
🧬 biology

Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory

This paper demonstrates that large language models' ability to distinguish self-generated content from user input (reality monitoring) is highly dependent on conversational memory structure, revealing that performance degrades under episodic delay and that existing benchmarks fail to capture critical dissociations between accuracy and confidence as models take on autonomous, multi-turn roles.

Original authors: Saurabh Ranjan, Konstantina Sokratous, Brian Odegaard

Published 2026-07-28
📖 5 min read🧠 Deep dive

Original authors: Saurabh Ranjan, Konstantina Sokratous, Brian Odegaard

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine you are walking through a giant, bustling library where the books can talk back to you. Sometimes, a book whispers a fact to you; other times, the book starts writing a new story on its own pages. In the world of human psychology, there is a superpower called "reality monitoring." It's the mental muscle that lets you know, "Wait, did I just hear that from a friend, or did I just imagine it?" If this muscle gets weak, people might start believing their own daydreams are real news, or they might forget that a doctor told them something important. This isn't just about remembering facts; it's about knowing where your thoughts came from.

Now, imagine we give this superpower test to a very smart computer program, a Large Language Model (LLM). These are the AI brains that write stories, solve problems, and chat with us. We know they are getting better at talking, but can they tell the difference between something they made up and something you told them? If an AI can't tell the difference, it might accidentally treat its own mistakes as facts you gave it, or it might get confused when you try to correct it. This paper asks a simple but crucial question: Can these AI brains keep track of their own thoughts versus the world around them, especially when they have to remember a long conversation?

The Great "Who Said What?" Test

The researchers set up a game to find out. They gave six different AI models a series of word pairs. Sometimes, the AI was shown a complete pair (like "Cat" and "Mouse") and just had to repeat the second word. This was like the AI "hearing" something from the outside world. Other times, the AI was shown only the first word ("Cat") and had to come up with the second word itself. This was the AI "imagining" something. Afterward, the AI had to play detective: "Did I just hear this word, or did I just make it up?"

The scientists wanted to see if the AI could do this when the conversation was short and sweet, and if it could still do it when the conversation was long and messy, like a real chat with a friend that goes on for hours.

The Surprise: It's All About the Memory Trick

Here is where things get wild. When the AI was tested in a quick, one-shot conversation (like a text message), it was a genius at knowing what it made up. It got 100% of its own "imagined" words right! It was like a kid who just drew a picture and immediately knew, "I drew that!"

But the moment the researchers made the AI wait and remember a long list of words from earlier in the conversation (a "delayed" test), the magic trick fell apart. Suddenly, the AI got confused. It started doing the opposite of what it did before: it became better at remembering what you told it, but terrible at remembering what it made up. It was as if the AI forgot it was the artist and started thinking it was just a mirror reflecting everything it saw.

The researchers found that this wasn't just about how "big" the AI's brain was. A massive AI with 70 billion parameters didn't necessarily do better than a smaller one. Instead, it seemed to depend on how the AI's internal "expert" teams were organized. It's not about having a bigger library; it's about how the librarians are arranged.

The Feedback Trap: When Correction Makes It Worse

The scientists also tried a new trick: they told the AI when it was right or wrong after every guess. You'd think this would help, right? Like a teacher correcting a student.

Surprisingly, for some of the AIs, this feedback made things worse. It didn't just fix their mistakes; it scrambled their confidence. Some models started getting the answers right but lost the ability to know that they were right. They became like a person who is 100% sure they are correct, even when they are completely wrong. In fact, for one specific model, the feedback actually flipped its brain inside out, making it confidently say the opposite of the truth.

The study suggests that as we start using these AIs for serious, long-term jobs—like helping doctors or lawyers—we can't just ask, "Do you know the answer?" We have to ask, "Do you know where you got that answer?" Because if an AI can't tell the difference between its own hallucinations and your real facts, it might confidently lead us down the wrong path. The paper suggests that simply making AI bigger isn't the solution; we need to figure out how to help them keep a clear diary of their own thoughts, especially when the conversation gets long and complicated.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →