← Latest papers
💬 NLP

Simulating Human Memory with Language Models

This paper demonstrates that while out-of-the-box language models possess unrealistically reliable memory compared to humans, applying specific prompting strategies and memory compactors can induce human-like forgetting, thereby improving their effectiveness as user simulators in educational contexts.

Original authors: Qihan Wang, Nicholas Tomlin, Michael Hu, Brian Dillon, Tal Linzen

Published 2026-05-26
📖 5 min read🧠 Deep dive

Original authors: Qihan Wang, Nicholas Tomlin, Michael Hu, Brian Dillon, Tal Linzen

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot how to be a student so you can test your teaching methods on it before trying them on real people. You'd want the robot to forget things just like a human does, right? If the robot remembers everything perfectly, it's not a very good test subject for a human teacher.

This paper is about a team of researchers who tried to build a "student robot" that forgets things the way humans do. They discovered that standard AI models are actually too good at remembering things, and they had to build special "forgetting machines" to fix that.

Here is the story of their findings, broken down into simple parts:

1. The Problem: The Robot with a Perfect Memory

The researchers set up a series of memory tests for both humans and AI models. Think of these tests like:

  • The Number Game: Remembering a list of numbers and repeating them backward.
  • The Story Game: Reading a short story and retelling it word-for-word.
  • The Map Game: Looking at a map of a city and then answering questions about how to get from point A to point B without looking at the map anymore.

The Result: The humans did okay. They remembered about 7 numbers, got a few details of the story, and knew the map pretty well. But the AI models? They were superheroes. They remembered 20 numbers perfectly, recited the entire story without a single typo, and navigated the map flawlessly.

Even when the researchers told the AI, "Hey, act like a human who is tired and forgetful," the AI just ignored the advice and kept remembering everything perfectly. It's like telling a camera to take a blurry photo; the camera just keeps taking sharp, high-definition pictures because that's what it's built to do.

2. The Solution: The "Four-Box" Backpack

The researchers realized that simply asking the AI to forget wasn't working. They needed to change the AI's "backpack."

In human psychology, there's a famous idea that our short-term memory is like a backpack with only four compartments (or "chunks"). If you try to put more than four things in there, the oldest ones fall out.

So, the researchers built a special tool called a COMPACTOR.

  • How it works: Before the AI tries to answer a question, it is forced to put all the information it just read into a "backpack" with only four slots.
  • The Catch: It has to summarize the information to fit. If the story is long, it has to squeeze the details into four tiny boxes. Anything that doesn't fit gets thrown away.
  • The Result: Suddenly, the AI started forgetting things! It made mistakes. It forgot parts of the map. It mixed up numbers. It finally started acting like a human student who is struggling to keep up.

3. The "Fake" Human vs. The "Real" Human

The researchers tried a few other tricks to make the AI act human:

  • The "Acting" Trick: They gave the AI examples of humans making mistakes and said, "Copy this."
    • Result: It didn't work well. The AI could copy the style of a mistake, but it didn't actually forget the information. It was like an actor pretending to trip but still walking perfectly.
  • The "Summary" Trick: They asked the AI to just write a summary of the text and answer from that.
    • Result: This helped a little, but the COMPACTOR (the four-box backpack) was the clear winner. It forced the AI to actually lose information, which is what real humans do.

4. Why Does This Matter? (The "Teaching" Test)

To see if this new "forgetting AI" was actually useful, the researchers ran a final test. They pretended the AI was a student in a class.

They showed the AI (and real humans) different versions of a biography:

  • Version A: Written in simple, easy language.
  • Version B: Written in very complex, hard language.
  • Version C: Full of repeated facts (redundant).
  • Version D: Full of distracting, useless information.

They asked: "Which version of the story will a human student remember best?"

  • The Standard AI (Perfect Memory): Said, "Oh, the human will remember all of them perfectly! It doesn't matter which one we use." (This is wrong because humans do struggle with complex or distracting text).
  • The COMPACTOR AI (Human-like Memory): Said, "The human will struggle with the complex one and the distracting one, but they'll do okay with the simple or repeated one."

The COMPACTOR AI was much better at predicting what a real human could handle. It realized that if you give a human too much hard stuff, they will forget it.

The Bottom Line

The paper concludes that if you want to use AI to simulate humans (for example, to test new teaching methods or design better user interfaces), you can't just use a standard AI. Standard AIs are too perfect.

You have to build artificial limitations into them—like a backpack with only four slots—to make them forget things the way we do. Only then can they be useful stand-ins for real people.

What the paper doesn't say:

  • It doesn't say this will cure memory loss in humans.
  • It doesn't say these robots will replace human teachers.
  • It doesn't claim the AI is now "perfectly" human; even the best "forgetting" AI still makes different kinds of mistakes than humans do. It's just a lot closer than before.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →