ForgetBench: Benchmarking Forgetting Dynamics of Long-Term Parametric Memory in Language Models
This paper introduces ForgetBench, a novel benchmark designed to systematically evaluate the long-term forgetting dynamics and retention stability of large language models under continual knowledge editing, revealing that current methods struggle to balance effective knowledge updates with the preservation of previously acquired information.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a world where your brain could learn a new fact every second, but every time you learned something new, you accidentally erased an old memory. This isn't just a sci-fi nightmare; it's a growing concern for the artificial intelligences we call Large Language Models (LLMs). These digital brains are incredibly good at reading books and having conversations, but they struggle with something humans do naturally: keeping old memories safe while making room for new ones. Scientists have long studied how human memory works, splitting it into "working memory" (what you're thinking about right now) and "long-term memory" (facts you've stored away for years). While we know how to teach these AI models new tricks, we don't fully understand how they handle the messy business of updating their internal knowledge over time without losing everything they knew yesterday. If we want AI to be a helpful lifelong companion rather than a forgetful friend who needs constant reminders, we need to figure out how to make their memories stick.
Enter ForgetBench, a new "stress test" designed by researchers to see exactly how and why AI models forget things when they are constantly being updated. Think of ForgetBench not as a simple quiz, but as a time-lapse camera for an AI's brain. Instead of just checking if an AI knows a fact right now, the researchers set up a scenario where the AI has to learn a stream of new information, one after another, like a student taking a new history test every day for a month. The big question they asked was: "If we teach this AI a new fact today, will it still remember the fact we taught it a week ago, or will the new lesson have overwritten the old one?"
To answer this, the team built two different kinds of memory challenges. The first, called Concept-based QA, is like a game of flashcards with isolated facts. Imagine you tell the AI, "Donald Rodriguez is 22 years old," and then later, "Actually, Donald is 79." The researchers wanted to see if the AI could swap the age without getting confused or forgetting other people's ages. The second challenge, Scenario-based QA, is much more like a complex role-playing game. Here, the AI has to track relationships between characters, like "Who followed Bill Gates and retweeted a post about space exploration?" This tests if the AI can hold onto a web of connected stories or if the new updates cause the whole story to collapse.
The researchers ran these tests on several different AI models using various methods to update their knowledge. What they found was a bit of a bummer for the current state of AI: existing methods are terrible at balancing the two. When the models learned new facts, they often forgot the old ones, or they became so confused that they started making up nonsense. The study measured this using something called a "forgetting curve," which tracks how quickly a memory fades as more new information is added. They found that for most current models, the "half-life" of a memory (the time it takes for the AI to forget half of what it learned) is very short. In other words, the AI is great at learning now, but it's terrible at remembering later.
The paper suggests that while we have gotten really good at teaching AI new things in a single session, we haven't solved the problem of how to keep those memories safe over a long period of constant updates. The researchers didn't just find a problem; they provided a new way to measure it, showing that the current tools for updating AI knowledge are like trying to write a new chapter in a book by erasing the previous pages. Until we can build better "memory mechanisms" that allow AI to grow without losing its past, these digital brains will remain brilliant but incredibly forgetful.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.