From Recall to Forgetting: Benchmarking Long-Term Memory for Personalized Agents
This paper introduces Memora, a long-term memory benchmark spanning weeks to months that evaluates personalized agents on remembering, reasoning, and recommending tasks using a novel Forgetting-Aware Memory Accuracy (FAMA) metric, revealing that current models frequently rely on obsolete information and struggle to reconcile evolving memories despite marginal improvements from dedicated memory agents.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a digital assistant, like a super-smart personal secretary. You've been working with this secretary for months. You've told them your favorite movies, your grocery budget, your fitness goals, and even that you used to hate a certain type of music but now love it.
Ideally, this secretary should remember all of that, update their notes when you change your mind, and forget the stuff you no longer care about. But in reality, most current AI assistants are like goldfish with a very short attention span. They remember what you said five minutes ago, but if you ask them about something you discussed three weeks ago, they often act like they've never met you. Or worse, they remember the old you, not the current you.
This paper, titled "From Recall to Forgetting," introduces a new way to test these assistants and a new metric to see if they are actually getting smarter or just getting confused.
Here is the breakdown in simple terms:
1. The Problem: The "Goldfish" vs. The "Human"
Current AI benchmarks (tests) are like asking a student to memorize a single fact from a textbook and recite it back.
- The Old Way: "I told you yesterday I like pizza. Do you remember?" -> AI: "Yes."
- The Real World: You tell the AI you like pizza. Two weeks later, you tell them you're on a diet and hate pizza. A month later, you tell them you're celebrating and want pizza again.
- The Failure: Most AI assistants get confused. They might still recommend pizza because they forgot the diet, or they might say you hate pizza because they forgot the celebration. They struggle to update their memory and forget the old stuff.
2. The Solution: Introducing "Memora"
The authors created a new test called Memora. Think of Memora not as a pop quiz, but as a long-term relationship simulator.
Instead of a quick chat, Memora simulates interactions over weeks, months, and even quarters. It creates a "life" for a digital persona (like a busy software engineer or a creative designer) and tracks how their life changes.
The test checks the AI on three specific skills, using a "Memory Grounding" approach:
- Remembering: Can you recall that I need to buy milk? (Basic recall)
- Reasoning: Can you look at my grocery receipts from the last month and tell me if I'm over my budget? (Connecting the dots)
- Recommending: Can you suggest a movie that fits my current taste, knowing I used to love horror movies but now prefer comedies? (Adapting to change)
3. The New Scorecard: "FAMA" (The "Forgetting" Test)
This is the most creative part of the paper. Previous tests only checked: "Did the AI get the answer right?"
But what if the AI got the right answer for the wrong reason? What if it recommended a horror movie because it remembered your old taste, even though you said you hate them now?
The authors introduced FAMA (Forgetting-Aware Memory Accuracy).
- The Analogy: Imagine a detective solving a case.
- Old Score: Did the detective catch the criminal? (Yes/No)
- FAMA Score: Did the detective catch the criminal using the right clues?
- If the detective catches the criminal but uses a clue that was proven false yesterday, FAMA gives them a penalty. It rewards the AI for forgetting outdated info and remembering the new truth.
4. What They Found (The Plot Twist)
The researchers tested 4 big AI models and 6 different "memory agents" (specialized tools designed to help AI remember). The results were surprising:
- The Longer the Time, the Worse the Memory: As the conversation stretched from a week to a quarter (3 months), the AI's performance tanked. It's like trying to remember a story told over 3 months; the details get fuzzy.
- The "Forgetting" Failure: The biggest issue wasn't that the AI couldn't remember; it was that it couldn't forget. It kept holding onto old preferences and outdated facts, leading to bad recommendations.
- Specialized Agents vs. Raw AI: The specialized "memory agents" were better at remembering facts (like a phone number), but they were still terrible at reasoning (connecting the dots) and adapting to changes.
- The "Reasoning" Gap: Even the smartest AI models struggled to do math or logic based on old conversations. They could tell you "I spent $50," but they couldn't figure out "I have $50 left in my budget" if that info was buried in a chat from two weeks ago.
5. The Takeaway
The paper concludes that building a truly "personal" AI isn't just about giving it a bigger brain or a longer memory bank. It's about teaching it how to manage a library.
A good AI needs to:
- Shelve new information correctly.
- Cross-reference old information with new information.
- Throw away (forget) the old, invalid books so they don't clutter the shelf.
Currently, our AI assistants are great at reading the book in front of them, but they are terrible at managing the library of their entire life with you. Memora is the first rigorous test to see if they can finally learn to be good librarians.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.