← Latest papers
🤖 AI

RECON: Benchmarking Agent Memory for Compositional Reasoning over Long Contexts

This paper introduces RECON, a novel benchmark comprising 24 long-context case files across criminal, medical, and financial domains, designed to evaluate the limitations of current LLM-based agents in performing complex compositional reasoning tasks such as multi-hop evidence reconstruction, conflict resolution, and counterfactual analysis, where even the strongest systems achieve only 22.4% accuracy.

Original authors: Mihir Shriniwas Arya

Published 2026-07-21
📖 3 min read☕ Coffee break read

Original authors: Mihir Shriniwas Arya

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to solve a mystery, but the clues aren't just scattered around a room; they are buried inside a library the size of a small city. This is the world of Large Language Models (LLMs), the super-smart computer brains behind chatbots and digital assistants. These models are amazing at reading and talking, but they have a tricky weakness: memory. Think of memory not just as a hard drive where facts are stored, but as a detective's ability to remember not just what happened, but how one event led to another. If a witness changes their story, or a lab test is proven fake, a good detective knows which old conclusions are now broken and which ones still stand. The big question researchers are asking is: Can these AI detectives actually track these complex, shifting stories over thousands of pages, or do they just get confused and make things up?

Enter RECON, a new "stress test" designed to see if AI agents can handle long, complicated stories where facts change, contradict each other, and ripple through the narrative like a wave. The researchers built 24 massive case files—some as long as 100,000 words—covering criminal investigations, medical mysteries, and financial fraud. These aren't just boring lists of data; they are dynamic stories where a clue found on Day 1 might be proven fake on Day 5, forcing the AI to figure out which of its earlier guesses were wrong. The test checks six specific skills, like connecting a chain of 15 clues, figuring out what would have happened if a key event occurred at a different time, or deciding which witness is lying when two people tell different stories.

The results of this experiment are a bit of a reality check for the AI world. Even the smartest AI systems tested struggled mightily. The best non-specialized AI got the right answer less than 23% of the time. To put that in perspective, if you asked a human to solve these same puzzles with the full text in front of them, they would do much better. The study found that the problem isn't just that the AI can't find the right page in the book (retrieval); the real bottleneck is reasoning. Even when the AI was given the "cheat sheet" (a perfect, structured map of all the facts and how they connect), it still only got about 55% of the answers right. This suggests that while AI is getting better at reading long books, it still hasn't mastered the art of connecting the dots when the story changes underneath it. The researchers conclude that for AI to truly act like a reliable agent in real-world jobs, it needs to get much better at understanding how facts depend on one another, rather than just remembering them.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →