← Latest papers
🤖 machine learning

Benchmarking In-context Experiential Learning Through Repeated Product Recommendations

This paper introduces BELA, a benchmark for experiential learning in repeated product recommendation scenarios using real-world products and synthetic customer personas, which reveals that while current LLM-based agents can adapt within individual interactions, they struggle to improve their strategies across multiple episodes by inferring shared latent structures.

Original authors: Gilbert Yang, Yaqin Chen, Thomson Yen, Hongseok Namkoong

Published 2026-08-11
📖 5 min read🧠 Deep dive

Original authors: Gilbert Yang, Yaqin Chen, Thomson Yen, Hongseok Namkoong

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery, but the clues are hidden inside a conversation. In the world of artificial intelligence, there is a big difference between solving a single puzzle and learning how to solve puzzles better over time. Most AI tests today are like giving a student a single math problem with all the numbers already written down; the goal is just to get the right answer right now. But real life is messier. It's more like a detective who talks to a witness, asks a few questions, gets an answer, and then has to use that experience to interview the next witness more effectively. This paper dives into a specific corner of AI research called "experiential learning." It asks a simple but tricky question: Can an AI get smarter at figuring out what people want just by talking to them over and over again? The researchers are testing if these digital agents can learn from their past conversations to become better detectives in the future, rather than just being good at solving one isolated case.

The authors of this paper, a team from Columbia Business School and Sun Yat-sen University, decided to test this idea using a scenario we all know: shopping for products. They built a giant, virtual playground called BELA (Benchmark for Experiential Learning and Active exploration). Think of BELA as a massive, infinite mall where the AI plays the role of a salesperson. In this mall, the AI has to talk to thousands of different "customers" (who are actually computer simulations of people with specific tastes) to figure out which product they should buy. The twist is that the AI doesn't just get one shot. It talks to a customer, makes a recommendation, gets feedback, and then moves on to a new customer. The big test is whether the AI remembers what it learned from the first customer and uses that knowledge to ask better questions of the second, third, and tenth customer.

To make this test fair and huge, the researchers didn't just make up a few fake people. They created a system with 71,000 real products from Amazon and 2,000 different groups of products (like a group of hair gels or a group of rice dishes). They then paired these with 1 million different customer personalities (simulated by other AIs) who have very specific, consistent likes and dislikes. This creates a mind-boggling 2 billion possible combinations of "customer + product group." It's like having a library of every possible shopping trip you could ever imagine, allowing the researchers to see if the AI can spot patterns across the crowd.

The researchers ran their experiments with the smartest AI models available today, including giants like GPT-5.4, Gemini-3.1, and Claude-Opus. They watched these models try to sell products over a series of 10 shopping trips (called episodes). They wanted to see two things: First, could the AI learn during a single conversation (asking the right follow-up questions)? Second, could the AI learn across conversations (getting better at the job as it gained more experience)?

Here is the surprising result: The AI models were actually pretty good at the first part. When talking to a single customer, they could learn from the conversation and ask better questions as the chat went on. However, they completely failed at the second part. After talking to nine different customers, the AI was no better at guessing what the tenth customer wanted than it was with the first one. It was like a salesperson who is great at figuring out what one person likes, but every time a new person walks in, they start from scratch, forgetting everything they learned from the previous nine people.

The paper suggests that these current "frontier" models are missing a crucial skill: they can't look back at their past experiences and say, "Oh, I see a pattern here; I should ask this type of question next time." Instead, they seem to treat every new customer as a completely fresh mystery, even if they just met someone with very similar tastes five minutes ago. The researchers also found that the AI didn't get better at guessing how confident it should be; it kept asking the same number of questions, even when it should have learned to ask fewer.

To prove that the task wasn't just too hard, the researchers created an "ideal path." They manually guided the AI through the best possible questions and showed that if the AI did use its past experiences correctly, it could have drastically improved its recommendations and asked fewer questions. This proves that the information was there to be learned, but the current models just weren't picking it up.

In short, this paper suggests that while our current AI is getting very good at having a single, smart conversation, it hasn't yet figured out how to be a "learning machine" that gets smarter over its entire career. It's a bit like a student who can ace a test if they study the night before, but forgets everything the moment the bell rings, unable to use what they learned on Monday to help them on Tuesday. The authors conclude that for AI to truly handle the messy, shifting real world, it needs to develop this ability to learn from experience across multiple episodes, a skill that today's most advanced models still struggle to master.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →