← Latest papers
💬 NLP

Can LLMs Replace Economic Choice Prediction Labs? The Case of Language-based Persuasion Games

This paper demonstrates that Large Language Models can effectively substitute for human subjects in generating training data for economic choice prediction within language-based persuasion games, often outperforming models trained on actual human data while revealing that capturing history-dependent decision patterns is crucial for accurate prediction.

Original authors: Eilam Shapira, Omer Madmon, Roi Reichart, Moshe Tennenholtz

Published 2026-08-18
📖 6 min read🧠 Deep dive

Original authors: Eilam Shapira, Omer Madmon, Roi Reichart, Moshe Tennenholtz

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the digital economy, every click and every scroll is a choice. When a traveler reads a hotel review on a booking site, they are weighing information against their own desires, trying to decide if a room is worth the price. This is a classic economic puzzle: how do people decide when faced with incomplete information and the possibility of being misled? For decades, economists have studied these moments by asking real humans to play games where one person tries to persuade another. These experiments are difficult to run; they require recruiting participants, building custom software, and waiting for people to make decisions, a process that is slow, expensive, and often limited in size.

Recently, a new tool has emerged that promises to change how we study these decisions: the large language model. These are artificial intelligence systems trained on vast amounts of text, capable of understanding and generating human language with remarkable fluency. Researchers have begun to wonder if these machines can do more than just chat. Could they act as stand-ins for real people in economic experiments? If a computer can mimic human behavior well enough, it might generate the massive amounts of data needed to train better prediction models, solving the problem of scarce human data. The question is not just whether machines can talk like humans, but whether they can think and decide like them in high-stakes situations.

A team of researchers at the Technion in Israel set out to answer this question by placing artificial intelligence agents into a complex game of persuasion. In their setup, one player, acting as an expert, tries to convince a second player, the decision-maker, to visit a specific hotel. The expert knows the true quality of the hotel but can only share a single written review from a list of options. The decision-maker must then choose whether to visit the hotel based on that review and their past experiences with the expert. If the hotel is actually good, the decision-maker wins; if it is bad, they lose. The expert, however, always wins if the decision-maker chooses to visit, regardless of the hotel's true quality. This creates a strategic tension where the expert might lie to get a better outcome, and the decision-maker must learn to spot the deception over time.

The researchers first gathered data from hundreds of real human players who engaged in this game over many rounds. They then replaced the human decision-makers with various large language models, instructing the machines to play the game using the same rules and facing the same hotel reviews. To make the simulations more realistic, the researchers gave each machine a specific personality, such as an optimistic traveler who always sees the bright side, or a price-sensitive shopper who cares deeply about cost. By running thousands of these simulated games, the team created a massive dataset of decisions made by machines, which they used to train a computer program to predict how a human would behave in the same situation.

The results were surprising. When the researchers trained a prediction model using only the data generated by the artificial intelligence players, the model became better at predicting human behavior than a model trained on the actual human data itself, provided the machine dataset was large enough. In fact, the machine-generated data outperformed the human data in almost every scenario tested. This suggests that the artificial intelligence agents were not just mimicking the words of humans, but were capturing the deeper strategic patterns of how people learn and adapt during repeated interactions. The machines learned to weigh the history of past reviews and the trust built over time, rather than just reacting to the immediate words on the screen.

However, the study also revealed a significant flaw in relying solely on machines. While the models trained on artificial data were highly accurate, they were poorly calibrated. In simple terms, this means that when the model said it was confident in a prediction, it was often wrong more often than a model trained on real humans would be. The machine data was so abundant that it taught the model the right answer most of the time, but it failed to teach the model how to gauge its own certainty. To fix this, the researchers found that mixing a small amount of real human data with the vast amount of machine data created the perfect balance. This hybrid approach produced a model that was both highly accurate and trustworthy in its confidence levels.

The researchers dug deeper to understand why the machines were so successful. They discovered that the key to predicting human behavior was not the sentiment of the review itself, but the history of the interaction. When the artificial intelligence agents were allowed to remember many past rounds of the game, they generated data that closely matched human decision-making patterns. But when their memory was restricted to just the last one or two turns, their ability to simulate human behavior collapsed. This finding highlights that human decision-making in these games is fundamentally about strategy and history, not just about reading the emotional tone of a single message. The machines succeeded because they learned to play the long game, just like people do.

To ensure their findings were not limited to just hotel reviews, the team tested their approach in a different, more abstract persuasion game where players could write any message they wanted, rather than choosing from a fixed list. Even in this freer environment, where the rules were less rigid, the data generated by artificial intelligence players proved effective at training models to predict human choices. This suggests that the ability of machines to generate useful training data is a robust phenomenon that extends beyond a single type of experiment.

The study concludes that while artificial intelligence cannot yet perfectly replace the nuance of human behavior, it can serve as a powerful engine for generating the data needed to understand it. By using machines to simulate the strategic complexities of human interaction, researchers can overcome the bottlenecks of traditional economic experiments. The path forward involves a careful partnership: using the scale of artificial intelligence to build broad models of behavior, while grounding those models with just enough real human data to ensure they remain reliable and true to life. This approach offers a new way to study the intricate dance of trust, deception, and decision-making that defines our economic world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →