Evolution of cooperation with Q-learning: how much information do we need?
By applying Q-learning to structured populations, this study reveals that cooperation levels follow an inverted U-shaped relationship with information availability, demonstrating that a moderate amount of information—rather than more—is optimal for fostering cooperation by balancing decision-making sufficiency with tractability.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the bustling world of nature and human society, cooperation is everywhere. Bees divide labor to keep their hives alive, and vampire bats share blood meals to help neighbors survive hunger. Yet, this willingness to help others at a personal cost seems to contradict the basic logic of survival, which suggests that individuals should act only in their own self-interest. For decades, scientists have tried to understand how cooperation can take root and thrive when selfishness appears to be the easier path. A central question in this field is whether having more information about one's surroundings leads to better decisions and, ultimately, more cooperation. Intuitively, it seems logical that knowing more about the people around you would help you make wiser choices. However, a new study suggests that this intuition might be wrong, proposing instead that there is a specific, limited amount of information that is actually best for fostering teamwork.
Researchers from Shaanxi Normal University and Ningxia University set out to test this idea using a computer model based on a classic scenario known as the prisoner's dilemma. In this scenario, two people must decide whether to cooperate or betray each other. If both cooperate, they both do well; if one betrays while the other cooperates, the betrayer wins big while the cooperator loses; if both betray, both lose out. The challenge is that betraying is always the safer individual choice, even though cooperating is better for the group. To see how people might learn to cooperate, the scientists used a method called Q-learning. This is a type of artificial intelligence where a computer agent learns by trial and error, trying different actions and remembering which ones lead to the best rewards over time. This approach mimics how humans and many animals learn from their own experiences rather than simply copying others.
The team placed these learning agents on a grid, where each agent could only see and interact with a certain number of neighbors. They treated the number of neighbors an agent could see as a measure of how much information that agent had. By changing this number from just one neighbor up to twelve, they could test how different levels of information affected the group's ability to cooperate. They ran these simulations on a simple square grid and also on a more complex, web-like network that resembles real-world social connections.
The results revealed a surprising pattern. When the agents had very little information, seeing only one or two neighbors, they mostly chose to betray each other. As the number of neighbors increased to three or four, the level of cooperation rose to its highest point. However, if the agents were given even more information, seeing five or more neighbors, cooperation began to drop again. The most cooperative groups were not the ones with the most information, but the ones with a moderate amount. This "inverted U" shape held true for both the simple grid and the complex web-like network. In the complex network, the researchers found that the most connected individuals, who had the most information, actually became the least cooperative, often acting as persistent defectors, while those with fewer connections maintained higher levels of cooperation.
To understand why this happened, the scientists looked closely at how the agents made their decisions. When information was scarce, the agents were too rigid; they quickly learned that betraying was the only safe option and stuck with it. When information was overwhelming, the agents became confused. The sheer volume of data made it difficult for them to form a clear strategy, causing them to switch between cooperating and betraying chaotically. They could not settle on a reliable plan. It was only with a moderate amount of information that the agents found a sweet spot. They had enough data to recognize when cooperation was a good idea, but not so much that they became paralyzed by choice. This balance allowed them to maintain stable strategies long enough for cooperation to take hold, while still being flexible enough to adapt when the situation changed.
The study also explored how the speed of learning and the importance of future rewards influenced these outcomes. They found that if the agents cared too much about immediate rewards and ignored the future, cooperation vanished entirely. However, the core finding remained robust: having more information did not guarantee better results. In fact, the simulations showed that an optimal amount of information exists, and exceeding it can actually harm the group's ability to work together. This challenges the common belief that more data always leads to better decisions. Instead, it suggests that in a world of complex interactions, knowing just enough is often the key to making the right choice. The researchers propose that this balance between having enough information to make a decision and having enough mental space to process it is crucial for the emergence of cooperation, offering a new perspective on how we might design systems or environments to encourage teamwork.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.