High Volatility and Action Bias Distinguish LLMs from Humans in Group Coordination
This paper reveals that large language models struggle to achieve stable group coordination compared to humans, exhibiting excessive action switching and limited responsiveness to richer feedback, which highlights fundamental behavioral differences in adaptive learning and convergence strategies.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are part of a team playing a high-stakes guessing game. The goal is simple: everyone on the team must guess a number, and when you add all those numbers together, the total must match a secret "mystery number" chosen by the referee.
Here's the catch: You cannot talk to each other. You don't know what your teammates are guessing. You only get one piece of information after every round: "Too high," "Too low," or "Just right" (and sometimes, how much too high or too low).
This is the game the researchers used to test humans against Large Language Models (LLMs) like the ones powering this conversation. They wanted to see: Can AI teams coordinate as well as human teams?
The short answer? No. Not even close.
Here is the breakdown of what happened, using some everyday analogies.
1. The Human Team: The Calm Orchestra
When humans played this game, they acted like a seasoned jazz band.
- They listened: If the group was "too high," they didn't all panic and jump down 50 points. They listened to the feedback and made small, calculated adjustments.
- They found roles: As the game went on, some people naturally stopped changing their numbers. They became the "anchors," holding steady while others did the heavy lifting of adjusting. This reduced the noise and helped the group find the target quickly.
- They learned: If they played the game 10 times in a row, they got better every time. They figured out the rhythm.
2. The AI Team: The Hyperactive Hamster Wheel
When the researchers swapped humans for AI models (like DeepSeek, Llama, and Gemini), the vibe changed completely. The AI teams acted less like a jazz band and more like a room full of hamsters on a wheel that's spinning too fast.
Here are the three main ways the AI teams failed to coordinate:
A. The "Action Bias" (The Hamster Effect)
Humans have a natural tendency to stay still if things are going well. If the group sum is close to the target, a human might think, "I'll just keep my number the same and let my teammates do the work."
The AI, however, has a massive Action Bias. It feels like it must do something every single turn. Even when the group is close to the target, the AI keeps changing its number.
- Analogy: Imagine trying to balance a broom on your hand. If you make tiny, steady movements, you win. If you keep jerking your hand wildly left and right because you feel like you "need to move," the broom falls. The AI kept jerking the hand.
B. The "Overreaction" (The Pendulum Swing)
When the AI got feedback like "You are 10 points too high," it didn't just subtract 10. It often subtracted 20 or 30.
- The Result: The group sum would swing wildly. One round they were way too high, the next round they were way too low, and the next round way too high again. They were stuck in an endless loop of swinging back and forth, never landing on the target.
- Human vs. AI: Humans tend to be a bit cautious (under-reacting), which actually helps stabilize the group. AI tends to be aggressive (over-reacting), which creates chaos.
C. The "Amnesia" (No Learning)
This was the most surprising part. Humans got better the more they played. They learned, "Okay, when the group is this big, we need to be more careful."
The AI models? They didn't learn. Even after playing 10 games in a row, they performed just as poorly in Game 10 as they did in Game 1. They treated every new game as if it were the very first time they had ever seen the rules. They couldn't connect the dots from one experience to the next.
3. The "Richer Feedback" Surprise
The researchers tried giving the AI better information. Instead of just saying "Too high," they said, "Too high by 25."
- For Humans: This was like getting a map. They used the extra detail to zoom in on the target and solved the game much faster.
- For AI: It was like giving a map to someone who doesn't know how to read. The extra numbers didn't help them at all. They still swung wildly and failed to coordinate.
The Big Picture: Why Does This Matter?
We are starting to build systems where multiple AIs work together to solve complex problems (like writing code, managing supply chains, or running simulations).
This paper warns us: Just because AI is smart at answering questions doesn't mean it's good at working in a team.
- Humans are great at "group coordination" because we know when to speak up and when to stay quiet. We adapt our behavior based on the group's mood.
- Current AI is like a very eager, very loud intern who never stops talking, over-corrects every mistake, and forgets what happened yesterday.
The Takeaway
If you want to build a team of AIs that works like a human team, you can't just plug them in and hope for the best. You have to teach them restraint. You have to teach them that sometimes, the best move is to do nothing and let the team stabilize. Until AI learns the art of "strategic inaction," it will struggle to coordinate with us or with itself.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.