Can LLMs Be CEOs? Benchmarking Strategic Resource Reallocation with Multi-Role Agent Simulation
This paper introduces \textsc{CEO-Bench}, a multi-agent benchmark that evaluates LLMs on strategic resource reallocation under conflicting stakeholder advice, revealing that while frontier models can produce structurally valid plans, they struggle with strategic calibration due to systematic failures like advisor capture and a tradeoff between integrating diverse perspectives and taking decisive action.
Original paper dedicated to the public domain under CC0 1.0 (http://creativecommons.org/publicdomain/zero/1.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are the captain of a massive, complex ship. You aren't just steering; you have to decide how to split your limited fuel, food, and crew between four different departments: the engine room, the navigation team, the cargo handlers, and the passenger service crew.
This paper asks a simple but tough question: Can an AI (specifically a Large Language Model) act as that captain?
To find out, the researchers built a test called CEO-BENCH. Here is how they did it, explained in everyday terms:
The Setup: A Boardroom of Robots
Instead of asking the AI a simple math problem or a trivia question, they put it in a simulated corporate boardroom.
- The AI is the CEO: It has to make the final decision on how to move money around the company.
- The Advisors are also AI: The AI CEO has to listen to four other "robot executives" (a CFO, CTO, COO, and CMO).
- The Catch: These advisors are programmed to disagree.
- The CFO (money manager) is scared of risk and wants to save cash.
- The CTO (tech boss) wants to spend heavily on new gadgets.
- The COO (operations boss) is worried about breaking things if they change too fast.
- The CMO (marketing boss) sees a huge sales opportunity and wants to spend big now.
Crucially, each advisor only knows a little bit of the truth (information asymmetry). The CEO AI has to figure out who is right, who is wrong, and how to balance their conflicting advice without crashing the company.
The Test: 13 Different Scenarios
The researchers created 13 different "crisis" situations for the AI to solve. Some were easy (everyone agreed), and some were messy (everyone was fighting, and the company was in trouble).
- The Goal: The AI had to produce a plan to move money from one department to another.
- The Rules: The plan had to be legal (follow the rules), bold enough to matter, and consistent with what the company did in previous rounds (no forgetting history).
The Results: Good at Math, Bad at Gut Feelings
The researchers tested five different top-tier AI models. Here is what they found:
- They are great at following rules: Almost all the AIs could create a plan that didn't break the math or the rules. If you asked them to "move 10% of money," they could do the math perfectly.
- They struggle with "Gut Feelings" (Strategic Boldness): This was the hardest part. When the situation demanded a risky, bold move to save the company, the AIs tended to be too scared. They often chose the safe, boring option even when the situation called for a gamble.
- They get "Amnesiac": In a multi-round game, some AIs forgot what happened in the previous round. They would make a decision today that contradicted what they decided yesterday, like a captain who forgets they already burned half the fuel.
- They get "Captured" by one voice: Sometimes, instead of balancing all four advisors, the AI would just blindly follow the loudest voice (usually the one screaming for growth or the one screaming for safety) and ignore the rest.
The Big Trade-Off
The paper discovered a funny contradiction:
- The AIs that were really good at listening to everyone and trying to find a middle ground often ended up making weak, indecisive plans.
- The AIs that made bold, decisive plans often ignored important warnings from the other advisors.
The Verdict
Can an AI be a CEO?
- As a calculator? Yes. It's excellent at checking the numbers and making sure the plan is legal.
- As a leader? Not quite yet. It struggles to weigh conflicting human-like emotions, handle uncertainty, and make the "gut call" when the data is messy. It tends to be too cautious or too easily swayed by one loud voice.
The paper concludes that while AI is getting better at reasoning, the specific skill of "integrating conflicting advice under pressure" is still a human specialty that AI hasn't quite mastered.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.