← Latest papers
🤖 AI

AI Agents for Inventory Control: Human-LLM-OR Complementarity

This contribution introduces InventoryBench, a comprehensive benchmark demonstrating that the combination of operations-research algorithms, large language models, and human decision-makers creates a complementary, high-performance system for inventory control that surpasses any single method acting alone.

Original authors: Jackie Baek, Yaopeng Fu, Will Ma, Tianyi Peng

Published 2026-05-06
📖 4 min read☕ Coffee break read

Original authors: Jackie Baek, Yaopeng Fu, Will Ma, Tianyi Peng

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you run a lemonade stand. Your goal is simple: buy enough lemons to sell lemonade all day without running out, but also without buying so many that they spoil in the heat. This is the classic problem of "inventory control."

For decades, economists have used Operations Research (OR) to solve this. Think of OR as a stubborn, hyper-astute calculator. It looks at your past sales data and says, "Based on the last 30 days, you need exactly 50 lemons." This works great when things are predictable, but when a heatwave hits or a sudden rainstorm drives the crowd away, the calculator gets confused because it knows only mathematics, not the weather.

Recently, we introduced Large Language Models (LLMs) – the AI behind chatbots. Think of an LLM as a knowledgeable but sometimes scatterbrained friend. This friend knows that "lemonade sells better in July" or "a heatwave is coming," but they are not good at performing the precise mathematics to figure out exactly how many lemons to buy. They might guess 100 lemons when you only need 60, or they might get distracted by random noise in the data and think demand is changing when it is not.

This work asks a simple question: What happens when we bring the calculator, the friend, and a human manager together?

The Experiment: A New Benchmark

The researchers built a massive test environment called InventoryBench. It is like a video game with over 1,000 different scenarios. Some scenarios are invented (synthetic), while others are based on real sales data from H&M clothing. They tested three types of "players":

  1. The Calculator (OR): Mathematics only.
  2. The Friend (LLM): AI only.
  3. The Hybrid: The calculator makes a suggestion, and the friend decides whether to follow it or change it.

The Result: The hybrid team won.
When the calculator made a suggestion and the friend reviewed it (or vice versa), they achieved higher profits than either could alone. The friend could recognize that "bath season" had begun (something the calculator missed), and the calculator could perform the precise mathematics to ensure they did not over-order (something the friend struggled with). They were complements, not competitors.

The Human Element: The "Co-Pilot" Test

The researchers then added a real human into the mix. They conducted a classroom experiment with 69 students playing the lemonade game. They tried three different working modes:

  • Mode A (The old way): The calculator sets a number, and the human makes the final decision.
  • Mode B (The Co-Pilot): The calculator sets a number, the AI (friend) explains why it agrees or disagrees, and then the human makes the final decision based on both.
  • Mode C (The Autopilot): The AI makes the decisions, and the human only gives occasional high-level advice.

The Winner: Mode B (The Co-Pilot).
Teams where the human read the AI's reasoning and then made the final decision performed best. They outperformed the calculator alone and they outperformed the AI alone.

Why?
The work found that humans and AI are like two different types of detectives.

  • The AI is excellent at recognizing patterns in the data (like "demand is rising") and using general knowledge (like "it is summer, so swimwear sells").
  • The Human is excellent at spotting the AI's errors. For example, if the AI assumes a lemon delivery is on the way, but the human notices the delivery truck never arrived (a "lost order"), the human can override the AI to order more.

The "Magic" of Teamwork

A common fear is that AI only makes weak humans look better and strong humans look worse, or that humans will simply ignore the AI. The researchers proved this is not the case.

They used a special mathematical trick to show that at least 30% to 60% of the people in their experiment actually mastered the game better because they worked with the AI. It was not just that "weak" players followed the AI; even "strong" players improved because the AI caught errors they had overlooked, and the AI improved because humans caught errors it had overlooked.

The Conclusion

The work concludes that the best way to manage inventory is not to replace humans with AI or to ignore AI and stick with old mathematics. The best system is a three-way partnership:

  1. The Calculator (OR) provides solid mathematics and structure.
  2. The AI (LLM) brings world knowledge and context (such as seasons or trends).
  3. The Human acts as the final judge, catching errors and adding common sense.

When these three work together, they create a system that is smarter and more profitable than any of them could be alone.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →