Experimental Designs for Multi-Item Multi-Period Inventory Control
This paper investigates the systematic biases inherent in standard A/B testing designs for multi-item, multi-period inventory systems due to temporal carryover and cannibalization, proposing and validating a pairwise experimental design that effectively mitigates these interference effects.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are the captain of a massive, floating grocery store that sails through time. Every day, you have to decide how much of every single item—from bananas to batteries—to stock on your shelves. If you guess wrong and run out, customers leave empty-handed (a "lost sale"). If you guess too high, you waste money storing things that might rot. Now, imagine you want to test a new, smarter way to make these guesses. You can't just try it on the whole ship at once because if it fails, you lose everything. So, you decide to run a "test drive." You split your crew into two groups: one group keeps using the old guessing method (the "control"), and the other tries the new method (the "treatment"). This is called an A/B test, the gold standard for figuring out if a new idea actually works.
But here's the tricky part: your ship has a limited amount of space in its cargo hold. If the "new method" group orders too many bananas, they might take up all the space, leaving no room for the "old method" group's apples. Or, if the new method leaves a huge pile of unsold bananas today, those bananas might still be there tomorrow, messing up the test for the next day. In the world of science, this is called "interference"—where the test of one thing accidentally changes the results of another. The big question is: How do you run a fair test on a crowded, time-traveling grocery ship without the results getting messed up by the cargo hold's limits or the leftover fruit?
This paper dives deep into that exact problem, specifically for inventory systems with lost sales and capacity limits. The authors, Xinqi Chen and colleagues, act like detectives investigating three different ways to run these tests: Switchback (where the whole ship switches between old and new methods every few days), Item-Level Randomization (where specific items are permanently assigned to old or new methods), and Pairwise Randomization (a mix where every single item on every single day gets its own coin flip). They discovered that the "best" way to test depends entirely on what is wrong with the old guessing method.
If the old method was just consistently underestimating how many people want to buy things (a "mean bias"), the Switchback test is a trap. It tends to make the new method look worse than it really is because of leftover inventory from the "good" days carrying over into the "bad" days. Meanwhile, the Item-Level test tends to make the new method look too good because the new items hog the limited shelf space, starving the old items and making the new method look like a hero. In this case, the authors found that Pairwise Randomization is the sweet spot, balancing out these errors.
However, if the old method was actually guessing the right average number of sales but was just very "jittery" or unpredictable (a "dispersion" problem), the story flips. Now, the Switchback test makes the new, steadier method look even better than it is (because the new method leaves less messy leftover inventory), while the Item-Level test actually works perfectly and gives an unbiased answer. The Pairwise test, in this scenario, starts to suffer from the same "leftover" bias as the Switchback.
The authors didn't just guess these outcomes; they built complex mathematical models to prove the direction of these biases and ran thousands of computer simulations to verify them. They even tested their theories on real-world data from a fresh-food retail network (FreshRetailNet-50K), where they simulated stockouts and customers swapping items when their first choice was gone. The results held up: the same mechanisms that caused bias in their simple math models also caused bias in the messy, real world.
So, what's the takeaway for the captain of the floating grocery store? Don't just pick a testing method because it sounds cool. If you are trying to fix a method that consistently underestimates demand, avoid switching the whole ship back and forth; instead, try assigning different items to different methods, or use the clever "Pairwise" mix. But if you are testing a method that is just more consistent and less jittery, you can safely stick items to their methods for the long haul. The paper warns that blindly using the popular "Switchback" method in inventory systems is risky and often leads to systematic errors, proving that in the chaotic world of inventory, one size does not fit all.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.