← Latest papers
📊 statistics

Flexible Design Strategies and Position Modeling for Order-of-Addition Screening Experiments

This paper introduces flexible linear, quadratic, and second-order component-position screening models that utilize position zero to encode component omission, enabling efficient estimation of nonlinear absolute-position effects and the construction of optimal designs for order-of-addition experiments where only a subset of components is available per run.

Original authors: Bing Wen, Bo Hu

Published 2026-08-28
📖 5 min read🧠 Deep dive

Original authors: Bing Wen, Bo Hu

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the laboratories where new medicines are mixed, where flavors are blended, and where complex chemical reactions are triggered, the order in which ingredients are added often matters as much as the ingredients themselves. A scientist might find that adding a catalyst before a solvent produces a brilliant result, while reversing that sequence yields nothing but waste. This is the world of order-of-addition experiments, where the sequence is the variable. However, when a researcher has a dozen potential ingredients but can only add a few at a time, the number of possible sequences explodes into the millions. Testing every single combination is impossible; the time and resources required would far exceed any realistic budget. The challenge, then, is to figure out which few sequences to test so that the scientist can learn the most about the process without running every possible experiment.

For years, statisticians have tried to build mathematical maps to navigate this complexity. Some maps focused on the relative order of ingredients, asking simply whether one item came before another. Others looked at the absolute position, asking if an ingredient was added first, second, or third. While these methods worked in some cases, they often hit a wall when the number of ingredients grew large or when the budget for experiments was tight. The existing maps were either too detailed to draw with limited resources or too simple to capture the subtle, non-linear ways that position affects the final result. They struggled to handle the reality that sometimes an ingredient is left out entirely, a scenario that older models treated as a confusing gap rather than a clear signal.

Bing Wen and Bo Hu have proposed a new way to draw these maps, one that is flexible enough to fit the size of the available budget. They introduced a set of models that treat the absence of an ingredient not as a missing piece of data, but as a specific position in the sequence, effectively assigning it a "zero" spot. This simple shift allows the models to remain lean and manageable. They developed three versions of this approach: a basic linear version for straightforward relationships, a quadratic version to catch curved or accelerating effects, and a second-order version that can detect how two different positions might interact to change the outcome. By matching the complexity of the model to the number of experiments a researcher can actually afford, they created a system that avoids the need for excessive data or complicated mathematical corrections.

The researchers proved that their new framework works by showing that the most complete set of experiments, if one could run them all, would be perfectly efficient under these new rules. More importantly, they demonstrated that many existing experimental designs, which were previously thought to be optimal only for older models, remain just as powerful under these new, more flexible rules. They also built new, highly efficient designs for situations where the number of ingredients and the number of available slots do not fit the neat patterns required by older methods. These new designs allow scientists to run far fewer experiments while still capturing the essential behavior of the system. In one specific case involving six ingredients and five slots, their method found an optimal design using only sixty runs, a fraction of the seven hundred and twenty runs required by the full set of possibilities, and far fewer than the seventy-two runs needed by the best previous method for a different model.

To see if these theoretical designs actually work in the real world, the authors ran computer simulations that mimicked two distinct scientific challenges. The first was a constrained job-scheduling scenario, where the goal was to find the most efficient order to process a set of tasks. In this test, their new model, using a budget of thirty-three experiments, produced predictions that were significantly more accurate than a random selection of experiments. It reduced the error in predicting the best schedule by nearly half and cut the "regret"—the cost of choosing a suboptimal sequence—by more than forty percent compared to random guessing. The second simulation involved a multi-drug screening scenario, where five drugs out of eight candidates had to be administered in a specific order. Here, the relationship between position and effect was non-linear and involved interactions between drugs. The researchers used a model capable of detecting these complex curves and interactions with forty-eight experiments. The results showed that their design not only predicted the outcome better than a random approach but also identified the top-performing drug sequences with much higher reliability.

The work suggests that by rethinking how we model the absence of an ingredient and by allowing the complexity of the math to grow only as fast as the experimental budget allows, scientists can extract more value from fewer trials. The new models do not just offer a theoretical improvement in efficiency; they translate directly into better predictions and more reliable decisions in practical settings where time and resources are limited. Whether in chemical synthesis, pharmaceutical development, or operational planning, the ability to screen sequences effectively without running every possible permutation is a powerful tool. The authors conclude that while their current designs are a major step forward, future work will focus on refining these models further to handle even more complex mixtures of effects and to ensure they remain robust when the real world does not follow the perfect patterns of a simulation.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →