← Latest papers
📊 statistics

Multi-Experiment Analysis

This paper introduces Multi-Experiment Analysis (MEA), a methodology that enables consistent joint estimation of individual, combined, and conditional treatment effects in online controlled experiments with arbitrary overlaps, eliminating the need for restrictive traffic splitting or factorial pre-design while being validated through simulations and large-scale production deployment.

Original authors: Reza Hosseini

Published 2026-04-21
📖 5 min read🧠 Deep dive

Original authors: Reza Hosseini

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are the head chef of a massive, bustling restaurant. Every day, you want to try new recipes to make your dishes better. You have a team of sous-chefs, each working on a different part of the menu: one is tweaking the sauce, another is changing the plating, and a third is adjusting the spice level.

In the old days, to test these changes safely, you would have to run them one at a time.

  • Monday: Test the new sauce on 50% of customers.
  • Tuesday: Once Monday is done, test the new plating on 50% of customers.

This is slow. By the time you figure out the perfect sauce, the plating, and the spice, it's already next month. You've lost time and money.

To speed things up, you decide to let all three chefs work at the same time on the same group of customers. But here's the problem: The Kitchen Chaos.

If the sauce chef adds extra salt, and the spice chef adds extra heat, the customer gets a dish that is both salty and spicy. If you just look at the "Salty" group and the "Spicy" group separately, you get confused.

  • "The salty dish was great!" (But it was actually great because it was also spicy).
  • "The spicy dish was terrible!" (But it was terrible because it was also salty).

You end up with bad data. You might launch a combination that tastes awful, or miss a combination that tastes like heaven.

Enter: Multi-Experiment Analysis (MEA)

This paper introduces MEA, a smart new way to look at the data from these "chaotic" simultaneous experiments. Instead of waiting your turn, MEA lets you run everything at once and then uses a special mathematical "magic lens" to untangle the mess.

Here is how it works, using our restaurant analogy:

1. The "L-Shape" Map (Partitioning)

Imagine the restaurant floor is a grid.

  • Some customers only saw the Sauce change.
  • Some only saw the Plating change.
  • Some saw both.
  • Some saw neither.

Old methods tried to ignore the people who saw "both" or treated them as a separate, confusing group. MEA says: "No, let's map everyone." It creates a detailed map of every possible combination of who saw what. It treats the people who saw only the sauce as a control group for the people who saw both.

2. The "Weighted Scale" (Consistent Estimation)

Now, imagine you want to know: "What happens if we launch the New Sauce AND the New Plating together?"

If you just average the results of everyone, you get a distorted number because the groups are different sizes.

  • Maybe 80% of people saw the sauce but not the plating.
  • Maybe only 10% saw both.

MEA acts like a smart scale. It weighs the results of the "Sauce Only" group, the "Plating Only" group, and the "Both" group according to how many people were in each group. It calculates a "fair" average that tells you exactly what would happen if you launched the perfect combo tomorrow.

3. The "What-If" Crystal Ball (Scenario Analysis)

Sometimes, you don't just want to know the average. You want to ask specific questions:

  • "If we launch the New Sauce, but we decide to keep the Old Plating, how will that do?"
  • "If we launch the New Plating, but the New Sauce is delayed, what happens?"

MEA can answer these "What-If" questions instantly. It simulates different launch scenarios using the data you already have, so you don't have to wait weeks to run a new test.

4. The "Truth Detector" (Assumption Checking)

There is one catch: This only works if the chefs don't accidentally change the number of people coming to the table.

  • Bad Scenario: The New Sauce is so delicious that it makes people stay longer, which accidentally triggers the "Plating" experiment for more people than usual. This breaks the math.
  • The Fix: MEA has a built-in "Truth Detector." Before giving you the results, it checks: "Did the New Sauce change who showed up for the Plating test?" If the answer is "Yes," it raises a red flag and says, "Hey, the data is messy, don't trust the numbers yet."

Why This Matters

Before MEA:

  • Slow: You had to wait for one test to finish before starting the next.
  • Confusing: You didn't know if a feature was good on its own or only good because of another feature running at the same time.
  • Risky: You might launch a "Frankenstein" product that looks good in isolation but fails when combined.

With MEA:

  • Fast: You can run hundreds of tests at once.
  • Clear: You know exactly which combination of features wins.
  • Safe: You can launch the "Best Sauce + Best Plating" combo with confidence, knowing the math accounts for the chaos.

The Bottom Line

The paper argues that in the modern world of technology (like Facebook, Google, or LinkedIn), we can't afford to wait. We need to test everything at once. Multi-Experiment Analysis is the toolkit that lets us do that without losing our minds or making bad decisions. It turns a chaotic kitchen into a well-orchestrated symphony, ensuring that when you finally serve the dish to the world, it's the best one possible.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →