← Latest papers
📊 statistics

Blending Proxy Metrics with a North Star

This paper proposes an optimal blending framework for A/B testing that dynamically balances proxy metrics and a "north star" based on experiment power and proxy quality, offering actionable insights for experiment design and demonstrating its real-world application at Netflix.

Original authors: Winston Chou

Published 2026-06-23
📖 5 min read🧠 Deep dive

Original authors: Winston Chou

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a captain steering a massive ship (a company like Netflix) through foggy waters. Your ultimate goal is to reach a specific, distant island called "The North Star" (true business success, like keeping customers for years). However, the fog is so thick that you can't see the island clearly until you are almost there.

To navigate, you have two tools:

  1. The North Star Compass: It points directly to your destination, but it's very weak and jittery. You can only trust it if you sail for a very long time (run a huge experiment).
  2. The Proxy Radar: It detects small ripples in the water (user clicks, engagement) that happen before you reach the island. It's super sensitive and gives you a clear signal immediately, but it doesn't always mean you are actually heading toward the island.

The Dilemma
In the past, companies had to choose: "Do I trust the weak North Star, or the noisy Proxy?"

  • If you trust the Proxy too much, you might steer toward a ripple that looks promising but leads to a dead end.
  • If you wait for the North Star, you might sail for years without making any decisions, missing out on opportunities.

The Solution: The "Smart Blend"
Winston Chou's paper proposes a clever navigation system that blends these two tools. Instead of picking one, the ship's computer automatically adjusts how much it trusts each tool based on how much data it has collected.

Here is how the "Smart Blend" works, using simple analogies:

1. The "Volume Knob" Analogy

Imagine the decision-making process is a radio with two volume knobs: one for the Proxy (clicks) and one for the North Star (retention).

  • When the experiment is small (low data): The North Star signal is too quiet to hear. The system turns the Proxy volume up and the North Star volume down. You make decisions based on the immediate, sensitive signals because you don't have enough data to see the big picture yet.
  • When the experiment is huge (high data): You have sailed far enough that the North Star is finally visible. The system automatically turns the North Star volume up and the Proxy volume down. You stop guessing based on ripples and start steering by the actual destination.

The paper proves that there is a mathematically "perfect" way to turn these knobs at every stage of the journey.

2. The "Taste Test" Analogy

Imagine you are a chef trying to create a new dish.

  • The North Star is the final taste test by a food critic (did they love it?). This takes time and is rare.
  • The Proxy is the smell of the kitchen while cooking. It's immediate and tells you if something is burning or smelling good.

If you only have a tiny sample (one spoonful), you can't trust the critic's opinion yet, so you rely on the smell. But if you have cooked a thousand pots, the critic's opinion becomes the most important thing. The paper's method tells you exactly how much to weigh the "smell" versus the "critic's taste" based on how many pots you've cooked.

3. The "Traffic Jam" Analogy (Designing Experiments)

The paper also changes how you plan your experiments.

  • If your Proxy is excellent (like a very accurate smell), you don't need to cook huge batches to know if the dish is good. You can run many small, quick experiments. This is like taking many small side roads to get around traffic.
  • If your Proxy is poor (the smell is unreliable), you can't trust it. You have to wait for the critic. This means you need to run fewer, but much larger experiments to get a clear signal.

Real-World Proof: Netflix

The author tested this at Netflix.

  • The North Star: "Plays" (did the user actually watch the movie?). This is hard to measure quickly.
  • The Proxy: "Clicks" (did the user click the title?). This is easy to measure instantly.

In Netflix's data, "Clicks" were much more sensitive than "Plays."

  • For small tests, the system relied mostly on Clicks.
  • For massive tests (millions of users), the system shifted to rely mostly on Plays.
  • The Result: By blending them, Netflix got better results than if they had used only Clicks or only Plays. They made fewer mistakes and found more successful features.

The Bottom Line

You don't have to choose between "fast but noisy" data and "slow but accurate" data. By using a blended approach that automatically shifts its trust from the fast data to the accurate data as you gather more information, you can make better decisions faster.

Key Takeaways from the Paper:

  • Don't pick a side: Use both metrics, but weight them differently.
  • Size matters: The bigger your experiment, the more you should trust the "North Star" (the true goal).
  • Quality matters: If your "Proxy" is very good, you can run more, smaller experiments. If it's bad, you need fewer, bigger ones.
  • The math works: The paper provides a formula to calculate exactly how to mix these metrics to get the best possible outcome.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →