Doubly Robust Estimation of Causal Effect on CVR with Targeted Regularization
This paper proposes a novel doubly robust causal effect estimator for post-click conversion rates (CVR) that leverages semiparametric theory and targeted regularization to overcome sample selection bias and variance issues, offering superior convergence rates and stability compared to existing methods that combine loss debiasing with standard causal estimators.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to figure out if a new marketing trick actually makes people buy things. In the world of online shopping and advertising, there's a famous two-step dance: first, a user has to click on an ad (the "click"), and second, they have to actually buy the product (the "conversion"). The rate at which people who click go on to buy is called the Conversion Rate, or CVR. It's like measuring how many people who walk into a store actually leave with a bag.
But here's the tricky part: if you only look at the people who clicked, you might get the wrong answer. It's like trying to judge how good a restaurant is only by asking the people who walked through the door, while ignoring everyone who saw the menu, thought it looked too expensive, and walked away. Those people who didn't click might have been the ones who would have bought the most if they had been given a different coupon or a better deal. This is called "selection bias." Scientists have been trying to fix this for a long time using math that guesses what would have happened if everyone had clicked, but the old tools often break when the data gets messy or when the computer models are too flexible. This paper steps into that detective story to build a better, more reliable tool for solving the mystery of whether a strategy actually works.
The authors, Jiayi Dan, Bo Li, and their team, propose a brand-new way to calculate these effects that they call a "doubly robust" estimator. Think of it like a high-tech safety harness for a climber. In their method, they need to guess three different things about the data (like how likely someone is to click, how likely they are to buy, and how the ad was shown). Usually, if your guess on even one of those things is wrong, your whole calculation falls apart. But this new "safety harness" is special: it's designed so that even if your guesses on two of those things are a bit off, as long as at least one of them is right, the final answer is still accurate. It's like having a backup parachute that works even if the first one fails.
To make this math work in the real world, where things can get wobbly and unstable, the team added a clever trick called "targeted regularization." Imagine you are trying to balance a stack of plates while riding a unicycle. The old way of doing this math was like trying to balance the stack by making sudden, jerky adjustments that often knocked everything over. The new method is like having a gentle, automatic gyroscope that makes tiny, smooth corrections to keep the stack steady without throwing it off balance. They tested this idea on both made-up data and real-world data from a massive dataset called CRITEO-UPLIFTv2 (which has about 13 million samples). The results showed that their new method was much better at finding the true effect than the old ways, and it stayed stable even when the settings were tweaked.
The paper also tackles a common temptation: just taking the old "loss debiasing" tricks used for predicting sales and slapping them onto these new causal questions. The authors found that this "mix-and-match" approach doesn't work well. It's like trying to fix a broken car engine by just painting the hood; it might look nice, but the engine still won't run right. They proved that you can't just fix the math for the prediction part and expect the final answer to be correct; you have to build the whole engine from scratch with the right parts. Their new framework, which combines the safety harness with the smooth gyroscope, consistently beat other popular methods in their experiments, showing that it's a much more reliable way to figure out if a marketing strategy is truly helping or hurting the user experience.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.