MAC: A Conversion Rate Prediction Benchmark Featuring Labels Under Multiple Attribution Mechanisms
This paper introduces MAC, the first public conversion rate prediction benchmark featuring labels from multiple attribution mechanisms, along with the PyMAL library and the MoAE model, to address data gaps in multi-attribution learning and demonstrate significant performance gains through carefully designed knowledge integration.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to guess what a friend will buy next. In the world of online shopping, this is a high-stakes game played by computers. Every time you click on an ad, a digital detective tries to figure out if that click will lead to a purchase. This is called Conversion Rate (CVR) prediction. For years, these detectives have been working with a very narrow rulebook: they only give credit for a sale to the very last ad you clicked before buying. It's like saying if you walked past a bakery, then a shoe store, and finally a toy shop before buying a toy, the toy shop gets all the credit, and the bakery and shoe store get nothing.
But in real life, our decisions are messy. Maybe the bakery made you hungry, or the shoe store gave you a great idea. To get the full picture, we need Multi-Attribution Learning (MAL). This is a smarter way of thinking that says, "Hey, let's look at all the ads you clicked, not just the last one." It tries to understand the whole journey, giving credit to the first click, the last click, or even splitting the credit evenly among all of them. The big problem? Until now, scientists and companies only had data that followed the "last click only" rule. They were trying to learn a complex, multi-step story using a single-sentence summary.
This paper, titled "MAC: A Conversion Rate Prediction Benchmark Featuring Labels Under Multiple Attribution Mechanisms," is like opening a secret library that finally has the full story. The authors, a team from Alibaba and Nanjing University, created a brand new dataset called MAC. This isn't just a list of clicks; it's a massive collection of 79 million ad clicks where they recorded four different ways to give credit for a sale: Last-Click, First-Click, Linear (splitting credit equally), and a fancy Data-Driven method that uses math to figure out who really did the most work. They also built a free software toolbox called PyMAL so anyone can test their ideas on this new data.
When they tested their new ideas on this data, they found some fascinating things. First, looking at the whole journey (Multi-Attribution Learning) consistently made the computers smarter, especially for people who take a long, winding path before buying something. Second, they discovered that just throwing more clues at the computer doesn't always help. If you try to predict a "First-Click" sale, adding too many other types of clues actually confused the model. It turns out you have to be very careful about which clues you use. Finally, they built a new model called MoAE (Mixture of Asymmetric Experts). Think of MoAE as a team of detectives where some specialize in reading the whole journey, but there's a strict boss who makes sure all that extra knowledge is used specifically to solve the main case. This new model beat all the previous champions, improving the prediction accuracy by up to 0.39 percentage points. The authors suggest that by using this new data and these smart design rules, we can finally build advertising systems that understand human behavior much better than before.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.