← Latest papers
📊 statistics

Learning Consumer Preferences from Bundle Sales Data

This paper proposes a novel approach using an EM algorithm and Monte Carlo simulation to estimate consumer valuation distributions from bundle sales data, overcoming the limitations of classical discrete choice models when customers purchase multiple products.

Original authors: Ningyuan Chen, Setareh Farajollahzadeh, Qingwei Jin, Fanni Shen, Guan Wang

Published 2026-07-03
📖 4 min read☕ Coffee break read

Original authors: Ningyuan Chen, Setareh Farajollahzadeh, Qingwei Jin, Fanni Shen, Guan Wang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a store owner trying to figure out exactly how much your customers value your products. You want to know: "Is this customer willing to pay $10 for a shirt, or only $5?"

In a perfect world, you could just ask them. But in the real world, you can't ask. You only see what they buy. If they buy the shirt for $5, you know they value it at least $5, but you don't know if they would have paid $20. If they buy a "Combo Deal" (a shirt + pants for a discount), you know they like the combination, but you have no idea how much they liked the shirt versus the pants individually.

This paper is like a detective story where the authors invent a mathematical "mind-reading" tool to solve this mystery using only the receipt data.

The Core Problem: The "Bundle" Puzzle

Stores love selling bundles (like a burger, fries, and a drink together). It's great for sales, but it's a nightmare for data analysis.

  • The Old Way: Traditional math models treat every bundle as a totally new, unrelated item. It's like saying a "Burger+Fries" combo is a completely different species from a "Burger" and a "Fries." This ignores the fact that the combo is made of the parts you already know.
  • The Missing Data: When a customer walks out without buying anything, the store usually doesn't record that. The data is "censored" (hidden). The store only sees the people who bought something, which skews the picture.

The Solution: The "Guess-and-Refine" Machine

The authors propose a new algorithm (a step-by-step computer recipe) called the EM Algorithm (Expectation-Maximization). Think of it like a sculptor trying to carve a statue out of a block of marble, but they can't see the whole block at once.

  1. The Guess (Expectation): The computer starts with a wild guess about what the customers' "hidden values" look like. It imagines a crowd of invisible customers, each with a specific price they are willing to pay for every item.
  2. The Check (Maximization): The computer looks at the real sales receipts. It asks: "If my guess about the invisible customers were true, would they have made the choices they actually made?"
  3. The Refine: If the guess doesn't match the receipts, the computer tweaks the invisible customers' values slightly. It repeats this process thousands of times, slowly refining the "invisible crowd" until their behavior perfectly matches the real-world sales data.

The Special Tricks

The paper adds some clever "special effects" to make this machine work better in the real world:

  • The "Polyhedron" Map: Instead of just guessing a single price, the computer draws a complex, multi-sided shape (a polyhedron) in a virtual space. This shape represents all the possible price combinations a customer could have had that would lead them to make the specific purchase they made. It's like narrowing down a suspect's location from "somewhere in the city" to "inside this specific building."
  • The "Synergy" Effect: Sometimes, two things are better together than apart (like a printer and ink). The authors added a feature to detect this "synergy." If the data shows people buying the printer and ink together much more often than expected, the algorithm learns that there is a "magic bonus" value when they are paired up.
  • The "Ghost" Customers: To fix the problem of missing "no-purchase" data, the algorithm invents "ghost" customers who didn't buy anything. It estimates how many ghosts there must have been to explain why the real customers bought what they did.

Does It Work?

The authors tested their tool in two ways:

  1. Fake Data: They created a fake store with known customer values and let the algorithm try to find them. It succeeded, even when the data was messy or incomplete.
  2. Real Data: They used real transaction records from JD.com (a massive Chinese online retailer). They compared their new tool against standard methods used by marketing experts. Their tool was much better at predicting what customers would buy next and at figuring out the true value of individual items hidden inside bundles.

The Bottom Line

This paper provides a practical "decoder ring" for retailers. It allows them to take a pile of messy sales receipts—full of bundles, discounts, and missing "no-buy" records—and mathematically reverse-engineer exactly what their customers are thinking and how much they value every single item. This helps stores set better prices and create better deals in the future.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →