Synthetic Customer 360 Benchmark for Customer Data Quality, Identity Resolution, and Survivorship in Omnichannel Retail
This paper introduces a synthetic Customer 360 benchmark with auditable ground truth to rigorously evaluate and statistically validate the performance of identity resolution and survivorship rules in omnichannel retail, demonstrating reproducible condition separation while clarifying that these findings do not establish real-world operational superiority.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to solve a giant, messy puzzle where every piece is a tiny clue about a person's life. Some clues are written in perfect handwriting on a VIP card, while others are scribbled on a napkin in a crowded coffee shop. In the world of online and offline shopping, companies collect millions of these clues—emails, phone numbers, addresses, and names—from different places like websites, mobile apps, and physical stores. The big challenge is Identity Resolution: figuring out that the "John" who bought a shirt on a phone app is the same "John" who picked up a package at a store. Once you know they are the same person, you face Survivorship: if the phone app says his address is "123 Maple" but the store says "124 Maple," which one do you trust to be the "Golden Record" (the one true version)? The problem is that companies can't easily share their real customer data to test their systems because it's private and sensitive. So, scientists need a way to build a fake, but perfectly accurate, version of this puzzle to see if their rules work.
This paper introduces a clever new tool called a Synthetic Customer 360 Benchmark. Think of it as a "training simulator" for customer data, built by an independent researcher named Pradeep Aronkar. Instead of using real people's private information, the author wrote a computer program to generate a fake world of shoppers. This program creates thousands of fake people, gives them fake accounts across four different "worlds" (like an online store, a loyalty app, a physical shop, and a service center), and then deliberately breaks the data in specific, controlled ways. It introduces "defects" like missing names, typos, or conflicting addresses, and even creates "ambiguity" where two different people might accidentally share the same phone number. The magic of this tool is that the creator knows the exact truth: who is really who, and what the correct answer should be for every single fake person.
The paper doesn't just build the simulator; it puts it to the test. The author ran the program 335 times to see if different computer rules could solve the puzzle under different levels of difficulty. They tested two main ideas. First, they asked: "Can our system tell the difference between easy puzzles (where clues are clear) and hard puzzles (where clues are messy or missing)?" The answer was a resounding yes. When the data was made "hard" with more errors and confusion, the computer's ability to match people correctly dropped significantly, proving the system could actually feel the difference in difficulty. Second, they asked: "If we have conflicting information, do different rules for picking the 'winner' value lead to different results?" Again, the answer was yes. When the fake data had more conflicts, the different rules disagreed with each other much more often.
The study found that while the computer rules worked well on easy data, they struggled when the data was messy, and that different "survivorship" rules (like trusting the newest info vs. the most verified info) often produced different "Golden Records" when the data was messy. However, the author is very careful to say this isn't a magic bullet for real-world companies yet. Because the data is entirely synthetic (made up by a computer), it doesn't prove that these rules will work perfectly for a real store with real customers. The paper explicitly states it does not claim to have found the "best" rule for everyone, nor does it prove that these methods scale up to massive real-world populations. Instead, it offers a transparent, auditable way for researchers and engineers to test their own ideas against a known truth, ensuring that when they do build real systems, they are testing them on a level playing field where the answers are already known.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.