← Latest papers
📈 economics

IoT technology for e-commerce customer information segmentation

This paper proposes an improved RFM model enhanced by the entropy weight method and K-means clustering to effectively segment e-commerce customer value for differentiated marketing, as validated by empirical data from a Pinduoduo merchant.

Original authors: Qiuping Lu

Published 2026-09-08
📖 4 min read☕ Coffee break read

Original authors: Qiuping Lu

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the bustling world of online shopping, a merchant faces a challenge as old as commerce itself: how to treat every customer as an individual without losing their mind. In the past, a shopkeeper knew their regulars by name and face. Today, a digital store might serve millions of people, most of whom are just a username and a shipping address away. To survive, these businesses must sort their customers into groups, understanding who brings the most profit and who needs a little extra attention. This is the art of customer segmentation. Traditionally, businesses have looked at three simple things to judge a shopper's value: how recently they bought something, how often they buy, and how much money they spend. This trio of facts forms a basic map of a customer's habits. However, simply adding these numbers together often misses the nuance of real behavior. A customer who spends a lot but rarely visits might be different from one who visits daily but spends little. The question remains: how do we weigh these factors fairly to see the true value of a shopper?

A researcher at Henan Polytechnic Institute set out to answer this by refining the way online merchants analyze their data. Instead of guessing which of the three factors—recency, frequency, or spending—matters most, the researcher used a mathematical approach called the entropy weight method to let the data speak for itself. This technique calculates the importance of each factor based on how much the data varies, ensuring that the final score reflects the actual patterns in the sales records rather than a pre-set opinion. The study focused on a specific merchant selling cookies and cakes on the Pinduoduo platform, examining over 32,000 purchase records from August 2019 to January 2020. After cleaning the data to remove invalid orders and identifying customers by their shipping addresses, the researcher applied this new scoring system to evaluate every single buyer.

Once the customers were scored, the researcher used a clustering tool to group them. Imagine sorting a pile of mixed stones not by color or size alone, but by finding the natural clusters where stones of similar weight and texture naturally settle together. The computer tested different numbers of groups, from two to eight, and selected the arrangement that created the most distinct and clear categories. The result was a map of four distinct types of shoppers. The first group, making up about 27 percent of the customers, were the low-value shoppers: they had not bought anything recently, bought infrequently, and spent very little. The second group, representing nearly 23 percent, were customers who had spent a good amount and bought often in the past but had gone quiet recently; these were the shoppers the merchant needed to win back. The third group, another 27 percent, were recent buyers who were still exploring the store but had not yet committed to high spending or frequent visits. Finally, the fourth group, comprising about 22.5 percent of the total, were the high-value customers: they bought recently, bought often, and spent the most money.

When the researcher compared this new method against the traditional way of scoring customers, the difference was clear. The traditional approach, which often relies on fixed weights decided by experts, tended to lump the most valuable shoppers into a tiny fraction of the total, identifying only a tiny 0.8 percent as high-value. In contrast, the new method, which let the data determine the weights, identified a much larger and more realistic group of high-value customers, aligning closely with the well-known business observation that a small portion of customers generates the majority of profits. The new model showed that the groups were more distinct from one another, meaning the merchant could now see the differences between customer types with greater clarity. This suggests that by letting the data decide how much weight to give to recency, frequency, and spending, online businesses can create more accurate profiles of their shoppers. The study concludes that this approach offers a better way to segment customers, allowing merchants to tailor their strategies to keep their best shoppers happy and gently encourage others to return, all based on a clearer picture of who they really are.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →