← Latest papers
🤖 machine learning

FreshRetailNet-LT: A Stockout-Annotated Censored Demand Dataset for Latent Demand Recovery and Forecasting in Fresh Retail

This paper introduces FreshRetailNet-50K, the first large-scale benchmark dataset featuring 50,000 hourly store-product time series with precise stockout annotations, which enables researchers to reconstruct latent demand and significantly improve forecasting accuracy for perishable retail products by correcting systemic biases caused by censored sales data.

Original authors: Yangyang Wang, Jiawei Gu, Li Long, Xin Li, Li Shen, Zhouyu Fu, Xiangjun Zhou, Xu Jiang

Published 2026-06-10
📖 4 min read☕ Coffee break read

Original authors: Yangyang Wang, Jiawei Gu, Li Long, Xin Li, Li Shen, Zhouyu Fu, Xiangjun Zhou, Xu Jiang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are running a busy lemonade stand. You want to know exactly how many cups people want to buy so you can make the perfect amount of lemonade. If you make too much, it goes sour (waste); if you make too little, you miss out on sales.

The problem is, your sales record is a liar.

The Problem: The "Empty Cup" Illusion

In the world of fresh food (like strawberries, fish, or milk), stores often run out of stock. Let's say you have 100 customers who want strawberries, but you only have 50. You sell out by 11:00 AM.

Your sales record for the rest of the day says: "0 sales."

But that's not true! The demand was still there; you just ran out of inventory. This is called censored data. It's like a camera that stops recording the moment the battery dies. If you try to learn from this broken camera, you'll think, "Oh, nobody wants strawberries after 11 AM," and you'll order even fewer the next day. This creates a vicious cycle where you constantly run out of stock and lose money.

The Solution: FreshRetailNet-LT

The authors of this paper created a new, massive dataset called FreshRetailNet-LT to fix this broken camera. Think of it as a "super-detailed diary" for fresh food sales.

Here is what makes it special, using simple analogies:

1. The High-Definition Video vs. The Daily Summary
Most old datasets are like a daily summary note: "Sold 50 apples today." They don't tell you when you ran out.
This new dataset is like a high-definition video recorded every hour. It shows you exactly when the shelf went empty. It records 20,000 different product stories from over 1,000 stores across 18 cities, spanning two years. Because it tracks things hour-by-hour, it can spot that you ran out of strawberries at 11:00 AM, even if you sold nothing at 2:00 PM.

2. The "Truth Label"
Usually, computers have to guess if a store ran out of stock. This dataset comes with verified "stockout labels." It's like having a teacher who puts a gold star next to every time the shelf was actually empty. This allows researchers to teach computers the difference between "nobody wanted it" (true zero demand) and "we ran out" (censored demand).

3. The Context Clues
The dataset doesn't just count sales; it records the "weather report" for the sales. It includes:

  • Promotions: Did we have a 50% off sale? (This usually causes a huge spike in demand).
  • Weather: Is it raining? (People might buy more fresh veggies to cook at home).
  • Holidays: Is it Chinese New Year? (People buy different things).

How They Tested It: The Two-Step Fix

The researchers used this new dataset to test a two-step method to fix the "lying sales record":

  • Step 1: The Detective Work (Recovery)
    First, they used the "gold star" labels to reconstruct the missing sales. If the shelf was empty at 11 AM, the model estimates how many people would have bought strawberries if the shelf hadn't been empty. It's like filling in the missing frames of a movie.
  • Step 2: The Crystal Ball (Forecasting)
    Second, they used this "reconstructed truth" to teach a computer how to predict future demand.

The Results: A Clearer Picture

When they tested this method, the results were impressive:

  • Less Guessing: The computer stopped underestimating demand. Before, it was wrong by about 6.7% (thinking people wanted less than they actually did). After using the new method, the error dropped to near zero.
  • Better Accuracy: The overall prediction accuracy improved by 1.63%. In the world of selling millions of dollars of fresh food, that small percentage saves a lot of money and reduces waste.

Why It Matters

This paper isn't just about math; it's about fixing a broken system. By providing a dataset that admits "we ran out of stock" and tells you exactly when, it helps stores stop the cycle of running out of fresh food. It allows them to order the right amount of perishable goods, ensuring that when you go to the store, the strawberries are actually there.

The authors have made this "super-detailed diary" and their code available for anyone to use, hoping to help both scientists and store owners solve the puzzle of fresh food demand.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →