← Latest papers
🤖 machine learning

Sequential Multimodal Evidence Optimization for Product Media Ranking in E-Commerce

This paper introduces Sequential Multimodal Evidence Optimization (SMEO), a two-stage framework that optimizes the sequential ordering of heterogeneous product media in e-commerce by learning trajectory utilities from biased logs and training an autoregressive ranking policy to help customers reach purchase decisions with fewer interactions and higher estimated conversion rates.

Original authors: Prasenjit Dey, Frank McIntyre, Arnab Sinha

Published 2026-08-18
📖 6 min read🧠 Deep dive

Original authors: Prasenjit Dey, Frank McIntyre, Arnab Sinha

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

On the digital shelves of modern online stores, the path to a purchase is rarely a single glance. Instead, it is a journey through a carousel of images, videos, and interactive three-dimensional models that a customer must swipe through one by one. Because shoppers cannot touch or hold the product, they rely on this sequence of visual clues to build a picture of what they are buying. They accumulate evidence, piece by piece, until they feel confident enough to buy or decide the effort is too great and leave. The challenge for the platforms that sell these goods is that the order in which these clues appear matters immensely. If the most convincing information is buried at the bottom of the list, a customer might give up before ever seeing it. This is not just about grabbing attention; it is about guiding a person through a story that leads to a decision.

Researchers at Amazon, led by Prasenjit Dey, Frank McIntyre, and Arnab Sinha, have developed a new way to solve this puzzle. They call their approach Sequential Multimodal Evidence Optimization, or SMEO. Rather than treating each image or video as an isolated item competing for a click, they view the entire collection of media for a single product as a cooperative team working together to tell a complete story. Their goal was to find the specific order that helps a customer reach a "yes" decision with the least amount of effort. By analyzing millions of real shopping sessions, they discovered that existing systems often fail because they focus on short-term reactions, like how long a user lingers on a screen, rather than the final outcome of whether a sale actually happened.

The researchers found that current methods often make the mistake of thinking that a flashy image that gets many clicks is the most important one to show first. However, a click does not always mean a purchase. A customer might click on a bright picture out of curiosity but still leave without buying because they never saw the technical details or the video demonstration that would have convinced them. The new system, SMEO, learns from the history of what customers actually did. It looks at the sequence of media a customer viewed before they either bought the item or stopped looking. It then figures out which pieces of evidence were the most persuasive in that specific context. For example, it might learn that for a coffee maker, a video showing the machine in action is far more critical than a static photo of the box, and that this video should be shown early in the sequence to prevent the customer from leaving.

To achieve this, the team built a two-step process. First, they trained a model to understand the value of a sequence of media based on whether it led to a sale. This model learned to ignore the bias of where items were placed in the past, focusing instead on the actual content and how it helped the customer decide. Second, they used this understanding to teach a computer program how to arrange the media for any new product. This program acts like an editor, picking the best order for the images and videos. It prioritizes showing the most decision-making information as early as possible, ensuring that the customer sees the most important proof before they get tired of swiping.

The results of testing this system on a massive scale were significant. When the researchers compared their new method against the standard rules used by online stores, they found that the new system increased the estimated rate of successful purchases by 5.5 percent. More importantly, it helped customers reach that decision 15 percent faster, meaning they needed to swipe through fewer images to find what they were looking for. The study also revealed that the system learned to recognize patterns without being explicitly told what to look for. For instance, it discovered that for furniture, three-dimensional models were far more valuable than for other types of products, while for clothing, videos showing how the fabric moved were the most persuasive. This allowed the system to automatically tailor the presentation of evidence to the specific needs of different product categories.

One of the most powerful aspects of this work is that it solves a problem that has been difficult to measure: how much credit should a specific image or video get for a sale? In the past, it was hard to tell if a purchase was caused by the first photo or the fifth video. The new system provides a way to look back and understand the contribution of each piece of media, even without direct labels saying "this image caused the sale." It does this by simulating what would have happened if a piece of media were removed or moved, allowing the system to see how the customer's journey would have changed. This insight helps store owners understand which assets are truly valuable and which are redundant, allowing them to improve their product pages more effectively.

The researchers tested their ideas rigorously using data from millions of real shopping sessions across various categories, from electronics to beauty products. They compared their method against several other approaches, including simple rules based on popularity and other advanced ranking systems. Their method consistently outperformed these alternatives, proving that treating media as a sequential story is more effective than treating it as a list of competing items. The system is designed to work offline, meaning the order of images is prepared in advance and stored, so it does not slow down the website when a customer visits. This makes it practical for large-scale use, where millions of products need to be organized every day.

Ultimately, this work shifts the focus from simply capturing a customer's eye to respecting their time and intent. By arranging visual evidence in a way that builds confidence quickly, the system reduces the friction of online shopping. It acknowledges that a customer's attention is a limited resource and that the goal is to help them find the information they need with the fewest possible steps. The findings suggest that when digital stores are organized with the customer's decision-making process in mind, the result is a better experience for the shopper and a more efficient marketplace for the seller. The study confirms that the order of information is not just a cosmetic detail but a fundamental part of how people make choices in the digital world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →