M3TR: Temporal Retrieval Enhanced Multi-Modal Micro-video Popularity Prediction
The paper proposes M3TR, a temporal retrieval-enhanced multi-modal framework that leverages a Mamba-Hawkes Process to model self-exciting user feedback dynamics and a temporal-aware retrieval mechanism to capture popularity evolution, thereby achieving state-of-the-art performance in micro-video popularity prediction.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the vast, churning ocean of short-form video content, where millions of clips are uploaded every day, a single question drives the engines of recommendation algorithms: which video will capture the world's attention, and for how long? Predicting this popularity is not merely a matter of counting likes or shares; it is an attempt to forecast the future behavior of a crowd based on the fleeting, often chaotic reactions of individuals. For years, researchers have tried to build models that can see the future of a video's success by analyzing its content—what it looks like, what it sounds like, and what the text says about it. They have also tried to track how users interact with a video over time, treating those interactions as a simple line graph that goes up or down. However, these approaches often miss the deeper, more complex rhythm of human engagement. They fail to understand that a burst of likes might trigger a wave of shares, or that a sudden flood of negative comments can instantly kill a video's momentum, regardless of how beautiful the video itself is.
A team of researchers from Shanghai Jiao Tong University and Queen's University Belfast has proposed a new way to solve this puzzle, introducing a system they call M3TR. Instead of just looking at what a video is made of, or treating user feedback as a simple number that changes over time, this new approach treats popularity as a living, breathing sequence of events. The researchers realized that to predict the future of a video, you must understand the specific, intricate chain reaction of human reactions. They built a system that learns to recognize the unique "temporal DNA" of a video's success. This means the system doesn't just ask, "Is this video about cats?" It asks, "Does this video follow the same pattern of excitement and decline as other videos that became famous, even if those other videos were about cooking or dancing?" By combining a deep understanding of how user interactions influence one another with a smart search engine that finds videos with similar popularity histories, the team created a tool that sees the future more clearly than ever before.
The core of their discovery lies in how they modeled the way people react to videos. In the past, systems often treated a "like," a "share," and a "comment" as independent events, or simply added them up into a single score. The researchers knew this was wrong. They understood that these actions are deeply connected: a surge of likes can encourage more shares, while a wave of controversial comments can suppress further engagement. To capture this, they developed a new method that views user feedback as a series of self-exciting events, where one action naturally triggers the next, but where the nature of that trigger can change based on the context. They used a sophisticated neural network architecture to learn the long-term patterns in these sequences, allowing the system to see that a specific type of comment might signal the end of a video's popularity, even if the video is still receiving likes. This ability to distinguish between positive momentum and a false peak was a critical breakthrough.
Once the system could accurately map the complex history of a video's interactions, the researchers faced the next challenge: how to use that knowledge to predict the future. Traditional methods would search for other videos that looked or sounded similar. If you had a video about a cat, the system would look for other cat videos. The researchers found this approach flawed. A video about a cat might have a slow, steady rise in popularity, while another cat video might explode instantly and then fade away. The content was the same, but the story of their success was completely different. Their new system, M3TR, changed the rules of the search. Instead of just matching content, it matched the "story" of the popularity. It searched a massive library of historical videos to find examples that shared the same pattern of engagement—videos that rose and fell in the same way, regardless of whether they were about cats, cars, or cooking.
This shift from content-matching to pattern-matching proved to be incredibly powerful. When the researchers tested their system on real-world data from two large datasets, the results were striking. The new model significantly outperformed all previous methods, reducing the error in its predictions by up to 19.3 percent. In simpler terms, it was far more accurate at guessing how many people would watch, like, or share a video in the coming hours and days. The system was particularly good at handling the most difficult cases: videos that started strong but then faded, or videos that lay dormant before suddenly going viral. By learning from the specific trajectory of past videos, the model could spot the subtle signs that a video was about to stall or take off, something that older models missed entirely.
The researchers also demonstrated that their approach was practical for real-world use. While the system requires a significant amount of computing power to build its library of video patterns, this work is done in advance, offline. Once the library is built, the system can make predictions for new videos in real-time, finding the most relevant historical examples in a fraction of a second. This two-step process ensures that the system remains fast and efficient for users, even as the library of data grows to include millions of videos. The study confirms that understanding the dynamic, evolving story of how people interact with content is far more important for prediction than simply analyzing the content itself. By listening to the rhythm of human attention rather than just looking at the picture, the researchers have provided a new, more reliable way to navigate the unpredictable world of viral media.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.