Action-Aware Generative Sequence Modeling for Short Video Recommendation
This paper proposes Action-Aware Generative Sequence Network (A2Gen), a novel model that leverages the temporal patterns of user actions to overcome the limitations of traditional binary classification in short video recommendation, achieving significant improvements in watch time, interaction rates, and user retention through extensive offline and large-scale online A/B testing on Kuaishou.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are watching a short video on your phone. It's a mix of different things: maybe a funny joke, a serious news clip, a catchy song, and then a sudden ad.
The Old Way (The "One-Size-Fits-All" Approach)
Traditional recommendation systems treat the whole video like a single, solid block of cheese. If you "Like" the video, the system assumes you liked everything in it.
- The Problem: Let's say the video is an interview where the guest chooses "Messi" over "Ronaldo." You "Like" the video right at that moment because you love Messi. But the rest of the video is just highlights of Ronaldo. The old system thinks, "Oh, they liked the video, so they must love Ronaldo too!" and starts showing you Ronaldo videos. You get annoyed because you only liked the Messi part.
The New Insight (Timing is Everything)
The authors of this paper realized that when you click a button matters just as much as that you clicked it.
- The Analogy: Think of a video like a movie trailer. If you clap your hands (a "Like") right when the hero saves the day, you're reacting to that specific moment. If you clap at the very end, you might just be being polite. The timing tells the system exactly which part of the video made you happy.
- The Discovery: By looking at millions of videos, they found that "Likes" and "Follows" often happen in specific "peaks" (highlights). If you "Follow" a creator before you even finish watching the video, it means you love the person, not necessarily the specific content you just saw.
The Solution: A2Gen (The "Storyteller" Model)
The paper introduces a new model called A2Gen (Action-Aware Generative Sequence Network). Instead of guessing if you'll like a video, A2Gen tries to predict the entire story of how you will watch it.
Here is how it works, broken down into simple parts:
The Context-Aware Attention Module (CAM) – "The Smart Reader"
Imagine a reader who doesn't just read words but understands the mood of the room. CAM looks at your past actions and the specific video content together. It knows that a "Like" on a comedy video means something different than a "Like" on a news video. It pays attention to the context of the moment.The Hierarchical Sequence Encoder (HSE) – "The Memory Bank"
This part remembers your long-term habits. It looks at your history like a librarian organizing books. It doesn't just see "User X watched 100 videos." It sees patterns: "User X usually skips ads but loves music videos, and they tend to 'Follow' creators after 15 seconds." It builds a deep profile of your unique style.The Action-seq Autoregressive Generator (AAG) – "The Crystal Ball"
This is the magic part. Instead of just saying "Yes/No," AAG acts like a crystal ball that predicts the entire sequence of your future actions.- It asks: "If we show this video, will the user start watching? Will they 'Like' it at the 5-second mark? Will they 'Follow' the creator at the 10-second mark? When will they stop?"
- It predicts the order and the exact timing of every click, like predicting the plot of a movie before it happens.
Why This Matters (The Results)
The team tested this on Kuaishou, a massive short-video platform with hundreds of millions of users. They replaced their old system with A2Gen and ran a huge experiment (A/B testing).
- The Outcome: Because the system now understands which parts of a video you actually enjoy, it recommends better content.
- The Numbers:
- People watched 0.34% more video time (which sounds small, but with millions of users, that's a lot of time).
- People interacted (liked, followed, commented) 8.1% more.
- More importantly, more people came back the next day (retention went up by 0.162%), which translates to about one million extra daily users staying on the platform.
In Summary
The paper argues that we shouldn't treat a video as a single "Yes/No" item. Instead, we should treat it as a timeline of moments. By using a model that predicts when and in what order you will react, the app can finally understand your true taste and show you videos you actually want to watch, rather than just guessing based on a single click.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.