Navigating User Behavior toward Personalized Multimodal Generation
The paper proposes NaviGen, a framework that transforms user interaction history into executable instructions for personalized multimodal generation by employing dual identifiers and a two-stage SFT+RL pipeline to overcome challenges in behavior encoding and instruction writing.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a very talented, high-tech artist (an AI) who can paint beautiful pictures or create amazing videos based on your instructions. The problem is, this artist is incredibly literal. If you say, "Draw a fantasy scene," they might draw something generic that doesn't quite match your specific taste. Most people aren't professional artists, so they can't write the perfect, detailed instructions needed to get exactly what they want.
The Big Problem:
We have a huge gap between what users actually like (based on what they click, watch, or buy) and the specific instructions needed to tell the AI what to make. Users rarely say, "I want a cinematic wide shot of a Roman coliseum at dusk with volumetric fog." They just click on a few cool videos or buy a few games. The AI needs to figure out how to translate those clicks into a perfect instruction.
The Solution: NaviGen
The researchers created a system called NaviGen (Navigation + Generation) to act as a "translator" between your past behavior and the AI artist. Here is how it works, using some simple analogies:
1. The Dual-ID System (The "Name Tag" and the "Resume")
To understand you, NaviGen looks at every item you've interacted with (like a game or a video) and gives it two special "ID cards":
- The Collaborative Identifier (CID): Think of this as a behavioral fingerprint. It's a code that says, "People who liked this also liked that." It captures the pattern of your taste without needing to read the text. It's like a secret code that tells the AI, "This item belongs to the 'adventure' family."
- The Textual Identifier (TID): Think of this as a condensed resume. It's a short, organized list of keywords (like "fantasy," "romance," "action") that describes what the item is actually about.
By combining these two, NaviGen gives the AI a compact "behavioral substrate" (the code) and a "semantic bridge" (the keywords) all in one package. This helps the AI understand why you liked something, not just what it was.
2. The Training Process (Learning to Write)
The AI needs to learn two things: how to understand your taste and how to write a good instruction. NaviGen teaches this in two stages:
Stage 1: Supervised Fine-Tuning (The "Study Hall"):
The system uses a clever trick called Evolutionary Search. Imagine a group of writers trying to guess the perfect instruction for a user. They start with three different styles (conservative, balanced, exploratory). They mix and match their ideas (like breeding plants), and an "editor" (another AI) picks the best ones to create the next generation of instructions.
The model learns from these "evolved" instructions. It also learns to "think out loud" (Chain-of-Thought), explaining how a user's taste evolved from clicking on a funny elf video to wanting a romantic anime scene.Stage 2: Reinforcement Learning (The "Game Show"):
Once the model can write instructions, it plays a game to get better. It gets points (rewards) for:- Accuracy: Did it predict the right "behavioral code" (CID) for the next item the user might like?
- Quality: Is the instruction specific, creative, and actually possible for the image/video generator to make?
- Consistency: Does the instruction match the predicted item, and does the predicted item match the user's history?
3. The Results
The researchers tested this on three different worlds: Shopping (Products), Gaming, and Short Videos.
- Better Art: When NaviGen wrote the instructions, the resulting images and videos were more relevant to the user's specific taste, looked more aesthetic, and were more creative than instructions written by other methods.
- Better Guessing: It was also better at predicting what item a user would click on next, proving it truly understood the user's "behavioral fingerprint."
- The "Aha!" Moment: In a case study, a user's history went from "funny elf videos" to "romantic anime." Other systems just saw "anime." NaviGen saw the transition and wrote an instruction for a "romantic anime scene with a student couple," which was much closer to what the user actually wanted.
Summary
NaviGen is like a personal assistant who watches what you enjoy, figures out your hidden taste patterns, and then writes the perfect, detailed prompt for an AI artist to create content you will genuinely love. It bridges the gap between your silent clicks and the loud, detailed instructions an AI needs to create magic.
Note: The paper focuses entirely on generating images and videos based on user behavior. It does not claim to be used for medical diagnosis, clinical therapy, or any other non-creative applications.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.