OneSearch-V2: The Latent Reasoning Enhanced Self-distillation Generative Search Framework
The paper introduces OneSearch-V2, a latent reasoning-enhanced self-distillation generative search framework that improves query understanding, uncovers precise user intents, and aligns with personal preferences through three key innovations, resulting in significant offline and online performance gains in e-commerce search without increasing inference costs.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are shopping in a massive, chaotic digital department store (like Kuaishou's mall). You type a search query, and the store's "AI Shop Assistant" tries to guess what you want to buy.
The old assistant (OneSearch V1) was good, but it had three big problems:
- It was a bit literal: If you asked for "gifts for my girlfriend," it might just show you generic boxes instead of thinking, "Oh, she likes flowers, but maybe not roses because of her allergy."
- It relied too much on the past: It only remembered what everyone else bought, not what you specifically wanted right now.
- It got stuck in a loop: It kept showing you the same popular items, making it hard to find new or niche products (the "information bubble").
OneSearch-V2 is the new, upgraded assistant. It doesn't just guess; it thinks before it speaks, but it does so in a way that is incredibly fast and doesn't slow down the store.
Here is how it works, using simple analogies:
1. The "Brainstorming" Phase (Thought-Augmented Query Understanding)
The Problem: When you type a vague query like "something for a rainy day," the old AI just looked for umbrellas. It missed raincoats, hot chocolate, or cozy blankets.
The V2 Solution: Before the AI gives you an answer, it runs a quick, invisible "brainstorming session" in the background.
- The Analogy: Imagine a chef who, before cooking, writes a quick recipe card in their head: "User wants 'rainy day' -> implies: staying dry (umbrella), staying warm (sweater), or comfort (soup)."
- The Magic: In the old system, the AI had to say this reasoning out loud (which took too long). In V2, the AI uses a special "Keyword CoT" (Chain of Thought). It extracts just the keywords from that brainstorming session (e.g., "umbrella," "sweater," "soup") and uses them to sharpen its focus. It's like the chef reading their mental recipe card instantly without writing it down.
2. The "Shadow Training" Phase (Reasoning-Internalized Self-Distillation)
The Problem: Usually, to teach an AI to think deeply, you have to make it write out long explanations. But writing takes time, and in a busy store, you can't wait 5 seconds for the AI to think.
The V2 Solution: The system uses a technique called Self-Distillation.
- The Analogy: Imagine a master chef (the "Teacher") and their apprentice (the "Student").
- Old Way: The master chef writes a 10-page essay on how to make a dish, and the student reads it every time. This is slow.
- V2 Way: The master chef cooks a dish while reading their secret recipe notes. The apprentice watches the master cook without the notes. The apprentice tries to copy the master's movements perfectly.
- The Result: After practicing this way, the apprentice (the model) learns the intuition of the recipe. Now, when the apprentice cooks alone, they don't need the notes anymore. They just "know" what to do instantly.
- Why it matters: The AI has "internalized" the reasoning. It doesn't need to stop and think; the thinking is now part of its muscle memory. This means it's smart and fast.
3. The "Real-Time Feedback" Phase (Behavior Preference Alignment)
The Problem: The old system learned from a static report of what people bought last month. If a new trend started today, the system wouldn't know until next month. Also, it might get "hacked" by showing only the most popular items, ignoring what you actually clicked on.
The V2 Solution: It stops waiting for monthly reports and listens to real-time whispers.
- The Analogy: Imagine a waiter who used to only read the "Best Sellers" menu from last year. Now, the waiter watches your table while you eat.
- If you glance at the dessert menu, they note it.
- If you order a coffee, they remember you like coffee.
- If you ignore the expensive steak but buy the soup, they learn that you prefer soup, even if the steak is the "best seller."
- The Magic: The system adjusts its recommendations instantly based on your actual clicks and purchases, balancing what is popular with what you specifically want. It also uses a "hierarchical" approach: it first guesses the category (e.g., "Shoes"), then the style (e.g., "Running"), then the brand. If it gets the first step wrong, it stops wasting time on the rest, saving energy and improving accuracy.
The Final Result
Because of these three upgrades, OneSearch-V2 is like a shop assistant who:
- Understands context (knows "rainy day" means more than just umbrellas).
- Thinks instantly (has memorized the logic without needing to write it down).
- Adapts to you (learns your specific taste in real-time).
The Outcome:
- You find what you want faster.
- You discover new, cool items you didn't know existed (breaking the "bubble").
- The store sells more because the suggestions are actually helpful.
- Best of all: It does all this without making the website slow or laggy. It's a smarter brain that runs just as fast as the old one.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.