Interactor: Agentic RL oriented Iterative Creation for Ad Description Generation in Sponsored Search
This paper introduces Interactor, an agentic reinforcement learning framework that iteratively refines ad descriptions in sponsored search through multi-turn interactions with generative reward models, significantly improving knowledge richness and landing page consistency while achieving successful online deployment.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are running a lemonade stand, but instead of just shouting "Lemonade!" (which is like an ad title), you have a long, detailed menu board where you can explain why your lemonade is special.
This paper introduces a new system called INTERACTOR that acts like a super-smart, self-correcting marketing assistant to write those long, detailed descriptions for search ads.
Here is how it works, broken down into simple concepts:
1. The Problem: The "One-Shot" Mistake
Most current AI tools try to write an ad description in one single go, like a student taking a test without a scratchpad. They get it right sometimes, but often they make two big mistakes:
- They make things up: They might say your lemonade uses "organic lemons" when your supplier only sells regular ones (this is called being unfaithful to the landing page).
- They are too vague: They might just say "It tastes good" instead of explaining that the lemons are hand-picked at dawn (missing world knowledge).
Because the AI can't "see" the user's reaction immediately, it keeps making the same mistakes.
2. The Solution: The "Coach and Player" Loop
INTERACTOR changes the game. Instead of a one-shot test, it turns ad writing into a multi-turn conversation between a Player (the AI writing the ad) and a Coach (a set of specialized AI judges).
Here is the step-by-step loop:
- The Player Makes a Draft: The AI writes a first version of the ad.
- The Coach Reviews It: The Coach doesn't just give a score (like "8/10"). It acts like a strict editor with a red pen. It says:
- "You claimed 'free shipping,' but your website says 'delivery fee applies.' That's a lie. Fix it."
- "You didn't mention that black sesame helps with hair health, which is what the user is actually asking about. Add that."
- The Player Refines: The AI reads the Coach's notes, thinks about them, and rewrites the ad to fix the specific errors.
- Repeat: This happens several times until the ad is perfect.
3. The Secret Sauce: "Agentic RL"
The paper calls this Agentic Reinforcement Learning. Think of it like training a dog, but the dog is smart enough to understand why it got a treat or a correction.
- Traditional RL: The AI gets a "Good Job!" or "Bad Job!" signal. It guesses what to do next.
- INTERACTOR (Agentic RL): The AI gets a reasoning feedback. It knows exactly which sentence was wrong and why. This allows it to learn much faster and write much better descriptions.
4. The Results: From "Fluff" to "Facts"
The paper tested this on real data from Baidu's search engine.
- Before: AI ads were often "fluffy" (full of empty words) or "hallucinated" (making up facts).
- After: The INTERACTOR system produced ads that were:
- Knowledge-Rich: They actually answered the user's question with real facts.
- Faithful: They stuck strictly to what the advertiser's website actually offered.
- More Clicks: Because the ads were better, more people clicked on them, and the company made more money.
5. Real-World Impact
The paper states that since May 2026, this system has been running live on a major search engine. It isn't just a theory; it is currently helping advertisers write better ads and helping users find more relevant information, all while boosting the search engine's revenue.
In a nutshell: INTERACTOR is like giving an AI writer a team of expert editors who don't just grade the paper, but sit down and help the writer fix every single mistake until the ad is perfect.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.