← Latest papers
💻 computer science

AdsWorldEngine: A Self-Evolving Conversational Advertising Agent through Orchestrator and Tool Coevolution

AdsWorldEngine is a self-evolving conversational advertising framework that employs an iterative actor-tool coevolution procedure and label-grounded judgment modeling to dynamically infer commercial intent, optimize ad delivery tools, and significantly improve relevance, diversity, and revenue metrics in multi-turn assistant interactions.

Original authors: Simiao Zuo, Chenhui Xu, Yimeng Jia, Qiang Lou, Jian Jiao, Denis Charles

Published 2026-08-17
📖 5 min read🧠 Deep dive

Original authors: Simiao Zuo, Chenhui Xu, Yimeng Jia, Qiang Lou, Jian Jiao, Denis Charles

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are chatting with a super-smart digital friend who helps you plan your day, find a restaurant, or pick out a new video game. In the old days of the internet, if you typed "best pizza near me" into a search box, a list of ads would pop up immediately. That was easy because your request was short and clear. But now, conversations are different. You might say, "I'm thinking of a beach trip," and your friend suggests a few places. Then you say, "Not in the Caribbean," and they refine the list. Finally, you might ask, "Anything cheaper?" In this back-and-forth dance, figuring out when to show an ad and what ad to show is like trying to catch a fish while it's swimming in a dark, moving river. You can't just throw a net; you have to understand the whole story, guess what the person actually wants, and make sure the ad feels helpful, not annoying. This is the challenge of "conversational advertising," a field trying to make sure the right message appears at the right moment in a chat without ruining the vibe.

Enter AdsWorldEngine, a new system designed by researchers at Microsoft to solve this tricky puzzle. Think of it as a highly trained, self-improving team of digital agents working together to run a smart ad shop inside a conversation. Instead of just following a rigid script, this system has three main characters: a Gatekeeper, a Conductor, and a Judge.

The Gatekeeper is the first to speak. Before any ads are even considered, it looks at the conversation and asks, "Is this a good time to show an ad?" It's like a bouncer at a club who decides if the music is right for a commercial break. If the user is just asking for facts or the conversation is too sensitive, the Gatekeeper says "No" and keeps the ad away to protect the user's experience.

If the Gatekeeper gives the green light, the Conductor takes over. This is the main brain of the operation. It listens to the whole chat history, figures out what the user really wants (even if they haven't said it directly), and then calls upon a set of Tools—like a librarian, a shopper, and a sorter—to find the best three ads. The Conductor doesn't just pick the first three it finds; it checks them, makes sure they are different from each other, and ensures they actually fit the user's current mood.

Here is the really cool part: the system doesn't just stop there. It has a Judge that watches the whole show after it happens. The Judge scores how good the ad selection was. If the ads were perfect, the Conductor gets a high score. If the ads were boring or annoying, it gets a low score. But the magic of AdsWorldEngine is that it uses these scores to teach everyone how to get better. It's like a video game where the player (the Conductor) and the level designers (the Tools) learn from each other. When the Conductor does something great, the system saves that moment to show the Tools, "Hey, this is how you should work!" And when the Tools do a great job, the Conductor learns, "This is how I should ask for help!" They evolve together, getting smarter with every conversation.

The researchers found that this "self-evolving" approach works incredibly well. By training the system to learn from its own successes and failures, they saw massive improvements. In their tests, the system made the ads 60% more diverse (meaning you saw a wider variety of things, not just the same thing over and over) and 80% more relevant (the ads actually matched what the user was talking about). When they tried it out in the real world with Microsoft Copilot, a popular AI assistant, the results were even more impressive: the system increased the value of the ads by 22% and showed ads in 74% more conversations where they were actually helpful.

What makes this different from other attempts is that it doesn't just rely on a human telling the computer what to do, or a simple rule like "if the user says 'buy', show an ad." The researchers argue that simple rules and basic training aren't enough because conversations are too complex. Instead, they built a loop where the system practices, gets graded, and then rewrites its own rules to improve. They also invented a special way to teach the system how to make "subjective" decisions—like knowing when an ad would be annoying—by using human labels as a guide and filtering out bad reasoning.

In short, AdsWorldEngine is a step forward in making AI assistants feel more natural and less like a billboard. It suggests that the future of advertising isn't about interrupting you, but about being a helpful part of the conversation, learning from every interaction to get the timing and the content just right. The paper shows that when you let the "actor" (the one who speaks) and the "tools" (the ones who fetch the data) learn together, the whole system becomes much smarter and more effective than if they were trained separately.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →