Supplement Generation Training for Enhancing Agentic Task Performance
This paper proposes Supplement Generation Training (SGT), a cost-effective strategy that trains a smaller language model to generate dynamic supplemental text for large foundation models, thereby enhancing agentic task performance without the need for expensive post-training or model modification.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Problem: The "Overworked Genius"
Imagine you have a brilliant, world-class expert (let's call them The Executive). This person knows everything, can solve complex math problems, write code, and answer deep philosophical questions. However, there are two big catches:
- They are expensive: You can't hire them to work on every single tiny problem.
- They are locked up: You can't teach them new tricks or retrain them for every new job because they are a "closed-source" system (like a black box).
Currently, if you want to use this Executive for a new task, you have to write them a perfect set of instructions (a prompt). But writing the perfect instruction is hard, and if the task changes, you have to start over.
The Solution: The "Super Assistant"
The authors of this paper propose a clever workaround. Instead of trying to retrain the expensive Executive, they train a small, cheap, and fast Assistant (a smaller AI model).
The Assistant's job isn't to solve the problem itself. Its job is to prepare the Executive for the job.
Think of it like this:
- The Old Way: You hand a complex legal case directly to the Lawyer (The Executive) and say, "Solve this." The Lawyer might miss a detail because they didn't have the right context.
- The New Way (SGT): You hand the case to your Legal Assistant first. The Assistant reads the file, realizes the Lawyer needs a specific summary of the evidence, a warning about common mistakes, or a step-by-step plan. The Assistant writes these notes down, attaches them to the file, and then hands the whole package to the Lawyer.
The Lawyer solves the case much better because they were given the perfect "cheat sheet" tailored specifically for that one case.
How Does the Assistant Learn? (The Training)
The paper introduces a method called Supplement Generation Training (SGT). Here is how they teach the Assistant:
- The "Warm-Up" (SFT): First, they show the Assistant examples of different types of notes it could write (e.g., "Here is a summary," "Here is a list of mistakes to avoid," "Here is a step-by-step plan"). This teaches the Assistant the format of being helpful.
- The "Trial and Error" (DPO): This is the magic part. The Assistant tries to write notes for a problem. The Executive tries to solve it.
- If the Executive solves it successfully, the Assistant gets a "Gold Star" (a reward).
- If the Executive fails, the Assistant gets a "Red Flag."
- Over time, the Assistant learns: "Oh, when the problem is about coding, writing a 'list of mistakes' works best. When it's about math, a 'step-by-step plan' works best."
The Assistant learns to dynamically choose the right type of note for the specific problem, rather than using the same note for everything.
The Results: Why It Matters
The researchers tested this on five different difficult tasks (like writing database queries, coding, and answering complex trivia).
- The Result: By using this "Super Assistant" to prep the "Executive," the system got 21% better at solving problems compared to just asking the Executive directly.
- The Efficiency: They trained a tiny model (1.7 billion parameters) to do this prep work. This is much cheaper and faster than trying to retrain the giant model (which might have hundreds of billions of parameters).
The "Search and Focus" Strategy
One of the coolest findings was how the Assistant learned.
- At first: The Assistant tried everything. It wrote summaries, lists, plans, and rephrased questions randomly. It was like a chef tasting every spice.
- Later: The Assistant "focused." It realized that for coding tasks, "Mistake Warnings" were the best spice. For logic puzzles, "Comparing Right vs. Wrong answers" was the best spice.
It stopped guessing and started being a specialized expert at preparing the Executive.
Summary
Supplement Generation Training (SGT) is a way to make big, expensive AI models smarter without changing them. You train a small, cheap AI to act as a "context coach." This coach reads the user's question, figures out exactly what extra help the big AI needs (like a summary, a warning, or a plan), writes that down, and passes it along. The big AI then uses that extra help to give a much better answer.
It's the difference between sending a general into battle with a vague map versus sending them with a detailed, up-to-date tactical briefing written by a specialized intelligence officer.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.