← Latest papers
💬 NLP

When Does Generating More Help? Disentangling Fixed-Source Synthesis from Source Expansion in Synthetic Data Scaling

This paper isolates and analyzes Fixed-Source Synthesis (FSS) by holding seed materials and teacher models constant while varying response budgets, revealing that while FSS performance is predictable via a derived scaling law, Source Expansion ultimately outperforms FSS at large budgets, establishing FSS as a bounded but controlled axis for evaluating synthetic data protocols.

Original authors: Xu Guo, Jian Tong, Zhihui Lu, Qipeng Guo

Published 2026-07-03
📖 4 min read☕ Coffee break read

Original authors: Xu Guo, Jian Tong, Zhihui Lu, Qipeng Guo

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a student (an AI model) how to solve difficult math or physics problems. You have a very smart teacher (a large AI) and a notebook of practice questions (the "source").

The paper asks a simple but crucial question: When you want to get better results, should you buy more practice questions, or should you just have the teacher write more answers to the same few questions?

The authors call these two approaches:

  1. Source Expansion (SE): Buying a bigger notebook with more unique questions.
  2. Fixed-Source Synthesis (FSS): Keeping the same notebook but asking the teacher to write 1, 2, 4, or even 32 different answers for each single question.

Here is what they found, explained through simple analogies:

1. The "More Answers" Strategy Has a Ceiling (FSS)

If you stick to the same few questions and just ask the teacher to write more and more answers for them, the student does get smarter at first. However, this improvement hits a wall very quickly.

  • The Analogy: Imagine you are studying for a test using only three practice questions.
    • First answer: The teacher explains the solution clearly. You learn a lot.
    • Second answer: The teacher explains it a slightly different way. You learn a little more.
    • Tenth answer: The teacher is just repeating the same logic in slightly different words. You aren't learning anything new.
    • Hundredth answer: You are just memorizing the teacher's voice, not the actual math.

The paper found that after a certain point, writing more answers to the same questions gives you diminishing returns. The student hits a "ceiling" determined by how good the teacher is and how many unique questions you started with. No matter how many times you ask the teacher to re-explain the same three questions, the student will never learn the concepts hidden in the other 997 questions they didn't see.

2. The "More Questions" Strategy is Better (SE)

When the researchers compared spending their budget on more answers vs. more questions, the winner was almost always more questions.

  • The Analogy: You have a budget of $100.
    • Option A: Buy 100 copies of the same 3 practice questions and read the answers 33 times each.
    • Option B: Buy 100 different practice questions and read the answer to each one once.

The paper found that Option B (Source Expansion) makes the student much smarter. Even if you have fewer total "study sessions," seeing a wider variety of problems teaches the student how to handle new situations. Sticking to the same few problems, even with thousands of answers, leaves gaps in their knowledge.

3. "Repackaging" Doesn't Help Much

The researchers also tested fancy tricks to make the "More Answers" strategy work better. They tried:

  • Asking the teacher to pretend to be a specific character (Persona).

  • Picking the most "different" answers to ensure variety.

  • Letting the teacher make mistakes and then fixing them (Trace Repair).

  • The Analogy: It's like taking the same three practice questions and trying to make them look different by changing the font, the color of the paper, or the teacher's accent.

  • The Result: None of these tricks beat the simple method of just asking for more answers (Rejection Sampling). Once you are stuck with a small set of questions, changing how you ask for answers doesn't help much. The problem isn't the style of the answer; it's the lack of new questions.

The Big Takeaway

The paper concludes that synthetic data scaling has two different paths:

  1. The "More Answers" path (Fixed-Source): This is useful for fine-tuning and checking if a protocol works, but it is bounded. You will eventually run out of new information to learn from the same source.
  2. The "More Questions" path (Source Expansion): This is the real growth engine. If you want to build a smarter AI, you need to feed it more unique, real-world examples (new seeds), not just more variations of the same old examples.

In short: Don't just keep asking the same question over and over hoping for a better answer. Go find new questions to ask.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →