Procedural-skill SFT across capacity tiers: A W-Shaped pre-SFT Trajectory and Regime-Asymmetric Mechanism on 0.8B-4B Qwen3.5 Models
This paper demonstrates that while procedural-skill SFT yields uniform absolute gains across 0.8B–4B Qwen3.5 models, the resulting performance trajectory follows a W-shaped pattern driven by regime-asymmetric base capabilities, where SFT provides the most significant absolute improvement to models that initially struggle with the procedure.
Original paper dedicated to the public domain under CC0 1.0 (http://creativecommons.org/publicdomain/zero/1.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Picture: Teaching a Recipe to Different Cooks
Imagine you have three cooks of different experience levels:
- The Novice (0.8B): A kitchen helper who is eager but easily overwhelmed.
- The Apprentice (2B): A cook who knows the basics but sometimes gets lazy or skips steps.
- The Sous-Chef (4B): A very competent cook who can usually figure things out on their own.
The researchers wanted to see what happens if they give all three cooks a specific, written 5-step recipe (a "procedural skill") to follow for 200 different cooking challenges. They wanted to know: Does giving them the recipe actually help them cook better, or do they just pretend to follow it?
The Main Discovery: The "W" Shape
The researchers found something surprising. Before they gave anyone the recipe, the cooks' ability to handle the task followed a weird "W" shape when you looked at them from smallest to largest:
- The Novice (0.8B): When given the recipe before training, they actually did worse. The 5 steps were too complicated; they got confused at step 1 and messed up the whole dish. The recipe was a burden.
- The Apprentice (2B): When given the recipe, they did better. They were in the "sweet spot." The recipe acted like training wheels, helping them check their work and avoid mistakes they usually made.
- The Sous-Chef (4B): When given the recipe before training, they did slightly worse (or stayed the same). They were already so good at cooking that reading the recipe felt like extra paperwork. It slowed them down without adding value.
The Twist: After the researchers trained the cooks using the recipe (a process called SFT), something magical happened. The amount of improvement was almost the same for all three cooks.
- The Novice learned to follow the steps without getting confused.
- The Apprentice got a little sharper.
- The Sous-Chef learned to stop treating the recipe as annoying paperwork and start using it as a helpful tool.
The Lesson: The training (SFT) works hardest where the cook is struggling the most. It fixes the Novice's confusion and the Sous-Chef's laziness, resulting in a similar "boost" for everyone, even though they started from very different places.
The "Format" Trap (A Major Mistake They Fixed)
The paper highlights a huge mistake they almost made.
Imagine a strict judge who only accepts answers written on a specific line: "ANSWER: [The Dish Name]".
- The Novice and Sous-Chef often wrote the answer in a sentence like, "I conclude that the dish is Pizza."
- The strict judge (a computer script) couldn't find the word "ANSWER" and marked them as failures, even though they got the right answer.
- The Apprentice happened to write in the exact format the judge wanted, so they looked like the winner.
The researchers realized their "judge" was biased. It wasn't measuring cooking skill; it was measuring typing skill.
The Fix: They switched to a human-like judge (an AI that reads the whole sentence).
- Result: The Novice and Sous-Chef's scores jumped up!
- The Bias: The strict judge had been unfairly punishing the cooks who wrote in "free-form" sentences. The researchers found that the "curated" (recipe-following) group was actually penalized more by the strict judge because their detailed explanations often buried the final answer in the text.
The "Ceiling" and The "Frontier"
The researchers compared their trained Sous-Chef (4B) to a World-Class Celebrity Chef (called "Haiku-4-5").
- Before training, the Sous-Chef was good, but not quite at the Celebrity Chef's level.
- After training with the recipe, the Sous-Chef tied the Celebrity Chef. They both got 197 out of 200 dishes perfect.
- This proves that the training didn't just make the Sous-Chef "look" like they were following steps; it actually made them as good as the best in the world for these specific tasks.
What Didn't Work (The "Negative" Results)
The researchers tried to make the Novice (0.8B) better by changing the training recipe (changing the text, removing parts of the steps, etc.).
- Result: Nothing worked. The Novice still couldn't cook the dishes perfectly.
- Why: The Novice simply didn't have enough brainpower (capacity) to hold all the steps in their head at once. No matter how they tried to teach them, the "ceiling" was the Novice's own limits, not the teaching method.
Summary of Key Takeaways
- Training helps everyone, but for different reasons: It rescues the confused, sharpens the competent, and re-engages the over-confident. The "boost" is surprisingly similar across the board.
- Don't trust rigid judges: If you only check for specific formatting (like "ANSWER: X"), you might miss the fact that a model actually solved the problem correctly.
- Size matters, but not how you think: A tiny model can learn to use a procedure just as well as a big one, but it can't necessarily execute the whole task perfectly if the task is too hard for its brain size.
- The "W" Curve: Small models get confused by procedures; medium models love them; big models find them annoying until they are trained to see the value.
In short: The paper shows that teaching a step-by-step process is a powerful tool, but you have to measure it fairly (ignoring rigid formatting rules) and understand that a model's size determines how well it can execute the steps, even if it learns to follow them just as well as a bigger model.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.