Privileged Likelihood Is Not Automatically Value: Three Checks for Token Credit in On-Policy Self-Distillation
This paper argues that privileged token likelihood changes in on-policy self-distillation do not automatically constitute valid outcome credit, demonstrating through formal analysis and empirical experiments that such signals often fail to track better actions or reinforce desired behaviors compared to outcome-only controls.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a robot to solve a tricky math puzzle. You show it the problem, and it starts typing out its thoughts, step by step. At the very end, you can easily tell if the robot got the right answer: it's either correct or it's wrong. But what about the steps in the middle? Did the robot make a brilliant move at step three, or did it stumble at step five? In the world of Artificial Intelligence, this is a huge mystery. We know the final result, but we don't know which specific words (or "tokens") the robot wrote were the heroes and which were the villains.
To fix this, researchers have been trying a clever trick called "privileged self-distillation." Think of it like a student taking a test, then immediately getting a reference sheet that explains exactly what they should have done. The student then looks at their own answer, reads the reference sheet, and tries to figure out which words they wrote were good and which were bad. The hope is that by rewarding the "good" words and punishing the "bad" ones, the robot will get smarter faster. But here is the catch: just because a word looks good after you know the answer, doesn't mean it actually helped you get there. It might just be a word that happens to sound nice when you're looking back with hindsight.
This paper from Salesforce AI Research is like a group of skeptical detectives showing up to the classroom to check if this "reference sheet" method is actually working or if it's just a magic trick that fools the teacher. They set up a rigorous experiment to see if the robot is actually learning the right lessons or just memorizing the wrong ones.
The Great "Hindsight" Heist
The researchers decided to put the "privileged self-distillation" method to the ultimate test using a 20-billion-parameter AI model (a very smart robot) and a set of hard math problems called AIME 2025. Their goal was to answer three specific questions that determine if this teaching method is actually useful or just a waste of time.
Question 1: Does the "Good Word" detector actually find the right words?
Imagine a detective trying to find the culprit. If the detective points a finger at a suspect, does that suspect actually have a motive? The researchers checked if the AI's "score" (the number it gives to a word to say how good it was) actually matched up with correct answers. They found that the score was basically guessing. When they looked at thousands of math attempts, the score was only slightly better than a coin flip at telling a correct solution from an incorrect one. In fact, after adjusting for how long the answers were, the score actually liked the wrong answers slightly more than the right ones! It's like a teacher who, after grading a test, gives a gold star to the student who got everything wrong because their handwriting was pretty.
Question 2: Did the robot write its own report card?
This is the sneaky part. In many experiments, the AI generates an answer, and then the same AI (or a version of it) writes a critique of that answer. The researchers worried this was a loop of self-flattery. If you write your own report card, you might be too nice to yourself. To fix this, they tried a "cross-fitting" method: they made the AI write a critique for another AI's answer to the same problem. This breaks the loop. But even with this fix, the scores still didn't reliably predict which answers were correct. The feedback was just as confused as before. It turns out that just because you can write a nice story about a solution doesn't mean the solution is actually good.
Question 3: What happens when the robot tries to learn from these scores?
Finally, they asked: even if the scores were weird, what happens when the robot actually tries to use them to improve? They trained the robot using five different versions of this "token credit" method and compared them to a robot that only learned from the final right-or-wrong answer. The results were a massive disappointment for the "token credit" fans. Every single robot that tried to use the fancy token scores ended up performing worse than the robot that just stuck to the simple right-or-wrong feedback.
- The simple robot got about 64.2% of the answers right.
- The fancy robots using token scores only got between 24.2% and 33.9% right.
It's as if the robot that tried to analyze every single step of its thinking got confused and forgot how to solve the problem entirely, while the robot that just looked at the final result kept getting better.
The Verdict
The paper concludes that a "token likelihood score"—a number that says a word is more likely to be right because of a reference sheet—is not automatically valuable. Just because a word becomes more probable after you see the answer doesn't mean that word was the key to the solution.
The researchers found that:
- The scores didn't track success; they were basically random noise.
- Whether the feedback came from the same answer or a different one didn't magically fix the problem.
- Using these scores to train the model actually made it worse, pushing it away from the correct answers.
In short, the paper argues that we cannot simply assume that "what looks good in hindsight" is "what is actually helpful." Before we start giving AI models credit for every word they type, we need to check three things: does the score actually match the outcome? Is the feedback fair and independent? And does the training actually make the model smarter? In this specific experiment, the answer to all three was a resounding "no." The fancy method of giving credit to individual words failed, and the old-school method of just checking the final answer was the clear winner.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.