When Domains Interact: Asymmetric and Order-Sensitive Cross-Domain Effects in Reinforcement Learning for Reasoning
This paper provides the first systematic analysis of Group Relative Policy Optimization (GRPO) across multiple reasoning domains, revealing that training performance is characterized by pronounced asymmetry, order sensitivity, and strategy dependence, thereby necessitating domain-aware and order-aware training designs.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are training a brilliant but somewhat rigid student to solve different types of puzzles: Math, Science, Logic, and Brain Teasers. You have a special teaching method called GRPO (Group Relative Policy Optimization), which is like a coach who gives the student feedback after they try to solve a problem, helping them get better over time.
This paper asks a simple question: Does it matter which subject you teach the student first, and does it matter if you teach them one subject at a time or mix them all together?
The researchers found that the answer is a resounding "Yes!" and the results are surprisingly lopsided. Here is what they discovered, explained through everyday analogies:
1. The "Math Magnet" Effect (Asymmetry)
Think of Math as a superpower that is very easy to catch.
- The Finding: If you train the student on Science, Logic, or even Brain Teasers, their Math skills magically get better (by about 25%).
- The Catch: It doesn't work the other way around. If you train them on Math, it barely helps them get better at Logic or Brain Teasers.
- The Analogy: Imagine Math is a "universal solvent." Pouring Math training into the student's brain dissolves barriers and helps them solve Math problems, even if they were originally studying something else. But pouring Logic training into the brain doesn't dissolve barriers for Logic; it just stays in the Logic bucket.
2. The "Order of Operations" Trap (Order Sensitivity)
The sequence in which you teach the subjects changes the outcome dramatically. It's like baking a cake: if you mix the ingredients in the wrong order, the cake collapses.
The Math & Science Pair:
- Good Order (Math → Science): If you teach Math first, then Science, the student becomes a genius at both. They score 83% on Math and 41% on Science.
- Bad Order (Science → Math): If you teach Science first, then Math, the student gets confused. Their Math score drops, and their Science score crashes to 25%.
- The Analogy: Think of Math as a strong foundation. If you build the Science house on top of a solid Math foundation, it stands tall. If you try to build the Math house on top of a shaky Science foundation, both structures wobble and fall.
The Logic & Science Clash:
- The Finding: Science and Logic hate each other. No matter which one you teach first, if you try to teach the other one afterward, the student forgets the first one.
- The Analogy: It's like trying to learn two different languages that use completely opposite grammar rules. If you learn French (Science) then try to learn a made-up alien language (Logic), the French rules get scrambled.
- The Solution: The only way to save both is to mix them. Instead of teaching French then Alien, you teach a little bit of French and a little bit of Alien in the same lesson. This "mixed training" stops the confusion and helps the student learn both.
3. The "One-Size-Fits-All" Myth
The paper proves there is no single "perfect schedule" for teaching all subjects.
- For Math: A strict schedule (Sequential training) works best. You should teach subjects one by one in a specific order.
- For Science and Logic: A mixed schedule works best. You should throw all the data into a blender and teach them together.
- The Danger: If you pick the wrong schedule (like teaching Science before Math), you could lose a huge chunk of performance—dropping from a 70% average score down to 56%. That's the difference between an "A" student and a "C" student.
Summary
The paper concludes that when training AI to reason:
- Math is special: It helps other subjects, but other subjects don't help it much.
- Timing is everything: Teaching Math before Science is a winning strategy; doing it the other way is a disaster.
- Mix and Match: For some subjects (like Science and Logic), you must mix the training data to avoid them fighting each other.
The researchers built a "training map" to show exactly which order works for which subject, warning that ignoring these rules can lead to AI models that are much worse at reasoning than they could be.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.