← Latest papers
💻 computer science

Asymmetric Collapse in Model Merging: When Refusal Over- writes Recognition

This paper demonstrates that standard model merging techniques often cause an asymmetric collapse where safety recognition capabilities are lost while broad refusal behaviors are preserved, primarily because the refusal fine-tune induces significantly larger task-vector magnitudes that dominate the merging process despite the vectors being nearly orthogonal.

Original authors: Aarnav Choudhary, Matheus Fonseca Rocha, Jiwon Seo, Vasu Sharma, Maheep Chaudhary

Published 2026-07-31
📖 4 min read☕ Coffee break read

Original authors: Aarnav Choudhary, Matheus Fonseca Rocha, Jiwon Seo, Vasu Sharma, Maheep Chaudhary

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a super-smart robot that can learn new tricks. Sometimes, you want to teach it one specific thing, like how to spot a fake news headline. Other times, you want it to learn a different trick, like how to say "no" to dangerous requests. But what if you wanted the robot to do both at the same time? In the world of artificial intelligence, there's a clever shortcut called "model merging." Instead of retraining the robot from scratch, scientists take two different versions of the robot—one trained for trick A and one for trick B—and mathematically blend their brains together. It's like mixing two different flavors of ice cream to get a new, perfect swirl. The big question is: when you mix these two "flavors" of safety training, do you get a robot that's good at both, or does one flavor completely drown out the other? This is the mystery a new study set out to solve, looking at whether we can keep a robot's ability to understand danger while also keeping its ability to refuse it.

The researchers behind this study, published at the COLM 2026 conference, decided to test this with a specific robot model called Gemma-3-1B-IT. They created two special versions of this robot. The first version, let's call it "Detective," was trained on a massive dataset of medical questions to become an expert at spotting how harmful a prompt is. It learned to give a graded answer, like saying, "This is a little risky," or "This is very dangerous." The second version, "Guardian," was trained on a huge collection of tricky, adversarial prompts designed to break the robot's rules. Guardian learned one simple, blunt lesson: if it looks dangerous, just say "No" and refuse to answer, no questions asked.

The team then took these two specialists and tried to merge them back into one robot using four different mathematical blending methods. They expected the result to be a robot that could both detect the level of harm and refuse the bad stuff. But instead, they found a strange and lopsided outcome. In every single case, the "Guardian" behavior took over completely. The merged robots became excellent at refusing bad prompts, keeping their refusal rates high (between 81% and 85%). However, the "Detective" skills vanished almost entirely. The robots forgot how to grade the harm; their ability to classify the severity of a prompt dropped to near zero (falling to as low as 0.6% or 12.9%, depending on the method). It was as if the robot learned to slam the door shut on everything, forgetting how to peek through the peephole to see what was actually there.

Why did this happen? The scientists looked closely at the "DNA" of the two robots to find the culprit. They discovered it wasn't because the two robots were fighting each other or trying to go in opposite directions. In fact, their brains were nearly at right angles to each other, meaning they were trying to do totally different things without clashing. The real problem was a matter of volume. The "Guardian" robot had been trained on a dataset that was roughly 29 times larger than the "Detective's" dataset. Because of this, the changes made to the Guardian's brain were much "louder" and stronger. When the scientists mixed the brains, the mathematical tools they used were like a volume knob that turned up the louder signal and turned down the quieter one. The "Guardian's" big, bold "No!" drowned out the "Detective's" subtle, nuanced "Maybe."

Even when the researchers tried to fix this by shrinking the Guardian's training data to match the Detective's size, the problem didn't fully go away for all the merging methods. This suggests that while the size of the data was a major factor, the way the merging algorithms handle these "loud" versus "quiet" updates is the real key. The study suggests that standard merging tools are currently biased toward keeping the big, blunt behaviors (like refusing everything) while accidentally deleting the fine-grained, detailed skills (like understanding why something is bad).

So, what does this mean for the future? The paper suggests that if we want to build safe AI that can both refuse bad things and understand the nuances of safety, we can't just mix and match models using current standard tools. We might be losing the "Detective" inside the "Guardian." The researchers conclude that future methods need to be smarter about balancing these different "volumes" of learning, ensuring that the robot doesn't just learn to say "no" to everything, but also learns to understand the difference between a little warning and a big danger.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →