Multimodal Model Diffing for Feature Discovery and Control
This paper introduces MMDiff, a multimodal model-diffing framework that utilizes sparse autoencoders to isolate, detect, and causally control specific features in Multimodal Large Language Models, thereby enabling targeted improvements in spatial understanding, OCR, and safety while preserving general performance.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you've built a super-smart robot that can read books, look at pictures, and answer questions about both. You've taught it to be helpful, but sometimes it gets confused, hallucinates facts, or even says something unsafe when shown a tricky image. The problem is, inside this robot's "brain," everything is a tangled mess of math. It's like a giant library where all the books are glued together, and you can't tell which specific page is responsible for the robot's bad behavior or its amazing ability to spot a cat in a photo. Scientists want to fix this, but they need a way to peek inside and find the exact "switches" or "dials" that control these specific skills.
To do this, researchers use a tool called a Sparse Autoencoder (SAE). Think of an SAE as a magical translator that takes the robot's messy, tangled thoughts and breaks them down into a list of simple, distinct "ideas" or "features." Instead of a blurry cloud of data, the robot's brain is suddenly organized into a dictionary where each entry is a single concept, like "red," "danger," or "left." But here's the catch: when you teach a text-only robot to see images, it doesn't just add new features; it often reshuffles the old ones. It's like taking a dictionary of English words and suddenly using the word "apple" to mean "a red car." This makes it hard to know which features are actually doing the new visual work and which ones are just leftovers from the text training.
This is where the new paper, MMDiff, steps in. The researchers, Hunar Batra and their team from Oxford and Microsoft, realized that to find the specific "visual switches," you can't just look at the final robot; you have to compare it to its "text-only" self before it learned to see. They created a method called Multimodal Model Diffing. Imagine taking two identical twins: one who has only read books, and one who has also learned to see the world. By comparing their brains side-by-side, MMDiff highlights exactly which parts of the brain have changed or been repurposed for vision.
The team found that this "diffing" process is incredibly powerful. They discovered that by isolating these changed features, they could actually control the robot's behavior with surprising precision. For instance, they found specific features responsible for understanding spatial relationships (like knowing if a cup is above a table). When they turned these specific features "off," the robot's ability to answer spatial questions dropped by about 12%, but its ability to answer general questions stayed exactly the same. It was like turning off the "left/right" switch without breaking the "what is this?" switch.
They also found features linked to safety. When they identified and removed the features that made the robot vulnerable to unsafe image-based tricks, the robot's success rate at ignoring those tricks improved by 24%. Even better, they could "steer" the robot by turning these features up or down. By boosting the right features, they improved the robot's ability to read text in images (OCR) by 1.8% and its spatial reasoning by 3.6%, all without needing to retrain the whole robot from scratch.
The paper explicitly rules out the idea that you can just train a new dictionary from scratch on the multimodal robot and expect to find these specific visual features easily. They showed that without comparing the robot to its text-only version (the "diffing" part), the features get mixed up, and you can't isolate the specific skills you want to control. They also found that simply looking at which features fire often isn't enough; you have to check if those features actually rotated or changed their meaning during the training process.
In short, MMDiff acts like a high-tech "find and replace" tool for AI brains. It doesn't just tell us that the robot is doing something; it tells us exactly which tiny part of the brain is doing it, allowing us to tweak, fix, or improve specific behaviors like reading text, understanding space, or staying safe, all while leaving the rest of the robot's intelligence untouched. This suggests that in the future, we might be able to audit and control these powerful AI systems much more safely and effectively, ensuring they do exactly what we want them to do.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.