← Latest papers
💻 computer science

SafeNexus: Discovering and Steering Modality-Universal Safety Neurons in MLLMs

The paper introduces SafeNexus, a cross-modal safety alignment framework that identifies and steers modality-universal safety neurons through targeted suppression and activation amplification to robustly defend multimodal large language models against diverse cross-modal threats while preserving utility.

Original authors: Jian Yu, Fei Shen, Cong Wang, Jian Wang, Lu Jin. Xiaoyu Du, Jinhui Tang, Tat-Seng Chua

Published 2026-08-03
📖 4 min read☕ Coffee break read

Original authors: Jian Yu, Fei Shen, Cong Wang, Jian Wang, Lu Jin. Xiaoyu Du, Jinhui Tang, Tat-Seng Chua

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the internet as a giant, bustling library where books can talk back to you. For a long time, these talking books were just text, like a very smart chatbot. But recently, they've evolved into "Multimodal Large Language Models" (MLLMs). Think of these as super-librarians who can now read text, look at pictures, and listen to audio all at once. They are incredibly helpful, but this new superpower comes with a catch: it opens up new ways for troublemakers to trick them. Just as a security guard at a museum might know how to spot a fake painting but fail to notice a fake sound recording, these AI models often have safety guards that only work for one type of input. If you try to trick them with a picture instead of words, or a sound clip instead of a sentence, the guards might miss the danger entirely. The big question researchers are asking is: How do we build a safety system that works no matter how the troublemaker tries to sneak in?

This is where a new study called SafeNexus steps in. The researchers, a team from universities in China and Singapore, decided to stop looking at the AI's safety from the outside and instead went on a treasure hunt inside the AI's "brain." They discovered that the AI's ability to say "no" to bad requests isn't scattered randomly; it's actually controlled by a tiny, special group of internal switches called neurons.

Here's the fun part: The team found that while different types of inputs (like text, images, or audio) use different parts of the brain to process information, there is a secret, shared club of neurons that acts as the ultimate "danger detector" for all of them. They call these the US-Neurons (Modality-Universal Safety Neurons). It's like finding out that whether you are trying to sneak a knife, a bomb, or a fake ID past security, there is one specific, tiny security camera in the hallway that spots all of them. If you turn that camera off, the whole building becomes unsafe, no matter what the intruder is carrying.

The paper suggests that previous methods tried to patch safety holes by retraining the whole AI or adding external filters, which is like hiring a new security guard for every single door. SafeNexus takes a different approach: it finds those few critical "danger detector" neurons and gives them a superpower boost. The team tested two ways to do this:

  1. The Amplifier: During the AI's thinking process, they simply turned up the volume on those specific neurons, making them shout "DANGER!" louder and faster.
  2. The Calibrator: They gave those specific neurons a tiny, targeted tune-up (using a method called LoRA) so they became naturally better at spotting trouble, without needing to retrain the entire AI.

The results were impressive. When they tested this on models like Qwen and VITA, they found that boosting these few neurons made the AI much better at refusing harmful requests, whether the request came as text, a picture, or a sound. Crucially, they found that this didn't make the AI "dumb" or overly cautious. It didn't start refusing harmless questions (like "how do I bake a cake?"), and it didn't forget how to do its other jobs. In fact, in some tests, the new method was far better at stopping attacks than the current best methods, all while changing less than 0.05% of the AI's total brain power.

The researchers also showed that this trick works even on things they didn't explicitly train on, like video. Even though they never showed the AI video clips during their "tuning" phase, the boosted neurons were still good at spotting danger in video attacks. This suggests they really found a core, universal safety mechanism rather than just memorizing specific examples.

In short, SafeNexus suggests that we don't need to rebuild the AI's entire safety system to make it safe across all formats. Instead, we just need to find the few tiny, universal "brakes" inside the machine and make sure they are working at full strength. It's a bit like realizing that to stop a car from speeding, you don't need to replace the whole engine; you just need to make sure the brake pedal is connected firmly to the wheels.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →