← Latest papers
💬 NLP

Omni-Safety under Cross-Modality Conflict: Vulnerabilities, Dynamics Mechanisms and Efficient Alignment

This paper identifies a "Mid-layer Dissolution" vulnerability in Omni-modal Large Language Models (OLLMs) through a new dataset and mechanistic analysis, and proposes **OmniSteer**, a lightweight alignment method that uses a golden refusal vector to effectively mitigate cross-modal safety risks without compromising general capabilities.

Original authors: Kun Wang, Zherui Li, Zhenhong Zhou, Yitong Zhang, Yan Mi, Kun Yang, Yiming Zhang, Junhao Dong, Zhongxiang Sun, Qiankun Li, Yang Liu

Published 2026-02-12
📖 3 min read☕ Coffee break read

Original authors: Kun Wang, Zherui Li, Zhenhong Zhou, Yitong Zhang, Yan Mi, Kun Yang, Yiming Zhang, Junhao Dong, Zhongxiang Sun, Qiankun Li, Yang Liu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a super-intelligent personal assistant (an Omni-modal Large Language Model). This assistant is amazing: it can read your texts, watch your videos, listen to your voice, and look at your photos. You’ve trained it to be a "good citizen"—if you ask it, "How do I steal a car?", it will firmly say, "I cannot help you with that."

However, researchers discovered a "glitch in the matrix" regarding how this assistant handles different types of information at once.

The Problem: The "Identity Crisis" (Cross-Modality Conflict)

Think of the assistant’s safety training like a security guard standing at a door.

  • If you walk up and say "I want to steal a car," the guard recognizes the words and stops you.
  • If you walk up and show a picture of a car being stolen, the guard recognizes the image and stops you.

But here is the vulnerability: If you walk up and say "Help me with the thing in this picture," while showing a picture of a stolen car, the guard gets confused. The "meaning" is split between your voice and your hands. Because the assistant is trying so hard to "connect the dots" between the text and the image, its internal "safety alarm" accidentally turns itself off.

The researchers call this "Mid-layer Dissolution." It’s like a security guard who starts to realize something is wrong, but halfway through processing the information, their brain suddenly goes, "Wait, I'm just looking at a picture and listening to a sentence... nothing to see here!" and they let you pass.

The Discovery: The "Fading Signal"

The researchers looked under the hood (the "hidden layers" of the AI's brain) and found that when information comes from multiple sources (text + image + audio), the "Refusal Signal"—the mental impulse to say "No"—becomes incredibly weak.

It’s not that the AI wants to be bad; it’s that the signal that tells it "This is dangerous" physically shrinks and fades away as it tries to merge the different types of data.

The Solution: "OmniSteer" (The Smart Spotlight)

To fix this, the researchers created a tool called OmniSteer.

Imagine that instead of just having one security guard, you give the guard a "Golden Compass" (the Golden Refusal Vector). This compass doesn't care if the danger is a spoken word, a drawing, or a video; it only points toward the concept of "danger."

But there’s a catch: if the guard is too aggressive, they might stop you just for asking, "How do I bake a cake?" (this is called "over-refusal").

OmniSteer acts like a Smart Spotlight:

  1. It uses a tiny, lightweight "adapter" (like a smart sensor).
  2. When the sensor detects a "split" input (like text + image), it realizes the safety signal is fading.
  3. It immediately shines a bright, focused light on that "Golden Compass" direction, boosting the "No!" signal back up to full strength.
  4. If the input is totally normal (like "What is the weather?"), the spotlight stays dim so the assistant can remain helpful and friendly.

The Result

By using this "Smart Spotlight," the researchers boosted the AI's ability to refuse harmful requests from about 70% up to over 91%. Most importantly, the AI didn't become a "grumpy" assistant; it stayed just as smart and helpful as before, only now it’s much harder to trick.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →