← Latest papers
🤖 AI

Xuanwu: Evolving General Multimodal Models into an Industrial-Grade Foundation for Content Ecosystems

The paper introduces Xuanwu VL-2B, a 2B-parameter industrial-grade multimodal foundation model that leverages a compact InternViT-300M and Qwen3 1.7B architecture along with a three-stage training pipeline to achieve superior performance in content moderation and adversarial scenarios while balancing fine-grained visual perception, general capability retention, and deployment costs.

Original authors: Zhiqian Zhang, Xu Zhao, Xiaoqing Xu, Guangdong Liang, Weijia Wang, Xiaolei Lv, Bo Li, Jun Gao

Published 2026-04-01
📖 5 min read🧠 Deep dive

Original authors: Zhiqian Zhang, Xu Zhao, Xiaoqing Xu, Guangdong Liang, Weijia Wang, Xiaolei Lv, Bo Li, Jun Gao

Original paper dedicated to the public domain under CC0 1.0 (http://creativecommons.org/publicdomain/zero/1.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a brilliant, well-read librarian named Xuanwu.

In the world of Artificial Intelligence, most "general" librarians are like the ones in a massive, public library. They know everything about history, science, and literature. They can answer almost any question you ask. However, if you ask them to spot a specific, tiny, handwritten note hidden inside a complex painting, or to find a secret code disguised as a water stain, they often miss it. They are too busy looking at the big picture.

Furthermore, if you try to teach this librarian a new, very specific rule about "forbidden books" in a tiny, specialized branch of the library, they might accidentally forget how to read Shakespeare or do math. This is called "catastrophic forgetting."

The Problem:
Social media platforms are like giant, chaotic city squares. Bad actors (spammers, scammers, and rule-breakers) are constantly trying to sneak in illegal or harmful content. They don't just write "I am selling drugs"; they hide it. They use:

  • Tiny fonts that look like dust.
  • Distorted text that looks like a melted candle.
  • AI-generated images where the text is blended into the texture of a cloud.
  • Secret codes hidden inside Mahjong tiles or watermarks.

Standard AI models (the general librarians) are too slow to check every single post in real-time, and they are too "dumb" to spot these clever tricks.

The Solution: Xuanwu VL-2B
The team at Hello Group built a new kind of librarian called Xuanwu. Instead of trying to be the biggest, most expensive library in the world, they built a highly efficient, specialized detective that fits in a small backpack (only 2 billion parameters, which is small for an AI).

Here is how they built Xuanwu, using simple analogies:

1. The Architecture: The "Swiss Army Knife" vs. The "Tank"

Most AI models are like giant tanks: heavy, powerful, but slow and expensive to run.
Xuanwu is like a Swiss Army Knife.

  • The Eyes (Vision): They gave Xuanwu a pair of high-powered, zoom-in glasses (InternViT) that can see tiny details, like a single pixel of text in a corner.
  • The Brain (Language): They connected these glasses to a very smart, compact brain (Qwen3) that understands context and rules.
  • The Connector: They used a simple, fast bridge (MLP) to connect the eyes to the brain.
  • The Result: It's small enough to run fast on a phone or a server, but sharp enough to spot a needle in a haystack.

2. The Training: The "Three-Stage School"

You can't just hand a detective a case file and expect them to solve it. Xuanwu went through a strict three-stage training program:

  • Stage 1: Pre-Training (The University Degree)
    Xuanwu read millions of books and looked at millions of pictures. It learned what a cat looks like, how to read a menu, and how to solve math problems. This gave it a solid foundation of general knowledge.
  • Stage 2: Mid-Training (The Specialized Boot Camp)
    This is where the magic happened. The trainers stopped showing Xuanwu general books and started showing it real-world crime scenes. They showed it thousands of examples of spam, scams, and hidden codes.
    • Crucial Trick: To make sure Xuanwu didn't forget how to read Shakespeare while learning to spot spam, they mixed in a little bit of the old "University" books every day. This kept its general smarts alive while sharpening its detective skills.
  • Stage 3: Post-Training (The Field Internship)
    Xuanwu was put in a simulated environment where it had to practice catching bad guys.
    • The "CoT" (Chain of Thought): Instead of just guessing "Yes" or "No," Xuanwu was taught to think out loud. It learns to say: "I see a picture of a cat. Wait, there is tiny text in the corner that says 'Add me on WeChat.' That is a rule violation. Therefore, I will flag this." This makes the AI explainable and trustworthy.
    • The "Reward System" (RL): When Xuanwu caught a tricky, distorted code, it got a digital "high five." When it missed one, it got a gentle correction. This reinforced its ability to spot the hardest tricks.

3. The Superpower: Dynamic Vision

Imagine trying to read a sign that is 100 feet away, but also trying to read a tiny label on a bottle in the foreground. If you squint at the sign, the bottle blurs. If you focus on the bottle, the sign disappears.
Xuanwu uses a "Dynamic Tiling" trick. It doesn't just look at the whole image at once. It breaks the image into puzzle pieces (tiles). It looks at the big picture to understand the scene, then zooms in on specific tiles to read the tiny, distorted text. It does this without getting confused or running out of memory.

The Results: Why It Matters

The paper shows that Xuanwu is a champion in the "Content Moderation Olympics":

  • Speed & Cost: Because it's small, it's cheap to run. A platform can check millions of posts a second without breaking the bank.
  • Accuracy: It caught 94% of the bad content in tests, beating much larger models.
  • The "Adversarial" Win: In tests where bad actors tried to trick the AI with weird distortions and hidden text, Xuanwu caught 82.8% of them. This is better than even the most expensive, commercial "super-models" (like Gemini) that weren't specifically trained for this job.

The Bottom Line

Xuanwu proves that you don't need a giant, expensive AI to solve complex problems. By building a compact, specialized detective that is trained on real-world tricks and taught to "think step-by-step," you can create a system that is fast, cheap, and incredibly good at keeping the internet safe from hidden dangers.

It's the difference between hiring a giant, slow-moving elephant to find a lost ring in a forest, versus hiring a small, hyper-observant squirrel that knows exactly where to look.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →