← Latest papers
💻 computer science

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models

This paper introduces Q-Mask, a novel OCR framework that employs a causal query-driven mask decoder to explicitly ground text to spatial regions before recognition, supported by the TextAnchor-Bench benchmark and the large-scale TextAnchor-26M dataset to significantly improve text anchoring capabilities in vision-language models.

Original authors: Longwei Xu, Feng Feng, Shaojie Zhang, Xin Chen, Hang Li, Anan Du, Hailong Yu, Pei Fu, Zhenbo Luo, Jian Luan

Published 2026-04-22
📖 4 min read☕ Coffee break read

Original authors: Longwei Xu, Feng Feng, Shaojie Zhang, Xin Chen, Hang Li, Anan Du, Hailong Yu, Pei Fu, Zhenbo Luo, Jian Luan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to find a specific word in a crowded, messy room full of people shouting different things.

The Problem:
Current "smart" AI models (Vision-Language Models) are like brilliant librarians who can read any book perfectly. If you show them a picture of a receipt, they can tell you the total price or the date. However, they have a weird blind spot: they can't point to where the text is.

If you ask, "Where is the word 'Total'?", a standard AI might say, "It's 12.50," but it can't draw a box around it. It's like a librarian who knows the answer but refuses to point to the shelf where the book sits. This is a problem for things like smart glasses or robots that need to interact with the physical world—they need to know where to look, not just what to read.

The Solution: Q-Mask
The researchers at Xiaomi created a new system called Q-Mask. Think of it as teaching the AI a new way of thinking: "Look before you read."

Instead of trying to read the whole image and guess the answer at the same time, Q-Mask breaks the process into two distinct steps, similar to how a detective solves a case:

  1. The "Where" Step (The Search): First, the AI looks at your question (e.g., "Where is the price?") and scans the image to draw a mental "spotlight" or a mask over the specific area where that information likely lives. It ignores everything else.
  2. The "What" Step (The Reading): Only after it has isolated that specific spot does it zoom in and read the text inside that spotlight.

The Creative Analogy: The "Flashlight" vs. The "Floodlight"

  • Old AI (Floodlight): Shines a bright light over the entire messy room. It sees everything at once, gets overwhelmed, and sometimes confuses a word on a poster with a word on a menu. It guesses the answer based on the whole scene.
  • Q-Mask (Flashlight): You ask, "Where is the exit sign?" The AI first uses a flashlight to find the exit sign. Once the flashlight is locked onto the sign, then it reads the letters. This ensures the answer is grounded in reality.

How They Taught It (The Training Data)
To teach the AI this new skill, the researchers couldn't just use normal books. They needed a special training manual called TextAnchor-26M.

Imagine trying to teach a child to find hidden objects. You wouldn't just show them a picture and ask "What do you see?" You would show them a picture, point to a specific spot, and say, "The word 'Apple' is right here."

  • They created 26 million examples of this.
  • They used a trick called "De-stylized Mask Rendering." Instead of showing the AI the exact, fancy font from a real sign (which might be hard to read), they took the text and printed it in a plain, boring font on a white background, matching the shape of the original sign. This taught the AI to focus on the shape and location of the text, not the fancy design.

The Result: The "TextAnchor-Bench" (TABench)
The team also built a test called TABench to see if other AIs could do this. They found that even the most famous, powerful AI models (like GPT-4o or Gemini) were terrible at pointing to text. They could read it, but they couldn't anchor it to a location.

Q-Mask, however, passed the test with flying colors. It became much better at:

  • Reading: Getting the text right because it focused on the right spot.
  • Pointing: Accurately drawing a box around the text.

Why This Matters
This isn't just about reading better. It's about interaction.

  • Smart Glasses: If you wear glasses that tell you what a sign says, you want the glasses to highlight the sign for you, not just speak the words.
  • Robots: If a robot needs to pick up a box labeled "Fragile," it needs to know exactly where that label is on the box to grab it safely.

In Summary:
Q-Mask is a new way of teaching AI to stop guessing and start searching. By forcing the AI to find the location of the text before reading it, the researchers created a system that is more accurate, more reliable, and much better at understanding the physical world. It's the difference between a student who guesses the answer on a test and a student who shows their work by pointing exactly to where the answer is found.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →