← Latest papers
💻 computer science

INST-IT: Boosting Instance Understanding via Explicit Visual Prompt Instruction Tuning

The paper introduces Inst-IT, a comprehensive framework comprising a diagnostic benchmark, a large-scale instruction-tuning dataset, and a continuous training paradigm that leverages explicit visual prompts to significantly enhance both instance-level and generic understanding capabilities in Large Multimodal Models.

Original authors: Wujian Peng, Lingchen Meng, Yitong Chen, Yiweng Xie, Yang Liu, Tao Gui, Hang Xu, Xipeng Qiu, Zuxuan Wu, Yu-Gang Jiang

Published 2026-03-17
📖 4 min read☕ Coffee break read

Original authors: Wujian Peng, Lingchen Meng, Yitong Chen, Yiweng Xie, Yang Liu, Tao Gui, Hang Xu, Xipeng Qiu, Zuxuan Wu, Yu-Gang Jiang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very smart, well-traveled friend who loves to look at photos and watch videos. This friend is great at describing the "big picture." If you show them a photo of a busy street, they can tell you, "It's a sunny day in a city with lots of cars and people walking."

But, if you ask them, "What is the specific person in the red hat doing in the third frame, and how did they move compared to the second frame?" your friend might get confused. They might mix up the people, forget who was wearing what, or just give a vague answer like, "Someone moved."

This is the problem the paper Inst-IT is trying to solve. It's about teaching these "smart friends" (which are actually AI models called Large Multimodal Models) to stop just looking at the whole forest and start paying attention to every single tree.

Here is a breakdown of their solution using some everyday analogies:

1. The Problem: The "Blurry Glasses" Effect

Current AI models are like someone wearing blurry glasses. They can see the general shape of a crowd, but they struggle to pick out specific individuals. They often hallucinate (make things up) or lose track of who is who when things move around, especially in videos.

2. The Solution: "Name Tags" for Everything

The researchers came up with a clever trick called Explicit Visual Prompting.

Imagine you are at a crowded party, and you want to tell a story about a specific guest. Instead of saying, "That guy over there," you put a bright, glowing name tag (like a number: 1, 2, 3) on everyone's forehead.

  • The AI's Job: Now, when you ask, "What is person #3 doing?", the AI doesn't have to guess. It can clearly see the tag "3" and track exactly that person, even if they move behind a pillar or change direction.

The paper calls this Set-of-Marks (SoM). They take a video or photo, use a tool to draw these invisible number tags on every object, and then show this "tagged" version to the AI.

3. The Training: The "Super-Notebook"

To teach the AI to use these tags, the researchers built a massive training dataset (a giant textbook for the AI).

  • The Content: They didn't just write simple sentences. They created a "Super-Notebook" that includes:
    • Descriptions of individuals: "Person #1 is wearing a blue shirt."
    • Descriptions of the whole scene: "It's a wedding."
    • The "Story of Change": "Between frame 1 and frame 2, Person #1 walked from the left to the right."
    • Questions and Answers: "What did Person #2 hold in their hand?"
  • The Scale: They made this for 51,000 photos and 21,000 videos. It's like giving the AI a library of millions of stories where every character is clearly labeled.

4. The Result: From "Generalist" to "Detective"

After training with this "tagged" data, the AI transformed.

  • Before: It was a generalist who could describe a scene but failed at details.
  • After: It became a detective. It can now answer complex questions like, "Did the woman in the red dress drop her bag before or after the car passed?" with high accuracy.

Why This Matters (The "So What?")

The paper shows that by teaching the AI to focus on specific details (instances), it actually got better at everything else, too.

  • Analogy: Think of it like a student who practices solving very specific, hard math problems. You might think they would only get good at those specific problems. But in reality, by mastering the details, they actually get better at general math, reading comprehension, and logic.
  • Real World: This means future AI assistants won't just say, "There is a car in the video." They will be able to say, "The blue car on the left swerved to avoid the dog, while the red car on the right kept driving straight." This is crucial for things like self-driving cars, medical imaging (spotting a specific tumor), or helping the visually impaired navigate the world.

Summary

Inst-IT is a new method that teaches AI to stop squinting at the whole picture and start reading the fine print. By putting "name tags" on objects and training the AI with a massive library of detailed stories about those specific objects, they created a smarter, more observant AI that understands the world one detail at a time.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →