← Latest papers
🤖 AI

Look on Demand: A Cognitive Scheduling Framework for Visual Evidence Acquisition in Multimodal Reasoning

This paper introduces CSMR, a multimodal reasoning framework that enhances accuracy by employing a cognitive scheduling mechanism where a language model dynamically decides when to invoke an independent visual perception module to acquire task-relevant evidence, thereby overcoming the limitations of static conversion and linguistic dominance in existing approaches.

Original authors: Yang Zhang, Xiaoshuai Sun, Rui Zhao, Wujin Sun, Yidong Chen, Jiayi Ji, Qian Chen, Rongrong Ji

Published 2026-05-28
📖 4 min read☕ Coffee break read

Original authors: Yang Zhang, Xiaoshuai Sun, Rui Zhao, Wujin Sun, Yidong Chen, Jiayi Ji, Qian Chen, Rongrong Ji

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Problem: Two Flawed Ways to Solve Visual Puzzles

Imagine you are a detective trying to solve a mystery based on a single photograph. The paper argues that current AI detectives (multimodal models) usually try to solve this in one of two ways, and both have a major flaw:

  1. The "One-Shot Description" Detective:
    This detective looks at the photo once, writes a long paragraph describing everything they see, and then puts the photo away. They solve the mystery using only that paragraph.

    • The Flaw: When writing the paragraph, they have to summarize everything at once. They might miss tiny, crucial details (like a specific license plate number or a small stain) because they are trying to fit everything into a short summary. By the time they start reasoning, the fine details are already lost.
  2. The "All-in-One" Detective:
    This detective keeps the photo in front of them the whole time while they think. They look at the photo and the text of the question simultaneously.

    • The Flaw: The paper found that this detective gets "distracted" by their own thoughts. Because they are so good at reading and thinking in words, their brain starts to guess the answer based on what usually happens in stories, rather than what is actually in the photo. They might "hallucinate" (make things up) because their linguistic brain is louder than their visual brain. They stop trusting the evidence in front of them.

The Solution: CSMR (The "Smart Manager")

The authors propose a new framework called CSMR (Cognitive Scheduling for Multimodal Reasoning). Think of this not as a single detective, but as a Team of Two working together:

  • The Manager (The Language Model): This is the "brain" that does the thinking, planning, and logic. It never looks at the photo directly.
  • The Photographer (The Perception Module): This is the "eyes." It only looks at the photo when asked and describes exactly what is there.

How it works (The "Look on Demand" Strategy):

Instead of looking at the photo once or staring at it constantly, the Manager follows a smart process:

  1. Start Thinking: The Manager reads the question and starts thinking.
  2. Ask for Evidence: If the Manager realizes, "I don't have enough info to solve this yet," it asks the Photographer: "Hey, can you look at the photo and tell me specifically about the color of the car?"
  3. Get the Answer: The Photographer looks only at that specific part of the photo and sends back a text description.
  4. Update the Plan: The Manager takes that new fact, adds it to their notes, and thinks again.
  5. Repeat or Finish:
    • If they still need more info, they ask another specific question (e.g., "Now, what is the person holding?").
    • If they have enough info, they stop asking and give the final answer.

Why This Works Better

The paper claims this approach solves the problems of the other two methods:

  • No Lost Details: Because the Manager asks for specific details only when needed, the Photographer doesn't have to summarize everything at once. They can zoom in on the tiny clues that matter.
  • No Distracted Thinking: Because the Manager (the thinker) and the Photographer (the viewer) are separate, the Manager can't get "confused" by the photo. The Manager stays focused on logic, and the Photographer stays focused on facts. The Manager only trusts the facts the Photographer reports.
  • Efficiency: The Manager knows when to stop. If they have enough clues after two questions, they don't waste time asking a third. This is called "early termination."

The Results

The researchers tested this "Smart Manager" system on several difficult logic puzzles involving pictures (like science questions or complex scene descriptions).

  • Better Accuracy: The system got the right answer more often than the "One-Shot" or "All-in-One" detectives.
  • Fewer Mistakes: It made up fewer fake details (hallucinations) because it kept checking the photo for proof.
  • No Extra Training: Surprisingly, they didn't have to teach the AI anything new. They just changed how the AI was allowed to talk to itself. They simply gave the "Manager" a set of rules to follow.

In a Nutshell

The paper suggests that the best way for an AI to solve visual problems isn't to try to be a super-genius who sees and thinks at the exact same time. Instead, it should act like a smart project manager: think first, realize what information is missing, ask a specialist to get that specific piece of information, and then continue thinking. This "Look on Demand" approach keeps the AI grounded in reality and leads to smarter answers.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →