← Latest papers
🤖 AI

What Users Leave Unsaid: Under-Specified Queries Limit Vision-Language Models

This paper introduces HAERAE-Vision, a benchmark of real-world, under-specified visual queries from Korean communities, revealing that current vision-language models struggle significantly with natural user inputs due to missing context rather than inherent capability gaps, and that explicit query rewriting yields far greater performance gains than web search.

Original authors: Dasol Choi, Guijin Son, Hanwool Lee, Minhyuk Kim, Hyunwoo Ko, Teabin Lim, Ahn Eungyeol, Jungwhan Kim, Seunghyeok Hong, Youngsook Song

Published 2026-04-14
📖 5 min read🧠 Deep dive

Original authors: Dasol Choi, Guijin Son, Hanwool Lee, Minhyuk Kim, Hyunwoo Ko, Teabin Lim, Ahn Eungyeol, Jungwhan Kim, Seunghyeok Hong, Youngsook Song

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a brilliant, super-smart detective (the AI) who has been trained on millions of books, movies, and clear, well-written police reports. You are excellent at solving crimes when the clues are laid out perfectly.

Now, imagine a real person walks into your office. They are flustered, holding a blurry photo of a strange object, and they say: "What is this? How do I fix it?"

They don't tell you what the object is, where they found it, or what they were trying to do. They just assume you can read their mind because they are holding the picture.

This paper, HAERAE-Vision, is about testing how well these "super-detective" AIs can handle real people who talk like that, rather than how they handle perfect test questions.

Here is the breakdown of the story:

1. The Problem: The "Perfect Test" vs. Real Life

Most AI tests are like a standardized exam. The questions are clear: "Identify the animal in this picture."
But real life is messy. People ask: "My dog ate this, where do I buy a new one?" (while pointing at a torn piece of wood).

The researchers found that current AI models are like students who ace the textbook but fail the real-world application. They get confused when the user doesn't spell everything out.

2. The Experiment: The "Korean Community" Challenge

The team went to Korean online forums (like Reddit or Quora) and collected 86,000 real questions people asked with photos.

  • The Filter: They were very strict. They threw out 99.2% of the questions because they were too easy, too weird, or didn't need the photo to answer.
  • The Result: They ended up with 653 "gold standard" messy questions. These are the real deal: informal, vague, and relying heavily on context.

They created two versions of every question:

  1. The Messy Version: What the user actually typed (e.g., "What's this worm?").
  2. The "Cleaned-Up" Version: A rewrite that fills in the blanks (e.g., "What is the name of this marine snail, Dendropoma maxima, found in Jeju Island waters?").

3. The Big Discovery: "Clarification" is Magic

They asked 45 different AI models (from the biggest giants like GPT-5 to smaller ones) to answer both versions.

The Shocking Result:

  • On the Messy Version, even the smartest AIs got less than 50% right. They were basically guessing.
  • On the Cleaned-Up Version, the same AIs jumped up to 57% or higher.

The Analogy:
Think of the AI as a chef.

  • Messy Query: A customer says, "Make me that thing I had yesterday." The chef has no idea what they want.
  • Cleaned Query: The customer says, "Make me the spicy beef stew I had at the restaurant on 5th Street."
  • The Lesson: The chef didn't get better at cooking; the customer just gave better instructions. The difficulty wasn't the AI's brain; it was the user's vague request.

Who benefited most? The smaller, cheaper AI models. They were the ones who got completely lost without the extra details. The big models were a bit better at guessing, but they still struggled.

4. The "Google Search" Trap

The researchers thought, "Okay, if the AI doesn't know, let's give it a Google Search button!"

  • Result: It helped a little bit, but not enough.
  • Why? If you ask, "How do I fix this?" without saying what "this" is, Google can't help you. The AI needs to understand the intent first before it can search for the answer. You can't search for a question you haven't fully formed yet.

5. The Hidden Hurdle: Cultural Context

Even after fixing the vague questions, the AIs still got some wrong. Why?

  • The Cultural Gap: Some questions relied on things only a Korean person would know.
    • Example: A photo of orange bags on a rural road. An AI might guess "trash" or "safety cones." A Korean local knows immediately: "Those are sandbags for winter snow removal."
  • The AI isn't just missing facts; it's missing the "local vibe" and cultural assumptions that humans take for granted.

The Takeaway

This paper tells us that AI isn't as "dumb" as we think, but it's also not as "smart" as we hope.

  1. The Benchmark Lie: Current tests are too clean. They make AI look smarter than it is in the real world.
  2. The User's Job: We can't just expect AI to read our minds. We need to learn how to give better context, or we need tools that help us clarify our own questions.
  3. The Future: To make AI truly useful, we need to teach it not just to "see" and "read," but to understand the messy, cultural, and incomplete way humans actually talk.

In short: The AI isn't failing because it can't see the picture; it's failing because it doesn't know why you're showing it the picture.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →