LaViSA: A Language and Vision Structural Ambiguity Benchmark
This paper introduces LaViSA, a benchmark comprising ambiguous sentences and corresponding disambiguating images across seven categories, to evaluate Vision-Language Models' ability to resolve structural ambiguity, revealing that while current models show some capability, they still struggle with specific ambiguity types and subtle visual distinctions.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are looking at a picture of a horse and a bird. Someone hands you a note that says, "A horse and a bird flying."
Now, pause. What does that actually mean?
- Interpretation A: The horse is flying, and the bird is flying.
- Interpretation B: The horse is standing on the ground, and only the bird is flying.
This is what linguists call structural ambiguity. The sentence is a puzzle with two valid solutions, and without more clues, you can't know which one is right. Humans are pretty good at this; if we see a picture where the horse is on the ground, our brains instantly pick Interpretation B.
But can AI do the same? That is the question this paper, LaViSA, sets out to answer.
The Problem: The AI's "Blind Spot"
The authors created a new test called LaViSA (Language and Vision Structural Ambiguity). Think of it as a "gym" for AI models to practice their eyes and ears simultaneously.
In this gym, the AI is given:
- A tricky sentence (like the horse/bird example).
- A picture that clarifies the meaning.
- A multiple-choice list of what the sentence could mean.
The AI's job is to look at the picture, read the sentence, and pick the correct meaning. It's like a game of "Guess the Story" where the picture is the only hint.
The Experiment: Who Passed the Test?
The researchers tested a wide variety of AI models, from the super-smart, expensive "proprietary" ones (like the latest GPT and Gemini versions) to free, open-source models.
Here is what they found, using a simple analogy:
- The "Smart" Models (Proprietary): These models are like students who studied hard. They generally did very well, especially with the "easy" puzzles. For example, if the sentence was about where something was attached (like "The boy calls the girl with a whistle"), they could usually tell if the boy or the girl was holding the whistle just by looking at the picture.
- The "Struggling" Models (Open Source): Some smaller open-source models were like students who hadn't studied enough. They often guessed randomly or got confused, especially when the picture was a cartoon rather than a realistic photo.
- The "Thinking" Models: The researchers also tested models that are designed to "think" before answering (like a student who writes out their work). Surprisingly, these didn't always do better than the standard models. Sometimes, thinking too much actually led them down the wrong path.
The Big Reveal: Where AI Gets Stuck
While the AI models are getting better, they still have a major blind spot. The paper found that AI is great at spotting objects but terrible at understanding relationships.
Imagine a scene with a girl, a laptop, a pen, and a book.
- The Sentence: "The girl holds the laptop or the pen and the book."
- The Trap: In the picture, the girl is holding the pen and the book. The laptop is just sitting on a shelf nearby.
- The AI's Mistake: Because the laptop is visually present (it's right there in the photo), the AI gets distracted. It thinks, "Oh, the laptop is in the picture, so the girl must be holding it!" It fails to understand that "holding" is a specific action, not just "being near."
The paper calls this a struggle with Conjunction Scope (grouping words together) and Ellipsis (leaving out words).
- Conjunction Scope: The AI can't figure out which items belong to which action when there are "or" and "and" in the sentence.
- Ellipsis: If a sentence says, "The lion eats the chicken. Also the cat," the AI struggles to figure out if the cat is eating the chicken too, or if the cat is just being eaten. It gets lost in the visual details and misses the grammar.
The Verdict
The paper concludes that visual scenes are a powerful tool for helping AI understand language. When the picture is clear, the AI can often solve the puzzle.
However, the AI is still like a tourist who can recognize landmarks but doesn't understand the local culture. It sees the objects in the picture, but it often fails to connect them to the correct "story" in the sentence. It struggles to say, "Just because the object is there, doesn't mean it's doing the action the sentence describes."
The authors built this benchmark to show us exactly where these AI models are failing, so future researchers can teach them to look deeper than just the surface of the image.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.