← Latest papers
💻 computer science

WikiVQABench: A Knowledge-Grounded Visual Question Answering Benchmark from Wikipedia and Wikidata

The paper introduces WikiVQABench, a human-curated benchmark that combines Wikipedia images, captions, and Wikidata to evaluate the knowledge-grounded reasoning capabilities of vision-language models through multiple-choice questions that require external information beyond visual perception.

Original authors: Basel Shbita, Pengyuan Li, Anna Lisa Gentile

Published 2026-05-21
📖 4 min read☕ Coffee break read

Original authors: Basel Shbita, Pengyuan Li, Anna Lisa Gentile

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are playing a game of "Guess the Picture" with a friend. Usually, the game is simple: you show a picture of a red apple, and your friend says, "That's an apple!" or "It's round!" This is how most current AI vision tests work. They check if the AI can see what's in the picture.

But in the real world, looking isn't always enough. Imagine you show your friend a picture of a very specific, rare spider. If you ask, "What family of spiders is this?" your friend can't answer just by looking at the web or the legs. They need to know biology facts that aren't visible in the photo. They need to pull information from their brain's "library" of knowledge.

WikiVQABench is a new test designed to see if AI can do exactly that: combine what it sees with what it knows from a library of facts.

Here is how the paper explains it, broken down into simple parts:

1. The Problem: AI is Too "Surface-Level"

Current AI models are great at describing what is right in front of them (like "a man holding a bat"). But they struggle when a question requires outside knowledge.

  • The Old Way: Show a picture of a landmark and ask, "What is this?" (Answer: "A tall stone tower.")
  • The New Challenge: Show the same picture and ask, "Which ancient civilization built this?" (Answer: "The Phoenicians.") You can't see "Phoenicians" in the stone; you have to know the history.

2. The Solution: A "Library-Linked" Test

The researchers built a new test called WikiVQABench. Think of it as a quiz where every question is tied to a specific page in a giant encyclopedia (Wikipedia) and a structured database of facts (Wikidata).

  • The Ingredients: They took real photos from Wikipedia, read the articles next to them, and grabbed structured facts from Wikidata (like a digital filing cabinet of facts).
  • The Recipe: They used a smart computer program (an LLM) to mix these ingredients together to create multiple-choice questions.
  • The Human Touch: This is the most important part. The computer made thousands of questions, but it made mistakes. So, real humans acted as "editors." They checked every single question to make sure:
    1. The answer is actually true.
    2. The question matches the picture.
    3. Crucially: You cannot answer the question just by looking at the picture. You must use the outside knowledge.

3. The Test: "The Knowledge Gap"

The researchers tested 15 different AI models on this new quiz. These models ranged from tiny ones (like a pocket calculator) to massive ones (like a supercomputer).

The Results:

  • The Big Gap: The best AI got about 76% correct. The smallest AI got only 25% correct (which is basically guessing randomly).
  • The Reality Check: Even the biggest, smartest AI didn't get 100%. This proves the test is hard and that current AI still struggles to perfectly blend "seeing" with "knowing."
  • Size Matters (but not everything): Bigger models generally did better, but just making a model bigger didn't automatically make it a genius at this specific type of reasoning.

4. Why This Matters

Think of this benchmark as a "driver's license test" for AI.

  • Old Tests: Checked if the AI could see the steering wheel and the road.
  • WikiVQABench: Checks if the AI knows the traffic laws, the history of the road, and the rules of the city, not just what the road looks like.

The paper concludes that we need tests like this to find the "blind spots" in AI. It shows that while AI is getting better at seeing, it still needs to get much better at understanding the world behind the image. The researchers have made this test and the data available for everyone to use, hoping it helps build smarter, more knowledgeable AI in the future.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →