Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models
The paper introduces "Blind-Spots-Bench," a novel benchmark comprising 235 human-simple yet AI-challenging tasks that reveals significant performance gaps between closed-source and open-weight models while highlighting persistent weaknesses in current multimodal systems that existing benchmarks fail to detect.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you're playing a video game where the AI boss has just defeated every single level you've ever beaten. It's crushing high scores, solving math puzzles faster than a calculator, and writing poetry that makes you cry. You'd think, "Okay, this AI is basically a super-genius." But then, you ask it a simple question: "Can you draw a dog with exactly five legs?" or "Write a sentence that is exactly 31 characters long."
Suddenly, the super-genius AI trips over its own feet. It draws a dog with six legs, or it writes a sentence that's 32 characters long. It's like a world-class chess champion who suddenly forgets how to tie their shoelaces.
This is exactly what the researchers at EPFL discovered with their new test, called blind-spots-bench. They didn't just look at how smart these AI models are on the big, fancy exams everyone talks about. Instead, they created a "stress test" made of 235 tiny, tricky tasks that are super easy for humans but seem to trip up even the most advanced AI.
The "Blind Spot" Discovery
The team collected these tricky questions from students in a graduate AI course. The students were asked to find the specific things that the smartest AI chatbots of late 2025 just couldn't get right. They cleaned up the questions and built a system to grade the answers automatically.
When they ran their tests, they found a few surprising things:
1. The "Rich Kid" Advantage
The paper suggests that the super-expensive, "closed-source" models (the ones you can't download and run yourself) are significantly better at these blind spots than the "open-weight" models (the ones anyone can use). Even when both types of models seem to score about the same on regular benchmarks, the closed-source ones still beat the open ones by about 10% on this specific stress test. It's like two runners who look equally fast on a track, but when you ask them to run through a maze while juggling, the one with the expensive training gear wins.
2. Bigger Isn't Always Better
You might think that making an AI model bigger (adding more "brain power") would fix these silly mistakes. But the paper measured this and found that scaling up doesn't always help. In some cases, a smaller version of a model actually did better than its giant sibling. For example, in the Qwen3.5 family, the 35-billion-parameter model sometimes outperformed the massive 397-billion-parameter version on certain logic puzzles. It's like how a giant elephant might struggle to fit through a small door that a cat slips through easily.
3. Tools Don't Fix Everything
Some people thought, "If the AI can't count, let's just give it a calculator!" The researchers tested this by letting the models use code execution tools. The results were a mixed bag. For some models, using tools helped them get the right answer and saved them time. But for others, like GPT-5.4, using tools actually made them worse. It's like giving a chef a fancy new knife; sometimes they chop faster, but sometimes they just cut their thumb and ruin the meal.
4. The "Counting" Crisis
One of the biggest blind spots the paper measured is counting. Whether it's counting objects in a picture or generating an image with a specific number of items, the models struggled. Even the best models only got about 60% of these counting tasks right. It's a persistent weakness, like a musician who can play a complex symphony but can't keep a simple 4/4 beat.
What This Test Doesn't Say
The authors are very careful not to overhype their results. They explicitly state that this isn't a "final boss" test that proves AI is broken forever. They admit the dataset is relatively small (only 235 questions) and was created by students, which means it might be biased toward the specific weaknesses of the models they were testing at the time. They also note that they didn't include a human baseline to compare scores directly, so we don't know exactly how much "better" humans are, just that humans find these tasks trivial.
The Bottom Line
The paper suggests that while AI is getting incredibly good at general tasks, it still has "blind spots" on very specific, concrete skills like counting, spatial reasoning, and following strict character limits. The closed-source models are currently holding a slight lead in fixing these spots, but no single model has mastered them all yet.
Think of blind-spots-bench not as a report card that says "AI is failing," but as a diagnostic tool. It's like a mechanic running a specific test on a car that runs perfectly on the highway but stalls at a stop sign. It helps us see exactly where the engine is sputtering so we can fix it, rather than just assuming the car is perfect because it's fast.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.