Beyond wheelchairs and blindfolds: Investigating disability stereotypes in T2I models with INCLUDE-BENCH
This paper introduces INCLUDE-BENCH, the first large-scale benchmark for evaluating disability-related biases in text-to-image models, revealing that current systems systematically rely on narrow stereotypes like wheelchairs and exhibit stronger alignment with real-world prejudicial associations when depicting people with disabilities.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a magical art machine that can draw whatever you describe. You say, "Draw a chef," and it spits out a picture of a person in a white hat cooking. You say, "Draw a person with a disability," and... well, that's where the magic gets a little glitchy.
A team of researchers at Utrecht University decided to test this glitch with a new, giant test called INCLUDE-BENCH. They didn't just want to see if the machines were "biased"; they wanted to see how they were biased, using a massive dataset of 119,680 generated images. They asked 17 different art engines (both open-source ones anyone can use and two closed, secret ones) to draw people with disabilities in all sorts of situations.
Here is what they found, served up with a side of reality checks.
The "Wheelchair" Shortcut
The biggest surprise? The machines are obsessed with wheelchairs. When the researchers asked for a "mobility-impaired" person or just a generic "disabled" person, almost every single model immediately reached for the wheelchair. It's like if you asked a chef to draw a "baker," and they only drew someone holding a rolling pin, ignoring the ovens, the flour, or the actual bread.
The paper shows that for these specific prompts, the machines are incredibly consistent at drawing the wheelchair, but they are terrible at drawing anything else. It's a visual shortcut. The machines think, "Oh, disability? That means wheelchair," and they stop there.
The "Old and Sad" Filter
The machines also seem to have a weird filter for age and mood. When they drew people with mobility issues, they were overwhelmingly older. If the person was a woman, she was often drawn in a home, cooking or cleaning. If the person was a man, he was often outside, in a park or a museum.
The researchers found that the machines are basically playing a game of "stereotype bingo." They are picking the most obvious, recognizable clues (like a wheelchair or a blindfold) to make sure you know who the person is, but in doing so, they erase all the other details. They aren't showing a diverse group of people living their lives; they are showing a very narrow, very old, and very specific set of characters.
The "Alignment vs. Diversity" Trade-off
Here is a fun way to think about the math behind the magic. The researchers measured two things:
- Alignment: How well does the picture match the words you typed?
- Diversity: How different are the pictures from each other?
They found a funny trade-off. The pictures that matched the words "mobility impaired" or "disabled" the best (highest alignment) were the least diverse. They all looked almost exactly the same: an old person in a wheelchair.
But when they asked for "blind" or "deaf" people, the pictures were a bit more diverse. You might see a person with a cane, or someone wearing sunglasses, or maybe just someone looking thoughtful. But even then, the machines kept falling back on old tricks, like drawing blindfolds or specific hand gestures.
The "Warmth and Competence" Score
The researchers even invented a new way to score the pictures, called the SCM Score. Imagine a graph where one side is "How nice does this person seem?" (Warmth) and the other is "How capable do they seem?" (Competence).
In the real world, people often think of disabled folks as "nice but not very capable." The paper found that the art machines are copying this exact vibe. The pictures of people in wheelchairs scored low on "competence." They looked less capable than the pictures of people without disabilities. Even when the researchers asked the machines to draw these people doing active things (like walking or chatting), the machines still struggled to make them look capable.
What the Machines Got Wrong (and What They Didn't Test)
It's important to know what this test didn't do. The researchers explicitly said they did not test for cognitive or intellectual disabilities. Why? Because those aren't always visible in a picture. You can't really "draw" a thinking process the same way you can draw a wheelchair. So, the results we have are strictly about things you can see.
Also, the paper doesn't claim to have "fixed" the problem. They didn't find a magic switch to turn off the bias. Instead, they built a giant measuring tape (INCLUDE-BENCH) to show us exactly how broken the tape measure is right now. They found that even when they tried to change the context (asking for a person in a cafe vs. a park), the machines still leaned heavily on the stereotypes.
The Bottom Line
The paper suggests that these art machines aren't just accidentally making mistakes; they are actively reproducing the same narrow, stereotypical stories we see in the real world. They are over-weighting the "diagnostic" clues (the wheelchair, the cane) and ignoring the rest of the human being.
The researchers are pretty sure about this because they generated nearly 120,000 images and checked them with math and AI tools. They aren't guessing; they measured it. And the measurement says: if you want to see a real, diverse, capable person with a disability, you can't just ask the machine yet. The machine is still stuck on the "wheelchair and old age" loop.
So, while the technology is amazing, it's still learning that a person is more than just their disability. Until then, the magic art machine is mostly just drawing the same old story, over and over again.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.