← Latest papers
💻 computer science

"It's trained by non-disabled people": Evaluating How Image Quality Affects Product Captioning with Vision-Language Models

This paper evaluates how common image quality issues degrade the accuracy of Vision-Language Models in generating product captions for blind and low-vision users, revealing a significant drop in performance and emphasizing the need for disability-centered evaluation practices to ensure these tools meet real-world accessibility needs.

Original authors: Kapil Garg, Xinru Tang, Jimin Heo, Dwayne R. Morgan, Darren Gergle, Erik B. Sudderth, Anne Marie Piper

Published 2026-04-01
📖 5 min read🧠 Deep dive

Original authors: Kapil Garg, Xinru Tang, Jimin Heo, Dwayne R. Morgan, Darren Gergle, Erik B. Sudderth, Anne Marie Piper

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are blindfolded, trying to identify a jar of peanut butter in your pantry. You pick up your phone, snap a picture, and ask a super-smart AI assistant, "What is this?"

In a perfect world, the AI would instantly say, "That's a jar of creamy peanut butter, brand name 'Jif', with no added sugar." But in the real world, your hand might be shaking, the lighting might be dim, or you might have taken the photo from a weird angle. The AI might squint at the blurry, sideways image and guess, "I think that's a jar of... maybe jelly? Or perhaps a can of soup?"

This paper, titled "It's trained by non-disabled people," investigates exactly why this happens and how we can fix it. The researchers looked at how Vision-Language Models (VLMs)—the "eyes" and "brains" behind tools like Seeing AI and ChatGPT—perform when blind and low-vision (BLV) people try to identify everyday products like food and medicine.

Here is the breakdown of their findings, using some everyday analogies.

1. The "Perfect Photo" Problem

Most AI models are like students who only study for exams using perfect, textbook diagrams. They are trained on millions of high-quality, perfectly centered, well-lit photos taken by sighted people.

  • The Reality: When a blind person takes a photo, it's often shaky, blurry, or cut off at the edges (like trying to take a picture of a whole car while only seeing the front bumper).
  • The Result: The AI gets confused. The study found that while these AIs are 98% accurate on perfect photos, their accuracy plummets to 75% when the photo has common issues like blur or bad framing. If the photo has multiple problems (blurry AND sideways), the accuracy drops even further.

2. The "Hallucination" Hazard

When the AI can't see clearly, it doesn't just say, "I don't know." Instead, it often guesses confidently, a phenomenon known as "hallucination."

  • The Analogy: Imagine a student taking a test who doesn't know the answer. Instead of leaving it blank, they write down a completely made-up answer that sounds plausible.
  • The Danger: In the study, the AI mistook a bottle of lotion for a pen, or identified a can of pears as peaches. For a sighted person, this is a funny mistake. For a blind person with a severe allergy to peaches, this mistake could be life-threatening. The AI might say, "This is safe," when it's actually dangerous.

3. The "Missing Details" Issue

Even when the AI gets the general idea right, it often misses the critical details that blind people actually need.

  • The Analogy: Imagine asking a friend, "What's in this box?" and they say, "It's a box of cereal." You ask, "Which one? Is it the sugary kind or the healthy kind? Is it the one with raisins or the one with nuts?" They just shrug and say, "It's cereal."
  • The Finding: The study showed that AI often misses the brand name, the flavor, or the ingredients. It might say "soup" when you specifically need to know it's "low-sodium tomato soup." For someone managing diabetes or a food allergy, "soup" isn't good enough; they need the specific label.

4. The "Blind Photographer" Struggle

The paper also surveyed 86 blind people to understand their experience.

  • The Frustration: Taking a good photo is the hardest part. Even with apps that try to guide them (like beeping when the object is centered), users found it difficult to hold the phone steady or know if the lighting was right.
  • The Privacy Dilemma: Many users prefer using AI over asking a human helper because they don't want to share private details (like their medication or toiletries) with a stranger. But if the AI is unreliable, they are forced to choose between privacy and safety.

5. The Core Message: "Trained by Non-Disabled People"

The title of the paper hits the nail on the head. The AI models are like gym trainers who have never experienced a disability. They don't understand what it's like to take a photo with a shaky hand or in a dark room. They are optimized for "perfect" inputs, not "real life" inputs.

Because of this, the AI fails when it matters most: in the messy, imperfect reality of daily life.

What Does This Mean for the Future?

The researchers aren't just pointing out problems; they are offering a roadmap to fix them:

  1. Train on "Messy" Data: Instead of just feeding the AI perfect photos, we need to train it on thousands of blurry, sideways, and poorly lit photos taken by blind people. This is like teaching the student to take the exam with a shaky hand.
  2. Better Feedback: Instead of just saying "I can't see this," the app should say, "I can't see this because the photo is too blurry. Please try moving the light closer."
  3. Honesty: The AI should learn to say, "I'm not sure," rather than guessing confidently. It's better to admit ignorance than to give dangerous advice.

In short: This paper argues that for AI to be truly helpful to blind and low-vision people, it needs to stop acting like a sighted person looking at a perfect photo and start acting like a helpful assistant who understands the challenges of the real world. We need to build AI that is robust enough to handle the "blurry" moments of life.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →