Position: Assistive Agents Need Accessibility Alignment
This paper argues that assistive agents for Blind and Visually Impaired users must treat accessibility as a core alignment problem rather than a peripheral concern, proposing a new design pipeline to address the systematic failures caused by current sighted-centric assumptions in agentic AI.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are hiring a very smart, highly educated tour guide for a blind traveler. This guide has read every map in the world and can see everything perfectly. However, there's a catch: the guide was trained by watching people with full vision walk around. They assume that if they give a wrong turn, the traveler can just look up, see the mistake, and say, "Hey, that's a wall, not a path!"
The paper argues that blind and visually impaired (BVI) users cannot do this. They can't "look up" to check the guide. If the guide is wrong, the traveler might walk into traffic or drop a pill they can't see.
Here is the breakdown of the paper's argument using simple analogies:
1. The Core Problem: The "Sighted" Bias
Current AI assistants are like tour guides who think everyone has eyes. They are designed to be fast, confident, and efficient. But for a blind user, speed and confidence can be dangerous.
- The "Hallucination" Trap: If a sighted guide guesses a landmark is there, the traveler can check. If a blind traveler's AI guide guesses a crosswalk is clear, the traveler has no way to verify it before stepping out. The paper calls this a "silent failure"—the AI is confidently wrong, and the user has no idea until it's too late.
- The Cost of Mistakes: For a sighted person, a wrong turn is a minor annoyance. For a blind person, a wrong turn could mean physical injury, taking the wrong medication, or losing money. The paper says current AI doesn't understand that mistakes cost more for these users.
2. The Evidence: What Goes Wrong?
The authors looked at 778 real-world tasks that blind people need help with (like navigating streets, reading medicine labels, or using a microwave). They found that even the smartest AI systems fail because they don't fit the "blind experience."
They grouped these tasks into four main areas:
- Moving Around (Mobility): Walking safely, avoiding obstacles, and finding paths.
- Reading (Text Access): Reading menus, signs, emails, or complex charts.
- Daily Objects: Figuring out what a button does on a washing machine or if a stove is on.
- Goal Questions: Asking things like, "Where is my keys?" or "Is this exit safe?"
The Big Finding: The AI often fails not because it's "dumb," but because it's misaligned. It assumes the user can verify the answer, can handle a lot of talking (cognitive load), and that a mistake is easy to fix. None of these are true for blind users.
3. The Solution: "Accessibility Alignment"
The paper proposes a new rulebook called Accessibility Alignment. Think of this as changing the tour guide's job description. Instead of just "getting the traveler to the destination," the guide must now prioritize safety, verification, and trust.
They suggest a four-part checklist for building these agents:
- Goal Alignment: Success isn't just "arriving." It's arriving safely. If the route is risky, the AI should say, "I'm not sure, let's wait," rather than guessing.
- Interaction Alignment: The AI shouldn't talk in long, confusing paragraphs. It needs to speak in short, clear chunks that are easy to follow without seeing.
- Risk Alignment: The AI must know when to be scared. If the data is blurry (like a smudged medicine label), the AI should admit, "I can't read this clearly," instead of making up an answer.
- Lifecycle Alignment: The AI needs to learn from its mistakes over time without forgetting safety rules. It needs a way to be audited and fixed if it starts acting recklessly.
4. The New Blueprint: A Three-Step Pipeline
To build these safe agents, the authors propose a specific workflow:
- Phase 1: Design (The Blueprint): Before writing code, you define the "Red Lines." What is the AI never allowed to do? (e.g., "Never tell a blind person to cross a street if the traffic light is unclear.") You also design how the AI admits uncertainty.
- Phase 2: Deployment (The Test Drive): You test the AI in high-stress situations (like a busy intersection or a dark room) to make sure it hits the "pause" button when it's unsure, rather than guessing.
- Phase 3: Iteration (The Feedback Loop): Once the AI is out in the world, you watch closely. If it almost makes a mistake, you fix the rules so it never happens again.
5. Debunking Two Common Myths
The paper addresses two wrong ideas people might have:
- Myth 1: "Just make the AI smarter."
- Reality: Making an AI "smarter" (generalizing) doesn't fix this. A super-smart guide who assumes you have eyes is still dangerous. You need a guide specifically trained for blindness, not just a smarter version of a sighted guide.
- Myth 2: "This is just an interface problem."
- Reality: It's not just about making the screen reader speak louder. The problem is deep inside the AI's brain (its decision-making). If the AI decides to act too quickly without checking facts, changing the voice won't fix it. The thinking process itself needs to change.
Summary
The paper argues that we cannot just "add accessibility" to AI at the end. We need to build Accessibility Alignment into the DNA of the AI from the start. For blind users, an AI that is "mostly right" is not good enough; it needs to be safely right, knowing when to stop, when to ask for help, and when to admit it doesn't know.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.