Large Language Models for Web Accessibility: A Systematic Literature Review
This paper presents a systematic literature review of 38 studies analyzing how Large Language Models are currently applied to web accessibility tasks, highlighting their predominant focus on text-centric issues and WCAG guidelines while identifying significant gaps in addressing cognitive accessibility and involving users with disabilities in evaluations.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the internet as a giant, bustling city. For this city to be truly "accessible," it needs to be easy for everyone to navigate, regardless of whether they use a wheelchair, a guide dog, a hearing aid, or a different way of thinking. Unfortunately, many parts of this city are currently built with broken sidewalks, missing signs, or confusing layouts that block people with disabilities.
In recent years, a new type of "super-intelligent assistant" called a Large Language Model (LLM) has arrived in this city. Think of an LLM as a very well-read, fast-talking robot that can write code, describe pictures, and fix text. Researchers have been asking: Can this robot help us rebuild the city to be more accessible?
This paper is a Systematic Literature Review, which is like a detective's report. The authors (Wajdi Aljedaani and Rubel Hassan Mollik) gathered and analyzed 33 different studies (they mention 38 in the abstract, but the final count in the methodology is 33) to see exactly how these robots are being used to fix the internet's accessibility problems.
Here is what they found, broken down simply:
1. What Jobs Are the Robots Doing?
The researchers found that these AI robots are mostly being used as specialized helpers rather than full-time city planners. They are doing specific, repetitive tasks:
- Writing Descriptions: If a website has a picture without a description, the robot writes one (like describing a photo for someone who is blind).
- Finding Broken Signs: The robot scans a webpage to find things that are broken or confusing.
- Fixing Code: When a problem is found, the robot suggests how to rewrite the code to fix it.
The Metaphor: Imagine a construction crew. The robots are the workers who are great at painting walls or laying bricks (fixing specific text or code), but they aren't currently being asked to redesign the whole building's architecture.
2. What Rules Are They Following?
The internet has a "rulebook" for accessibility called WCAG (Web Content Accessibility Guidelines). This is like the city's building code.
- The Good News: Almost all the studies use this rulebook. The robots are trained to check if a website follows these rules.
- The Missing Piece: There is another set of rules for cognitive accessibility (helping people with learning disabilities or those who get overwhelmed by complex text), called COGA. The researchers found that the robots almost ignore these rules. They are great at fixing visual problems (like color contrast) but often forget about making things easier to understand for the brain.
3. Which Robots Are They Using?
The studies mostly rely on the big, famous, general-purpose robots (like GPT-4, ChatGPT, and Claude).
- The Metaphor: It's like trying to fix a city using only a Swiss Army Knife. These tools are powerful and can do many things, but the researchers aren't building custom tools specifically designed for accessibility. They are just using the general tools they already have and asking them to do the job.
- How they talk to them: Most researchers just ask the robot a question (a "prompt") like, "Fix this website." They aren't building complex systems where the robot talks to other robots or learns over time.
4. What Problems Are They Solving?
The robots are mostly fixing the "easy" visual problems:
- Missing Alt Text: Describing images.
- Bad Colors: Making sure text stands out against the background.
- Confusing Headings: Organizing text so screen readers can navigate it.
What they are missing: They are rarely fixing problems related to time (like videos without captions), complex navigation, or content that is too hard to read. They also sometimes make up facts (called "hallucinations"), like describing an object in a photo that isn't actually there.
5. How Do We Know They Work?
This is where the report gets a bit critical. The researchers looked at how these tools were tested in the studies.
- The Problem: Many studies tested the robots using other computers or by asking experts to look at the results. Very few studies actually asked people with disabilities to use the tools.
- The Metaphor: It's like testing a new ramp for a wheelchair by having a carpenter look at it and say, "It looks good," instead of actually having a person in a wheelchair try to roll up it.
- The Result: We know the robots can fix code, but we don't have enough proof that the fixes actually help real people in the real world.
The Bottom Line
This paper concludes that Large Language Models are promising assistants for making the web more accessible, but they are currently being used in a limited way.
- They are great at fixing visual and structural text issues.
- They are mostly following the standard visual rules (WCAG) but ignoring the mental/cognitive rules (COGA).
- They are being tested by computers and experts, not enough by the people they are meant to help.
The authors suggest that in the future, we need to build better systems that understand cognitive needs, stop relying only on the "big" general robots, and, most importantly, test these tools with real people who have disabilities to make sure they actually work.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.