Do VLMs See What Sensors Feel? A Scalable Expert-Guided Design for Wheelchair Accessibility Assessment from Street View
This paper proposes an expert-guided retrieval-augmented framework using vision-language models and Google Street View imagery to assess wheelchair accessibility at scale, demonstrating that while model ratings show partial alignment with real-world mobility friction derived from GPS dwell times, they effectively capture structural features like curb ramps but struggle with subtle or transient barriers.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to figure out if a city is easy for a person in a wheelchair to navigate. Traditionally, you'd have to send a team of people to walk (or roll) every single street, checking for broken sidewalks, steep curbs, or blocked paths. It's slow, expensive, and hard to do everywhere.
This paper asks a simpler question: Can a super-smart computer program, looking at photos of the street, figure out the same things a person in a wheelchair would feel?
Here is the story of how they tried to answer that, explained simply:
The Problem: The "Invisible" Bumps
For a wheelchair user, a city isn't just a picture; it's a physical experience. A tiny crack in the pavement, a curb that's too high, or a parked car blocking the path can stop them dead in their tracks. These "friction points" are hard to map because they are everywhere, they change, and they are often too small for a standard map to notice.
The Experiment: The Computer vs. The Wheelchair
The researchers set up a test at the University of Florida. They did two things at the same time:
- The "Feel" Test: They attached a GPS tracker to a manual wheelchair and had it drive around campus. They looked for places where the wheelchair stopped or slowed down for a long time (called "dwell time"). They figured that if the wheelchair stopped, it was probably because the path was hard to navigate.
- The "See" Test: They took Google Street View photos of those exact same spots and fed them into a Vision-Language Model (VLM). Think of a VLM as a robot that has read every book about city planning and can "see" images.
The Secret Sauce: The Expert Guide
The researchers realized that just asking the robot, "Is this accessible?" wasn't enough. The robot might guess wrong or miss details. So, they gave the robot a cheat sheet.
They built a system where the robot could look up rules (like the Americans with Disabilities Act) and advice from real accessibility experts while it was looking at the photo. It's like giving a student a textbook and a study guide right before a test, rather than just letting them guess.
What They Found
The results were a mix of "Great job!" and "Not quite yet."
- The Good News: The robot got better as it got smarter. The biggest, most powerful robot models (the "78B" and "32B" versions) started to agree with the wheelchair's GPS data. When the wheelchair struggled (stopped for a long time), the robot gave the spot a low score. When the path was smooth, the robot gave it a high score.
- Analogy: It's like a weather app that isn't perfect, but if you ask the smartest version, it can usually tell you when it's going to rain just by looking at the clouds.
- The Bad News: The robot still missed some things. It was good at spotting big, obvious problems like missing ramps or blocked paths. But it struggled with subtle issues, like a slightly bumpy surface or a temporary obstacle that wasn't clearly visible in the photo.
- Analogy: The robot is like a person looking at a map; they can see a bridge is out, but they can't feel the pothole in the road unless it's huge.
The "Aha!" Moments
The researchers discovered a few tricks to make the robot work better:
- Don't use the 360-degree view: Surprisingly, the robot worked better when shown two specific, straight-ahead photos (like looking forward and to the side) rather than a full 360-degree panorama. The 360 view was too confusing for the robot to focus on the details.
- Ask more questions: Instead of asking one big question ("Is this accessible?"), they asked the robot ten specific questions (e.g., "Is the curb ramp there?", "Is the path wide enough?", "Are there cracks?"). This "checklist" approach made the robot much more accurate.
The Bottom Line
The paper concludes that these smart computer programs are not ready to replace human inspectors. You can't just let a robot decide if a building is legal or safe.
However, they are excellent scouts. They can scan thousands of streets quickly and say, "Hey, these 50 spots look really bad. Send a human expert to check those first." It's a way to find the trouble spots in a city without having to send a human to check every single block.
In short: The robot can see what the sensor feels, but only if you give it a good guidebook and ask it the right questions. It's a powerful tool for finding problems, but it still needs a human to confirm the details.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.