SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models
This paper introduces SafetyALFRED, a new benchmark extending the ALFRED environment with real-world kitchen hazards to evaluate multimodal large language models, revealing a significant gap between their ability to recognize safety risks in static question-answering settings and their capacity to actively mitigate those risks through embodied planning.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have hired a very smart, highly educated robot butler named "AI." You've taught it everything about the world: how to cook, how to clean, and even the rules of safety from thousands of books and websites.
You give it a simple job: "Go to the kitchen and wash the butter knife."
The robot looks around, sees a smartphone sitting in the sink, and says, "I see a phone in the sink! That's dangerous because water and electricity don't mix. I should move the phone."
So far, so good. The robot is smart.
But then, you ask the robot to actually do the job. It looks at the phone, nods, and immediately turns on the faucet to wash the knife, completely ignoring the phone. Splash! The phone gets ruined.
This is exactly the problem researchers at the University of Michigan and Boise State University discovered in their new paper, SafetyALFRED.
The Big Idea: "Knowing" vs. "Doing"
The researchers built a test called SafetyALFRED. Think of it like a driving test for robots, but instead of a car, they are testing AI agents in a virtual kitchen. They created six types of "kitchen disasters" waiting to happen, such as:
- A phone in the sink (Water + Electronics = Bad).
- A metal spoon in the microwave (Metal + Microwaves = Fire/Explosion).
- A cabinet door left open (Trip hazard).
- A stove left on (Fire hazard).
They tested 11 of the smartest AI models available (like Qwen, Gemma, and Gemini) to see if they could handle these situations.
The Two Tests
The researchers ran two different tests to see how the robots behaved:
1. The "Quiz Show" Test (Question & Answer)
They showed the robot a picture of the kitchen and asked, "Is there anything dangerous here?"
- Result: The robots were amazing! They got about 90% of the answers right. They could spot the phone in the sink or the open stove almost perfectly. They were like students who aced the safety exam.
2. The "Real Life" Test (Embodied Planning)
They gave the robot the same kitchen scene but told it, "Now, go wash the knife." The robot had to plan its steps and actually move things around.
- Result: The robots failed miserably. Even though they knew the phone was dangerous, they often forgot to move it and just tried to wash the knife. Their success rate dropped to less than 60%, and for some tricky hazards, it was near 0%.
The Analogy: The "Know-It-All" Student
Imagine a student who studies hard for a fire safety exam.
- In the classroom (The Quiz): The teacher asks, "What do you do if you see a fire?" The student shouts, "Call 911 and use an extinguisher!" They get an A+.
- In the real world (The Embodied Task): A trash can actually catches fire. The student, panicked and focused on finishing their homework, accidentally kicks the trash can, making the fire worse, because they were so focused on the "homework" (the task) that they forgot the "safety rule" (the hazard).
The paper calls this the "Alignment Gap." The AI's knowledge of safety doesn't match its actions when it's busy trying to get a job done.
Why Did They Fail?
The researchers found a few reasons for this failure:
- Task Tunnel Vision: When the robot is told to "wash the knife," it gets so obsessed with that goal that it ignores the phone in the sink. It's like a driver so focused on getting to the store that they run a red light.
- Perception Issues: Without extra help (like a text description of what objects are), the robots sometimes couldn't "see" the danger clearly enough in the image to act on it.
- Prioritization: The robots seem to think, "I must finish the task first, safety can wait."
The Solution: The "Safety Co-Pilot"
To fix this, the researchers tried a clever trick. Instead of asking one robot to do everything, they used two robots working together:
- The Safety Judge: A robot whose only job is to look at the scene and shout, "Hey! There's a phone in the sink!"
- The Doer: The robot that actually washes the knife.
When the "Safety Judge" told the "Doer" about the danger, the "Doer" suddenly got much better at fixing it. It's like having a co-pilot in a car who yells, "Brake!" while you are focused on the GPS.
The Takeaway
The main message of this paper is simple but scary: Just because an AI can pass a safety test doesn't mean it will be safe in the real world.
We can't just ask robots, "Do you know what's safe?" and assume they will act safely. We need to build systems where safety is a constant, active part of their decision-making, not just a fact they memorized for a quiz.
The researchers have released their code and data (SafetyALFRED) so other scientists can build better, safer robots that won't accidentally ruin your phone or start a fire while trying to make you dinner.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.