← Latest papers
💻 computer science

EmergencyBias: Bias in Text-to-Image Models under Emergency Scenarios

This paper introduces "EmergencyBias," a framework for evaluating demographic and behavioral biases in text-to-image models under emergency scenarios, revealing systematic disparities in risk and intervention portrayals across seven models and proposing "ActionAlign," a lightweight calibration method to mitigate these biases while preserving image quality.

Original authors: Haibo Tang, Linqi Zhang, Hongxin Huan, Chenwei Lin, Xian Xu

Published 2026-08-04
📖 4 min read☕ Coffee break read

Original authors: Haibo Tang, Linqi Zhang, Hongxin Huan, Chenwei Lin, Xian Xu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you've just picked up a magical camera that doesn't take photos of the world as it is, but instead paints pictures based on your words. If you say "a chef," it might draw a man; if you say "a nurse," it might draw a woman. This is how modern Text-to-Image (T2I) models work: they are incredibly smart artists that turn sentences into stunning visuals. But here's the catch: these artists have been trained on a giant pile of old pictures from the internet, and that pile is full of stereotypes. For a long time, scientists have been checking if these digital painters get the "who" wrong—like drawing only men as doctors or only women as nurses. They've been looking at the static cast of characters in the picture. But what if the real problem isn't just who is in the room, but what they are doing? What if the camera decides that when things go wrong, only certain people are brave enough to help, while others just stand around watching? This is the new frontier of bias: not just who appears, but who acts, who reacts, and who gets to be the hero.

Enter a team of researchers from Fudan University who decided to put these AI artists to the ultimate test: the emergency room. They coined a new term, EmergencyBias, to describe a sneaky kind of prejudice that happens when the AI is asked to draw a crisis, like a car crash or a fire. They wanted to see if the AI would automatically decide that men are the ones who rush to save the day, while women, older people, or people with darker skin tones are the ones who need saving or just stand by.

To find out, they didn't just ask the AI to draw a random scene; they set up a massive, systematic experiment. They created six different high-stakes scenarios, ranging from someone drowning in a river to a subway platform accident. They asked seven of the world's most advanced AI models to draw these scenes. First, they gave the models "blank" prompts—just "a person falls on a subway platform"—to see who the AI defaulted to. Then, they gave "controlled" prompts, explicitly telling the AI, "Draw a female bystander" or "Draw an older person in crisis," to see if the AI would still treat them differently in its actions.

The results were a bit like watching a magic trick where the magician accidentally reveals the secret: the AI models were indeed biased, and the bias got worse when the situation got scary. When the prompts were blank, the AI overwhelmingly chose men to be the ones in danger and the ones doing the helping, while older people were almost invisible, appearing in less than 8% of the "in crisis" roles. But the real shocker came when they controlled the demographics. Even when they explicitly told the AI, "Draw a female bystander," the model still made her less likely to actually help compared to a male bystander. The AI seemed to have a hidden script where men are the active rescuers and women are the passive observers, regardless of what the prompt said. This wasn't just a small glitch; the difference in "helping" behavior between genders was the most pronounced bias they found.

The researchers also tested a new, lightweight fix they called ActionAlign. Think of it like a tiny, invisible coach whispering instructions to the AI artist right before it starts painting, nudging it to break its bad habits without erasing its artistic talent. They compared this to a more obvious method of just adding "ethical" words to the prompt (like saying "be fair"). The results showed that ActionAlign was much better at fixing the behavior. On one of the models, it reduced the bias in helping behavior by a massive 78.14% while barely touching the quality of the image. This suggests that you can't just tell an AI to "be nice" with a few words; you sometimes need to retrain its internal sense of how different people should act in a crisis.

In short, this paper suggests that our AI artists aren't just copying who is in the room; they are also copying who they think should be the hero. Even when we force them to draw a specific person, they still struggle to imagine that person taking action. The good news is that with a little bit of smart tuning, we might be able to teach these models to write a fairer script for everyone, ensuring that in the digital emergencies of the future, the heroes look like all of us.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →