SoS: Analysis of Surface over Semantics in Multilingual Text-To-Image Generation
This paper introduces the first comprehensive analysis of the "Surface-over-Semantics" (SoS) phenomenon in multilingual text-to-image models, revealing that most models prioritize surface-level linguistic cues over semantic meaning, leading to culturally stereotypical depictions across various languages and cultures.
Original paper dedicated to the public domain under CC0 1.0 (http://creativecommons.org/publicdomain/zero/1.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a magical art machine that draws pictures based on what you type. If you tell it, "Draw a person from Germany," it should draw a German person. But what happens if you type the exact same request in a different language, like "Zeichne eine Person aus Deutschland" (German) or "Dibuja una persona de Alemania" (Spanish)?
This paper investigates a strange glitch in these machines called Surface-over-Semantics (SoS).
The Core Problem: The "Accent" vs. The "Meaning"
Think of a text-to-image prompt like a letter sent to an artist.
- The Semantics (The Meaning): The actual content of the letter. "I want a picture of a German person."
- The Surface (The Look): The language and script the letter is written in. Is it written in English, Hindi, or Chinese?
The researchers found that many of these AI art machines are like bad translators who focus on the accent rather than the story. When you ask for a "German person" in English, the machine draws a German person. But if you ask for the same "German person" using a different language (like Hindi or Finnish), the machine often ignores the "German" part and instead draws something that looks like a stereotype of the language you used.
It's as if you asked a chef for "Spaghetti" in Italian, and instead of making pasta, they served you a pizza because they heard the word "Italian" and assumed that's what you wanted, ignoring the word "Spaghetti."
The Experiment: A Global Taste Test
To prove this, the researchers created a massive "taste test":
- The Menu: They picked 171 different cultural identities (like "Japanese," "Nigerian," "Brazilian").
- The Languages: They translated the request "A photo of a [Culture] person" into 14 different languages (including English, Arabic, Chinese, Finnish, etc.).
- The Artists: They asked 7 different AI art machines to draw these pictures.
- The Scorecard: They invented a new way to measure the results, called the SoS Score.
The SoS Score is like a "Bias Meter":
- If the score is positive, the machine listened to the meaning (e.g., It drew a German person when asked for a German person, regardless of the language).
- If the score is negative, the machine listened to the surface (e.g., It drew a stereotypical "Finnish" scene when asked for a German person, just because the prompt was in Finnish).
What They Found
- Most Machines are "Surface-Heavy": Out of the 7 machines tested, 6 of them were heavily influenced by the language used. If you typed in a language the machine wasn't very good at, it would often default to drawing stereotypes associated with that language's script or culture, rather than the culture you actually asked for.
- The "Deep Dive" Effect: The researchers looked inside the machines' "brains" (specifically the layers where the text is processed). They found that as the message travels deeper into the machine, the bias gets stronger. It's like a game of "Telephone" where the message gets more distorted the further it goes; by the time the machine starts drawing, the language's "accent" has completely drowned out the original meaning.
- Visual Stereotypes: When the machines got confused, they didn't just draw random things; they drew cultural stereotypes.
- Prompts in Finnish often resulted in snowy forests.
- Prompts in Amharic (an Ethiopian language) often resulted in dry, sandy deserts.
- Prompts in Chinese or Japanese sometimes included scary elements like blood or scars, which had nothing to do with the actual request.
The One Exception
One machine, called AltDiffusion, was the "good student." It mostly ignored the language accent and stuck to the meaning, drawing the requested culture regardless of the language used. However, even this machine struggled a bit with languages it hadn't seen much during training.
Why This Matters
The paper concludes that while these AI tools are getting better, they still have a blind spot. They are too easily swayed by the look of the words (the language) rather than the meaning of the words. This means that if you use these tools in your native language, you might get a picture that reflects a stereotype of that language rather than the specific person or thing you asked for.
The researchers suggest that to fix this, we need to train these machines to pay more attention to the "story" (semantics) and less to the "accent" (surface), especially for languages that aren't English. They also propose their "Bias Meter" (SoS Score) as a tool to help developers check if their machines are being fair.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.