← Latest papers
🤖 AI

Investigating Social Bias in Narrative Image Generation

This paper demonstrates that social biases in text-to-image generation models are not only more prevalent but also more explicitly manifested in narrative formats like storyboards and comics compared to single images, highlighting the critical need to evaluate these systems across diverse visual contexts.

Original authors: Junyeong Park, Sowon Min, Euna Jang, Soobin Kim, Jiho Jin, Hyunseung Lim, Gahyeon Bae, Hwajung Hong

Published 2026-08-04
📖 1 min read☕ Coffee break read

Original authors: Junyeong Park, Sowon Min, Euna Jang, Soobin Kim, Jiho Jin, Hyunseung Lim, Gahyeon Bae, Hwajung Hong

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Technical Summary: Investigating Social Bias in Narrative Image Generation

Problem Statement

While Text-to-Image (T2I) generation models are increasingly deployed in media, education, and design, concerns persist regarding their tendency to reproduce social biases. Prior research has extensively documented that T2I models associate specific demographics with professions, traits, and social roles. However, existing bias evaluations predominantly focus on single-shot photorealistic image generation. This narrow scope leaves a critical gap: it remains unclear whether and how social biases manifest in narrative visual formats (such as storyboards and comics), where meaning is constructed through character continuity, event sequencing, temporal progression, and the interplay between visual and textual elements. The authors posit that biases which remain subtle or invisible in static photos may become explicit or amplified in sequential storytelling formats.

Methodology

The study adapts the BBG (Bias Benchmark for Generation) framework, originally designed for text generation, to evaluate image generation models. The methodology involves the following components:

  • Prompt Construction: The authors selected 140 seed prompts from the BBG dataset, covering seven social bias categories: age, disability, gender, physical appearance, race/nationality, religion, and socioeconomic status (SES). These prompts were intentionally contextually ambiguous, leaving key attributes (e.g., which character possesses a specific trait or outcome) unspecified. The dataset includes 70 English and 70 Korean prompts to investigate cross-linguistic bias.
  • Models Evaluated: Six state-of-the-art T2I models were tested:
    • Proprietary: GPT-Image-1.5, GPT-Image-2, Gemini-3.1-Flash-Image, and Gemini-3-Pro-Image.
    • Open-Source: FLUX.1-Dev and SDXL.
  • Generation Settings: For each prompt, outputs were generated in three distinct formats:
    1. Photorealistic Image: A single static image.
    2. Storyboard: Rough sketches focusing on shot framing and camera coverage without text.
    3. Four-Panel Comic: A sequential narrative integrating visual elements with speech bubbles, captions, and narration.
  • Evaluation Protocol: Four authors manually annotated the generated images (2.4K total) based on the implied answers to the BBG questions. Outputs were labeled as biased, counter-biased, or neutral. The study also employed thematic coding to identify recurring visual and narrative patterns. A case study on video generation (Sora-Pro-2 and Veo-3.1) was conducted to extend findings to temporal media.

Key Results

Quantitative Findings

  • Narrative Formats Amplify Bias: Proprietary models generated significantly more biased outputs in narrative formats compared to photos. On average, proprietary models produced 25.9% biased outputs in photo generation. This rate increased by 9.6 percentage points (pp) in storyboard generation and 18.2 pp in comic generation.
  • Linguistic Disparities: Bias prevalence was higher in Korean prompts than in English prompts, with Korean settings showing a 7.2 pp higher biased-output rate on average. The largest gap appeared in photo generation (e.g., Nano Banana 2 showed a 17.4 pp increase).
  • Open Model Performance: Open models (FLUX.1-Dev and SDXL) generated neutral outputs in 93.6% of cases on average. However, the authors attribute this to a weak instruction-following capability rather than effective bias mitigation, noting these models often failed to incorporate demographic cues or generated culturally mismatched content.

Qualitative Findings

  • Explicit vs. Implicit Bias: In photo generation, biases are often encoded through subtle visual cues (e.g., specific clothing, accessories, or expressions). In contrast, storyboards and comics reveal biases more explicitly through event sequencing, character positioning, narrative resolution, and textual elements (e.g., speech bubbles explicitly stating stereotypes).
  • Persistence of Visual Associations: Stereotypical visual associations (e.g., "creative" individuals depicted with casual/hippie fashion or lightbulb symbols) persisted across all formats, even when the output style changed.
  • Layered Stereotypes: Narrative formats exposed biases beyond simple role assignment. For instance, the reasoning for a character's failure or dependence differed by stereotype (e.g., a female nurse's dependence framed as emotional vulnerability vs. a blind person's dependence framed as a practical need).
  • Surface-Level Mitigation: Models sometimes attempted to avoid bias by introducing anti-stereotypical elements or moralizing endings (e.g., reconciling conflicting neighbors). However, the authors note this does not necessarily remove underlying associations, as subtle visual cues may still reinforce stereotypes.
  • Cultural Misalignment: In Korean contexts, models frequently failed to align with cultural nuances, generating text in English or mixed languages, or producing Western/overly traditional East Asian imagery that did not match the prompt's cultural context.

Significance and Contributions

The paper claims three primary contributions:

  1. First Study of Narrative Bias: It presents the first systematic investigation of social bias in narrative image generation formats, comparing photorealistic images with storyboards and four-panel comics.
  2. Evidence of Format-Dependent Bias: It demonstrates that proprietary models exhibit significantly higher bias in narrative formats (storyboards and comics) than in single-image generation, suggesting that single-image evaluations are insufficient for capturing the full scope of model bias.
  3. Qualitative Mechanisms: It provides qualitative evidence on how biases manifest differently across formats—shifting from subtle visual cues in photos to explicit narrative and textual reinforcement in comics.

The authors conclude that narrative image formats serve as a stronger probe for model bias than photo generation alone. They argue that evaluating T2I systems requires moving beyond single-image outputs to account for diverse generation formats, as biases that are less visible in static images may surface and become more harmful in sequential storytelling contexts. The study also highlights the need to evaluate multilingual and cross-cultural capabilities, noting that current models often fail to ground prompts in their specific cultural contexts.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →