← Latest papers
💻 computer science

What Gets Lost When Memory Becomes Media? Evaluating AI-Generated Oral History Visualization

This paper proposes a failure-mode-based evaluation framework and an empirical analysis to assess AI-generated oral history visualizations, revealing that narrative preservation often conflicts with scene planning and demonstrating that the source testimony's narrative structure strength is the primary predictor of this conflict, thereby informing a routing protocol for selecting between multi-agent and single-summarization pipelines.

Original authors: Kwangsuk Park, Jaehyun Koo, Jiyeon Lee, Anjung Tan, Hyoungchul Park

Published 2026-07-29
📖 6 min read🧠 Deep dive

Original authors: Kwangsuk Park, Jaehyun Koo, Jiyeon Lee, Anjung Tan, Hyoungchul Park

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a time traveler with a camera, but instead of snapping photos of the present, you are trying to capture memories from someone else's past. This is the tricky world of oral history, where people tell stories about their lives, often from decades ago. The challenge is that human memory is messy, emotional, and full of jumps between "what happened" and "how it felt." When we try to turn these spoken stories into pictures or videos, we have to make a magic trick: we have to turn a first-person voice ("I was scared") into a third-person scene ("A person is running").

Enter Generative AI, the digital artist that can create images from text. But here's the rub: AI is great at making things look pretty, but it's also great at making things up. If you ask an AI to draw a "sad refugee," it might just show you a generic, sad-looking person in a dark room because that's what it has seen a million times in other movies. But the real story might be about a specific red backpack, a broken watch, or a specific street corner in 1980s Seoul. The big question for scientists is: When we let AI turn a memory into a picture, what parts of the real story get lost, and what parts get made up?

This paper dives into that exact question. The researchers built two different "AI chefs" to cook up visual stories from 82 real interviews with people from diaspora communities (people who moved away from their home countries). One chef, called SSP, tries to summarize the whole story into one smooth, continuous paragraph before drawing. The other chef, MAS, breaks the story down into tiny, separate scenes first, like a comic book, before drawing. They wanted to see which chef kept the "flavor" of the real memory and which one just served up a generic, tasty-looking meal that didn't taste like the original ingredients.

The Great AI Art Showdown

The researchers set up a massive taste test. They took 82 interviews, each about an hour long, and picked a key three-minute segment from each. Then, they fed these stories to both the SSP (Single Summarization Pipeline) and the MAS (Multi-Agent Scene-decomposition Pipeline) systems. Both systems were tasked with creating a sequence of 6 images that told the story.

To judge the results, they didn't just ask, "Is this pretty?" They created a strict scorecard with 15 different metrics based on the rules of oral history. They looked for three main ways the AI could mess up:

  1. Transition Dissolution: Did the AI smooth over the jumps in the story so much that the distinct moments blended into a boring blur?
  2. Genericization: Did the AI replace specific details (like a "blue 1990s jacket") with clichés (like a "generic sad person")?
  3. Preservation Damage: Did the AI change the facts, the person's identity, or the timeline?

The Surprising Results: A Trade-Off, Not a Winner

Here is the twist: There was no single "best" system. Instead, the researchers found a structural conflict, like a see-saw.

  • The MAS Chef (The Comic Book Approach): This system was amazing at breaking the story into distinct scenes. It kept the "texture" of the memory alive, avoiding generic clichés and showing specific details. However, in doing so, it often weakened the overall flow of the story. It was like having a comic book with great individual panels, but the story jumped around so much you forgot how the characters got from point A to point B.
  • The SSP Chef (The Summary Approach): This system kept the narrative flow perfect. The story moved smoothly from start to finish, preserving the big picture and the emotional arc. But, because it tried to keep everything smooth, it sometimes lost the specific sensory details, making the images feel a bit more generic or "stock-photo" like.

The data showed that in about 68.6% of the cases, you had to choose: you could have better scene details (MAS) or better story flow (SSP), but rarely both at the same time.

The Secret Ingredient: How the Story is Told

The most exciting discovery was why this trade-off happened. The researchers found that the answer lay in the original story itself.

They measured the "narrative-structure strength" of the interview.

  • Strong Stories: If the person telling the story had a very clear, logical chain of events (like "I left home, then I got lost, then I found help"), the SSP system was already doing a great job. If you tried to use the MAS system on these strong stories, it actually hurt the result. It was like taking a perfectly written novel and chopping it up into tiny, disconnected sentences; you just broke the magic.
  • Weak Stories: If the original story was a bit jumpy, emotional, or hard to follow, the SSP system struggled to make sense of it. In these cases, the MAS system was a hero. By breaking the story into small, manageable scenes, it actually helped reconstruct the missing structure, making both the details and the flow better.

The Solution: A Smart Switch

So, what's the takeaway? The paper suggests a "routing protocol," which is basically a smart switch for AI systems. Before you generate images, you should first run the story through a quick check to see how strong its structure is.

  • If the story is strong and clear, use the SSP (Summary) system to keep the flow smooth.
  • If the story is jumpy or complex, use the MAS (Scene-decomposition) system to help organize the chaos.

The researchers tested this on their 82 interviews and found that this simple rule could predict which system would work best. They didn't find a "magic bullet" that solves everything, but they did find a way to manage the tension between keeping the story true and making it look good.

Why This Matters

This isn't just about making pretty pictures. For communities that have been displaced or forced to leave their homes, these visual stories are a way to preserve history that text alone can't capture. If an AI turns a specific, painful memory of a border crossing into a generic image of a sad person, it loses the truth of that person's experience.

The paper suggests that we can't just let AI do whatever it wants. We need to be careful, to check the "structure" of the story first, and to choose our tools wisely. It's a reminder that in the age of AI, the human element—the messy, specific, sometimes confusing way we tell our stories—is still the most important part of the recipe. The AI is just the chef; we have to be the ones who decide which recipe to cook.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →