High-Risk Memories? Comparative audit of the representation of Second World War atrocities in Ukraine by generative AI applications
This paper empirically audits how generative AI applications handle high-risk memories of Second World War atrocities in Ukraine, identifying significant risks of historical misrepresentation, hallucinations, and inconsistent moralization that could distort collective memory.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a new, incredibly fast librarian who can write stories about history in seconds. You ask, "What happened in Ukraine during World War II?" and this librarian instantly writes a detailed page for you. Sounds helpful, right?
This paper is about testing three of these "super-librarians" (the AI chatbots Bing/Copilot, Google Bard, and ChatGPT) to see if they are actually telling the truth about some of the most painful and controversial events in history: the atrocities committed in Ukraine during World War II.
Here is the breakdown of their findings, using some everyday analogies:
1. The Problem: The "Guessing Game" Librarian
Human memory is like a library where books are carefully cataloged, checked, and debated by historians. AI, however, doesn't "remember" facts the way we do. It's more like a super-fast parrot that has read millions of books and learned to guess the next word in a sentence based on probability.
- The Risk: If the parrot has read a lot of books that say "X happened" and a few books that say "Y happened," it might just guess a mix of the two, or invent a third option that sounds plausible but is completely made up. This is called a "hallucination."
- The Danger: When it comes to "high-risk memories" (like mass murders, the Holocaust, or ethnic cleansing), getting the facts wrong isn't just a typo; it can be used to hurt people, deny crimes, or stir up political hatred.
2. The Test: Asking the Same Question in Three Languages
The researchers treated the AI like a student taking a test. They asked 74 different questions about WWII atrocities in Ukraine (covering victims who were Jewish, Polish, and Ukrainian).
- They asked the same questions in English, Ukrainian, and Russian.
- They compared the AI's answers against a "gold standard" created by real human historians.
3. The Results: The Librarian is Failing the Test
A. The Accuracy Score: A C-Grade Student
Overall, the AI got the facts right only about 50% of the time.
- The "Goldilocks" Effect: The AI was much better at answering questions in English (the language with the most training data) but struggled significantly in Ukrainian and Russian.
- The Analogy: Imagine a chef who is amazing at making Italian pizza because they have a million recipes, but when you ask for a traditional Ukrainian dish, they guess the ingredients and serve you a pizza with pickles on it.
- Specific Failures:
- Bard (Google): Was the worst at "hallucinating." It invented fake dates, fake numbers of victims, and even made up quotes from people who never existed. In Ukrainian, it was particularly unreliable.
- Bing (Microsoft): Was okay with general questions but got lost when asked about specific, lesser-known massacres.
- ChatGPT: Was the most consistent, but it still got half the facts wrong.
B. The "Moralizing" Problem: The Preachy Robot
The researchers also looked at how the AI talked about these tragedies.
- The Issue: Sometimes the AI acts like a moral judge. It would say things like, "This was a dark chapter in history, and we must learn from it," or "We must condemn this."
- The Inconsistency: This was the weirdest part. The AI would be very preachy and moralizing in one language, but completely silent in another.
- Example: If you asked in English, ChatGPT might lecture you on the importance of remembering. If you asked the same question in Ukrainian, it might just give a dry list of facts.
- The Metaphor: Imagine a tour guide at a museum. In one room, they stop and give a 5-minute speech about the tragedy and what we should learn. In the next room, they just whisper the facts and walk away. It's confusing and inconsistent.
4. Why This Matters
The paper warns us that we are handing over the job of "remembering history" to machines that don't actually understand what history means.
- The "Fake Past": The AI isn't just making mistakes; it's creating a "past that never existed." It can generate a version of history where the facts are slightly off, or where the moral lessons are applied unevenly.
- The Language Trap: Because the AI is trained mostly on English data, it knows the "Global North" version of history better than the local version. For people in Ukraine or Russia trying to learn about their own history, the AI is often giving them a distorted, foreign version of their own past.
The Bottom Line
If you ask an AI about the Holocaust or WWII in Ukraine, do not trust it blindly.
- It might get the numbers wrong.
- It might invent fake stories.
- It might lecture you on morals in one language but stay silent in another.
The authors suggest that until these AI "librarians" are fixed, we need to be very careful. We might need to tell them, "If you don't know the answer, just say 'I don't know' instead of making something up," and we need to teach them how to be consistent moral guides, not just random guessers.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.