Whose story gets told? ChatGPT, collective memory, and ethical narratives in the age of Artificial Intelligence
This paper argues that while AI tools like ChatGPT can democratize access to historical knowledge, they risk distorting collective memory through algorithmic biases and factual inconsistencies, necessitating ethically grounded human oversight to ensure historical narratives remain products of critical reflection rather than automated representation.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the digital age, our understanding of the past is no longer stored solely in dusty archives or the memories of elders. It lives in the vast, interconnected networks of the internet, where artificial intelligence systems act as new gatekeepers. These systems, known as large language models, are trained on enormous collections of human writing. They do not simply retrieve facts; they synthesize information to create stories and answers. This raises a profound question for historians and sociologists: when we ask a machine to tell us about a war or a conflict, whose version of history do we get? The concept of "collective memory" refers to the shared story a group tells itself about its past, a story that shapes who they are. If an algorithm decides which details to include, which to leave out, or how to frame a tragedy, it effectively reshapes that shared memory. The concern is not just about getting a date wrong, but about whether these tools might smooth over the rough, painful edges of history, turning complex human conflicts into neutral, sanitized summaries that serve the interests of the technology's creators rather than the truth.
A team of researchers set out to test exactly how these artificial intelligence tools handle the messy reality of history. They focused on a specific, powerful version of the technology called ChatGPT, asking it to recount the stories of four major multinational conflicts: the wars following the breakup of Yugoslavia, the long-standing conflict between Israel and Palestine, the colonial wars fought by Portugal in Africa, and the Vietnam War. To see if the machine told the same story to everyone, the researchers did not just ask in English. They posed the same questions in the native languages of the people involved in each conflict, such as Croatian, Serbian, Arabic, Hebrew, Portuguese, and Vietnamese. They treated the chatbot not just as a search engine, but as an active participant in a conversation, watching closely to see how it constructed its answers, what it chose to remember, and what it seemed to forget.
The results revealed that the machine is far from a neutral observer. Instead of providing a single, consistent version of history, the AI produced different narratives depending on the language used to ask the question. When the researchers asked about the same event in different languages, the answers often contradicted one another in significant ways. For instance, when asked to identify the most significant final event of the wars in the former Yugoslavia, the English version pointed to a peace agreement in North Macedonia from 2001. The Croatian version stopped at a 1995 peace deal, while the Serbian version extended the timeline to the 2022 trial of a military leader. These were not just minor variations in wording; they were fundamentally different stories about when the conflict ended and what mattered most. The study found that the AI often mirrored the official national narratives of the language it was speaking, effectively reinforcing the specific political perspectives of that region rather than offering a balanced, global view.
Beyond these inconsistencies, the researchers identified a pattern of errors and omissions that distorted the historical record. The AI frequently got dates and names wrong, sometimes inventing events that never happened or mixing up the order of battles. More troubling was what the machine left out. In its attempts to be helpful, it often skipped over crucial details, such as the specific victims of atrocities, the role of certain military forces, or the complex causes of a war. In one instance, the AI described the wars in Portuguese colonies as a single, simple story, merging distinct struggles in different countries into one flat narrative. This "simplification" stripped away the unique political and social realities of each conflict. The study also noted a tendency toward "neutralization," where the AI used vague language to avoid assigning blame. It would describe a genocide as a "tragic event" or a "crime" without naming the perpetrators, effectively washing away the specific responsibility of the actors involved.
The researchers concluded that these large language models are not merely reflecting the past; they are actively reshaping it. Because the AI learns from the data it is fed, it inherits the biases and blind spots of that data. When it generates a story about a war, it is not pulling from a perfect, objective record. Instead, it is predicting the next likely words based on patterns in its training, often prioritizing the most common or dominant voices in its database. This means that for users who rely on these tools for historical knowledge, the past becomes a product of algorithmic logic rather than critical reflection. The study suggests that without human oversight and a strong ethical framework, we risk losing the complexity of our shared history to a machine that prefers smooth, simplified, and sometimes contradictory stories. The power to decide which stories get told, and how they are told, has shifted from historians and communities to the code and data that drive these digital systems.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.