Auditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes
This paper presents the first quantified, scene-level audit of an LLM-generated autobiography against a ground-truth record, revealing a 96.7% verification failure rate dominated by "grounded drift" (invented scenes using real people and settings) and demonstrating that while grounding generation in the subject's actual corpus improves accuracy, substantial hallucination persists.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the quiet corners of artificial intelligence research, a new kind of experiment is taking place. It asks a question that feels simple but cuts deep: if you ask a computer to tell a person's life story, how much of that story is actually true? For years, scientists have known that these systems, often called large language models, can sometimes make things up. They might invent a fact about a historical event or claim a scientific study exists when it does not. This is usually called a "hallucination," a word that suggests the machine is seeing things that aren't there. But when the machine is asked to write a memoir, a biography, or a personal diary, the stakes change. The story is no longer about general knowledge; it is about a specific human life. If the machine invents a childhood memory or a career milestone, it is not just making a mistake; it is rewriting a person's history. This is where the concept of "confabulation" becomes important. Unlike a random error, confabulation is a fluent, confident fabrication. The machine fills in the gaps with a story that sounds perfect, feels real, and fits the tone of a memoir, even though the events never happened. As we move toward a future where digital versions of ourselves might write our biographies, we need to know if those stories are grounded in reality or built on a foundation of fiction.
A researcher named Heather Renze decided to find out by becoming her own subject. She commissioned a year-long book, a "page-a-day" diary, to be written by an artificial intelligence. The goal was to create a gift book of 366 daily entries, each telling a short story from her life in the first person. The process was straightforward: the computer was given a template, two example days to show it the style, and a single quote for each day to serve as a prompt. Crucially, the computer was not given her actual diaries, her memoirs, or her personal records. It had to rely on its own internal knowledge to fill in the details of her life. Before the book was ever published, Renze decided to audit every single entry. She treated the project like a scientific experiment, checking every anecdote against a massive collection of her own verified documents, including her published memoirs, her professional records, and her personal archives. She wanted to see, with absolute precision, how many of the stories the machine told were true.
The results were stark. Out of the 366 days in the book, only 12 entries contained a scene that could be positively confirmed by the records. That means that for 354 days, the story the computer told could not be verified as true. In the language of the study, this is a failure rate of nearly 97 percent. The machine did not just get small details wrong; it constructed entire scenes that had no basis in reality. In some cases, the computer was so confident that it invented events that were actually impossible or directly contradicted by the facts. For example, the machine claimed she had children when she did not, or that she survived a specific car accident that never happened. These were not vague guesses; they were specific, named events that the records proved were false. About 5 percent of the days fell into this category of being actively contradicted by the truth.
The most common type of error, however, was something more subtle. The machine would get the setting right but invent the story. It would correctly identify that she worked at a specific company or lived in a certain city, but then it would create a fictional scene around that fact. It might say she had a dramatic meeting in an office that never happened, or that she had a specific conversation with a real colleague that never took place. This is what the researchers call "grounded drift." The machine was anchored to the truth in some ways, but it drifted away from it to create a narrative that sounded better or more dramatic. It knew the names of her employers and the cities she lived in, but it used those real details as a stage for a play that never occurred. This is particularly dangerous because it makes the lies harder to spot. A story that gets the facts right but invents the emotional core feels more trustworthy than one that gets everything wrong.
To be sure of these findings, the study included a second layer of checking. The original audit was performed by an artificial intelligence system, which raised the question of whether the machine was judging itself fairly. To test this, two independent raters from a different model family than the original auditor reviewed a random sample of 60 days using the exact same rules. They agreed with the original findings almost completely. When asked to decide if a story was true or not, they matched the original verdict 95 to 98 percent of the time. This confirmed that the high failure rate was not a mistake in the checking process; the stories really were mostly false. The researchers also tried to see if newer, more powerful computer models would do better. They ran the same prompts through the latest versions of the technology, but the result was the same: even though the models had access to her digital records, they were only loosely directed to use them and still failed to tell the truth, inventing stories for every single day they were asked to write.
The study did offer a path forward, though it was not a perfect fix. When the researchers explicitly and specifically instructed the computer to ground its output in the subject's existing digital records, the failure rate dropped significantly. The number of days with verified scenes rose, and the number of completely false stories fell. However, even with this help, the machine still struggled. It continued to invent details and drift away from the facts, suggesting that simply giving the machine the right books is not enough to make it tell the truth. The machine still seems to have a strong urge to create a narrative, even when the facts are right in front of it.
This experiment reveals a fundamental limit in how we might use artificial intelligence to tell our stories. If we ask a machine to write our biography without giving it our own history, it will likely fill the gaps with fiction. It will create a version of us that is coherent and engaging, but it will be a version that never existed. The study suggests that for any system that claims to simulate a person's life, we cannot trust the details of their daily experiences unless those details are checked against a verified record. The machine can mimic the voice of a person, but it cannot yet remember their life. The trust we place in a digital twin or a synthetic biography must be limited to the surface level of attitudes and general facts, because the specific moments that make up a life, the scenes, the conversations, the small dramas, are where the machine most often fails. The lesson is clear: if we want a true story, we must provide the facts, and we must check the work, because the machine will happily invent a life for us if we let it.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.