← Latest papers
💬 NLP

TimeCapsule: Generative Hallucination as a Method for Historical Sensemaking

The paper introduces TimeCapsule, a generative model trained exclusively on Victorian texts that leverages its temporal isolation to transform modern hallucinations into interpretive probes for historical sensemaking, effectively challenging contemporary perceptions of authenticity while demonstrating superior perplexity on period-specific prose.

Original authors: Hayk Grigorian, Hamed Yaghoobian

Published 2026-07-29
📖 7 min read🧠 Deep dive

Original authors: Hayk Grigorian, Hamed Yaghoobian

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to understand the past, but every time you ask a question, the answer comes back with a spoiler alert from the future. This is the problem with today's most powerful computer brains, known as Large Language Models (LLMs). These systems are like time-travelers who have read every book in the library of human history, including the ones that haven't been written yet. Because they are trained on everything from the internet today, they carry modern knowledge, modern ethics, and modern facts in their "minds." When you ask them to pretend to be someone from the 1800s, they often slip up, accidentally mentioning things like airplanes or the internet, or judging historical events with today's moral compass. They are great at mimicking the style of the past, but they can't truly inhabit it because they know too much.

To fix this, researchers are exploring a new idea: what if we built an AI that is deliberately ignorant? Instead of feeding it the whole internet, what if we fed it only books, newspapers, and letters from a specific time period, and then cut the power so it could never learn anything that happened after that date? This paper, titled "TimeCapsule," dives into this concept. It treats "not knowing" not as a bug, but as a feature. By forcing a computer to think only with the tools, words, and ideas available in the 19th century, the researchers hope to create a machine that can genuinely simulate how people back then made sense of the world—even when that world was confusing or full of things the computer couldn't possibly understand.


The Time-Traveling Computer That Forgets the Future

Meet TimeCapsule. It's a 1.2-billion-parameter computer model (think of it as a digital brain with about 1.2 billion tiny connections) that was built with a very strict rule: it was only allowed to read texts from 1800 to 1875.

The creators, Hayk Grigorian and Hamed Yaghoobian from Muhlenberg College, wanted to see what happens when you build an AI that has an "epistemological event horizon." That's a fancy way of saying the model has a hard wall in its memory. Everything before 1875 is its world; everything after 1875 is a blank void. It doesn't just pretend it doesn't know about the future; it literally cannot know.

The Experiment: A Victorian Brain in a Digital Body

To build this, the team gathered a massive library of 136,302 documents from the 19th century. This included everything from parliamentary laws and medical journals to poetry and fiction. They cleaned up the text just enough to make it readable by a computer but kept the old-fashioned spellings and quirks. Then, they trained the model exclusively on this data.

The result? A computer that speaks like a Victorian but thinks like one, too.

The "Hallucination" Twist
Usually, when an AI makes up a fact, we call it a "hallucination" and treat it as a mistake. But TimeCapsule does something different. When you ask it about something that didn't exist in 1875—like a computer or an airplane—it doesn't just say "I don't know." Instead, it tries to figure out what those things must be using only the concepts it has.

For example, when asked to describe a "computer," TimeCapsule didn't say "a machine that calculates." Since the word "computer" in 1875 referred to a person who did math for insurance companies, and since the model was trained on medical texts, it described a computer as a "hypertrophied lung" (an enlarged, diseased lung). It connected the dots between "calculation," "statistics," and "vital signs" to create a weird, but historically logical, explanation.

The researchers call this "ontological repair." Instead of failing, the model is creatively trying to fit a modern object into a 19th-century worldview. It's like if you showed a person from 1800 a smartphone; they might describe it as a "magic mirror that holds the voices of the dead" because that's the only vocabulary they have to make sense of it.

The Numbers: How Good Is It?

The team tested TimeCapsule to see how well it understood the language of the 1800s compared to modern models.

  • They measured something called perplexity (a score that tells you how surprised a model is by the text it's reading; lower is better).
  • A standard modern model (GPT-2) got a score of 68.83 on Victorian text.
  • TimeCapsule got a score of 37.59.
  • This is a 45.4% reduction in confusion, meaning TimeCapsule is much more fluent in the specific language of the era than even larger, smarter modern models.

However, the paper notes that if you just want the lowest possible score, a massive modern model like Mistral-7B can get an even lower score (16.50). But that's because the modern model knows the future! It's cheating. TimeCapsule's "lower" score is actually a sign of its honesty: it is truly speaking the language of the past, not just guessing based on modern knowledge.

The "Uncanny Valley" of History

The most surprising part of the study happened when the researchers asked two experts in literature and history to guess which texts were written by humans and which were written by TimeCapsule.

The results were a bit spooky.

  • The experts misclassified about 40% of the real Victorian texts as being written by the machine.
  • One expert thought a real passage by Charles Dickens was fake because it used the phrase "brick-and-mortar," which the expert mistakenly thought was a modern term (it wasn't!).
  • Another expert thought a real passage was authentic simply because it was "boring as hell," which turned out to be exactly what the AI produced. The AI was so good at mimicking the mundane, everyday writing of the 1800s that it fooled the experts.

This created a "crisis of authenticity." The AI was so good at being average that it made the real, interesting, or even the real, boring stuff look fake. It suggests that in the age of AI, we might lose our ability to tell what is truly human history and what is a digital imitation.

The Ethics of "Bad" AI

The paper also tackles a tricky ethical question. Modern AI is usually trained to be "safe" and polite, refusing to say racist or imperialist things. But TimeCapsule was not trained to be safe. It was trained to be honest to its time period.

Because the 1800s were a time of empires and strict gender roles, TimeCapsule reflects those biases.

  • In the model's "mind," the word "progress" is tightly linked to words like "empire," "conquest," and "dominion."
  • In a modern model, "progress" is linked to "invention" and "improvement."

The researchers argue that this is actually good for historians. If you "fix" the AI to remove these biases, you erase the history of how people actually thought. By keeping the biases, TimeCapsule acts as a mirror, showing us exactly how the 19th century viewed the world, including its ugly parts. It's not a tool for chatting with a friendly Victorian; it's a tool for studying the Victorian mind, flaws and all.

The Takeaway

TimeCapsule suggests that to truly understand the past, we might need to build computers that are limited. By cutting off their access to the future, we force them to think like people from the past. When they get things "wrong" about modern technology, they aren't failing; they are showing us how a 19th-century mind would try to make sense of our world.

The paper concludes that this "generative hallucination" isn't a bug to be fixed, but a new way to do historical research. It turns the AI's ignorance into a window, letting us see the hidden logic of a time that no longer exists. As the authors suggest, the model doesn't just mimic history; it reconstructs it, one "hypertrophied lung" at a time.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →