Time Present and Time Past: Benchmarking Large Language Models on Temporally Evolving Document Understanding
This paper introduces TIDE, a new expert-verified benchmark for evaluating large language models on temporally evolving document understanding, revealing that current models struggle significantly with version resolution and often prioritize confident parametric knowledge over authoritative, time-specific text.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to find the right rule for a game, but the rulebook is a living thing that changes every day. In the world of artificial intelligence, there are two main ways knowledge is usually handled. The first is like a giant encyclopedia: if a fact changes, the old one is simply erased and replaced with the new one. If you ask about the capital of a country that changed its name, the encyclopedia just shows you the new name. The second way, which is much trickier, is like a stack of legal contracts or a video game with patches. Here, an old rule doesn't disappear; it stays perfectly correct for the time it was active, while a new rule takes over later. If you ask, "What was the speed limit in 2015?" you need the 2015 rule, not the 2025 one. If the computer gets this wrong, it's not just a silly mistake; it could be like telling someone to pay the wrong tax or follow a medical guideline that was canceled years ago. This is the challenge of "version resolution": figuring out exactly which version of a rule was in charge on a specific date.
A team of researchers has built a new test called TIDE (Temporal Information Drift in Evolving Documents) to see if today's smartest AI computers can handle this tricky time-traveling logic. They didn't just make up fake questions; they dug into 644 real, official documents from the Bangladesh government, covering customs laws, tax codes, and regulations from 1969 all the way to 2025. These documents are messy, mixing English and Bengali, and they contain a web of amendments where one rule cancels another, which then gets canceled by a third. The researchers created 3,050 questions based on these documents, asking the AI to act like a time-traveling lawyer. They tested nine of the most powerful AI models available, giving them three different ways to find answers: relying only on what they memorized during training (like a closed-book exam), reading the exact right documents provided to them, or searching a database to find the right pages.
The results were a bit of a wake-up call. Even the best AI models, when given the perfect documents to read, only got about 68.5% of the answers right. That means they were still failing more than one out of every three times, even when the answer was right in front of their "eyes." The study found that these models are surprisingly good at finding the right text when they know exactly where to look, but they are terrible at realizing when a text they are reading is actually the wrong version for the date in question. For example, if you give an AI a document that says "The tax is 10%" but the question asks about a time when the tax was 5%, the AI often confidently ignores the document and sticks with its memory or the wrong text. In fact, when the researchers tried to trick the AI with a confident but wrong statement, the models were very good at flagging that something was wrong (detecting errors about 79% of the time), but they were terrible at pinpointing exactly what was wrong (identifying the specific error only about 19% of the time). This means they often sense a problem but confidently blame the wrong detail.
The paper suggests that simply giving AI more text to read or better search tools isn't enough to solve this problem. The models struggle most with tasks that require them to piece together a timeline, like sorting events in order or calculating what the rule was "three years later." They tend to assume that rules only get stricter or change in one direction, missing the fact that laws often get repealed or reverted to old versions. While the researchers showed that having the right documents helps, it doesn't guarantee success. The study concludes that "version resolution" remains a major, unsolved hurdle for AI. Until these models can truly understand the flow of time in legal documents, they might be too risky to trust with real-world tasks like tax filing or customs clearance, where getting the date wrong could lead to serious financial or legal trouble.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.