← Latest papers
💬 NLP

LexKairos: Benchmarking Legal Temporal Capabilities in LLMs

This paper introduces LexKairos, a comprehensive benchmark designed to evaluate the temporal capabilities of large language models in the Chinese legal context across statutory knowledge, case modeling, and reasoning, revealing that even top-performing models still struggle with precise time-sensitive legal tasks.

Original authors: Chenyang Li, Zejia Feng, Yuqin Huang, Yuxiao Ye, Huiyuan Xie

Published 2026-08-11
📖 4 min read☕ Coffee break read

Original authors: Chenyang Li, Zejia Feng, Yuqin Huang, Yuxiao Ye, Huiyuan Xie

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a world where the law isn't just a giant library of rules, but a living, breathing machine that only works if you turn the right knobs at the exact right moment. In this world, a law might be a powerful tool today, but yesterday it was a broken toy, and tomorrow it might be a completely different machine. This is the realm of Legal Temporal Capabilities. While computers have gotten really good at reading laws and understanding what they say (the "what"), they are still struggling to understand when they apply (the "when"). Think of it like a GPS that knows every street name but gets confused about traffic lights and construction zones that change every hour. If a computer can't tell the difference between a rule that is active now versus one that expired last year, it can't give legal advice without causing a massive mess. This is why researchers are racing to teach Artificial Intelligence how to keep perfect time with the law.

Enter LexKairos, a new "time-travel test" designed by researchers from Oxford, Sun Yat-Sen University, and Tsinghua University to see how well modern AI brains can handle the ticking clock of the Chinese legal system. They didn't just ask the AI to guess; they built a rigorous obstacle course with nine different challenges, ranging from simple memory quizzes about when laws started, to complex puzzles where the AI must mix facts from a court case with strict deadlines from the rulebook. They put eight different AI models through this gauntlet, letting some think silently before answering and others just spit out an immediate guess.

The results? It's a bit of a mixed bag, like a race where the runners are fast but keep tripping over their own shoelaces. The top performers, like Gemini-3-Flash and GPT-5.4, managed to score around 84.57% and 82.75% respectively when allowed to use their "thinking" modes. That sounds impressive, but the researchers found that even the smartest models stumble badly when asked to recall specific details about which version of a law was active on a specific day. It turns out that while AI is great at figuring out the order of events in a story (like who did what first in a court case), it is surprisingly bad at remembering the exact "birth dates" and "expiration dates" of the laws themselves.

One of the most interesting discoveries was how these AI models fail. When asked to identify the correct version of a law, the models didn't just get it wrong randomly; they had distinct personalities in their mistakes. One model, GPT-5.4, was like a shy student who was afraid to guess, often leaving the answer blank or missing the version tag entirely (a "missing tag" error). Another model, LegalOne-8B, was like an overconfident student who made up facts, inventing version numbers that didn't exist (a "hallucinated tag" error). When the researchers turned on the "thinking mode"—letting the models pause and reason through the problem—the shy model started guessing more (reducing its missing tags) but started making up more fake versions, while the overconfident model became more cautious but started missing tags more often. It's as if giving them more time to think didn't fix their memory; it just swapped one type of mistake for another.

Furthermore, the study suggests that simply letting an AI "think" out loud isn't always the magic bullet. For some models, the thinking process was so long and winding that they ran out of space to finish their answer, cutting off right before the solution. The researchers found that by giving the AI a specific, structured checklist of legal steps to follow (a "task-specific prompt"), they could get just as good results but with answers that were up to 3.4 times shorter. This suggests that for legal time-travel, a guided tour might be better than letting the AI wander aimlessly.

Ultimately, LexKairos suggests that while AI is getting better at legal reasoning, the "when" part of the law remains a stubborn challenge. Even the best models are still prone to errors when the stakes involve precise dates and procedural deadlines. The paper doesn't claim this problem is solved; rather, it highlights that we need to build AI that is not just smart, but also incredibly precise with time, because in the legal world, being a day late can mean the difference between justice and a lost case.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →