← Latest papers
💬 NLP

TempViz: On the Evaluation of Temporal Knowledge in Text-to-Image Models

This paper introduces TempViz, the first dataset designed to holistically evaluate temporal knowledge in text-to-image models, revealing that current models exhibit weak temporal competence and that existing automated evaluation methods fail to reliably assess their ability to handle time-dependent visual cues.

Original authors: Carolin Holtermann, Nina Krebs, Anne Lauscher

Published 2026-01-22
📖 4 min read☕ Coffee break read

Original authors: Carolin Holtermann, Nina Krebs, Anne Lauscher

Original paper dedicated to the public domain under CC0 1.0 (http://creativecommons.org/publicdomain/zero/1.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a magical painter that can create pictures from your words. If you say, "Draw a dog," it paints a dog. But what happens if you say, "Draw a baby dog" or "Draw a very old dog"? Or if you ask for "A forest in winter" versus "A forest in summer"?

This paper, called TEMPVIZ, is like a report card for these magical painters. The researchers wanted to see if these AI artists actually understand how time changes the way things look.

Here is the story of their findings, broken down simply:

1. The Problem: The AI is "Time-Blind"

We know that time changes the world. A building might be standing in 1990 but destroyed in 2005. A map might show different borders in 1938 compared to today.
The researchers built a giant test called TEMPVIZ. Think of it as a massive quiz with nearly 8,000 questions. They asked five different AI painters to draw things like:

  • Animals: A baby alpaca vs. an elderly alpaca.
  • Landscapes: A lake in January (frozen) vs. July (green).
  • Buildings: The Berlin Wall before it fell vs. after.
  • Maps: Europe's borders in 1938 vs. today.
  • Art: A painting in the style of 1732 vs. 1880.

2. The Results: The Painters Struggle with Time

When humans looked at the pictures the AI made, the news wasn't great.

  • The Score: Even the best AI painter got less than 75% of the time-based details right. In some tricky categories (like historical maps), they got it right only 15% of the time.
  • The Confusion: Sometimes the AI drew a perfect dog, but it looked like a 10-year-old dog when you asked for a 1-month-old puppy. Sometimes it drew a winter scene with green leaves.
  • The Surprise: The AI that was best at making pretty pictures wasn't necessarily the one that understood time best. Being a good artist doesn't mean you are good at understanding history or biology.

3. The Robot Judges Failed, Too

Since checking 8,000 pictures by hand is exhausting, the researchers tried to use other AI tools (automated judges) to grade the pictures for them. They hoped these "robot graders" could spot the mistakes.

  • The Reality: The robot graders were terrible at this. They couldn't reliably tell if a picture was "winter" or "summer," or if a building was "before" or "after" a disaster.
  • The Metaphor: It's like asking a robot to grade a history essay by only looking at the font size and paper quality, ignoring the actual facts written inside. The automated tools missed the subtle clues that humans could see.

4. What They Did (The "TEMPVIZ" Toolkit)

To prove this, the team created a new toolkit:

  • The Prompts: They wrote thousands of specific instructions (e.g., "Draw a map of Europe in 1938").
  • The Reference Photos: For many categories, they gathered real photos to show what the "correct" answer should look like (like a teacher's answer key).
  • The Human Check: They had real people look at the AI's work to see if it followed the time instructions.

The Bottom Line

The paper concludes that while our AI painters are getting very good at making images, they are still quite confused about time. They don't naturally "know" that seasons change, animals age, or borders shift.

Furthermore, we currently don't have a good way to automatically test if they are getting the time right. The tools we use to check AI quality are missing the "time" part of the picture. The researchers hope this new test (TEMPVIZ) will help scientists build better AI that truly understands the flow of time.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →