Can LLMs Perceive Time? An Empirical Investigation
This empirical study demonstrates that large language models fundamentally lack the ability to perceive or estimate their own inference time, consistently producing inaccurate pre-task predictions, post-hoc recalls, and relative ordering judgments across various tasks and model families due to a disconnect between their propositional knowledge and experiential grounding.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: The "Time-Blind" Chef
Imagine you hire a brilliant, super-fast chef (the AI) to cook a meal. This chef knows everything about food. They know that a steak takes 10 minutes to grill, that baking a cake takes an hour, and that chopping onions is quick. They have read every cookbook ever written.
However, there is a catch: This chef has never actually stood in a kitchen. They have never felt the heat of the stove, never heard the timer beep, and never felt the weight of time passing while they chop. They only know what the books say, not what reality feels like.
This paper investigates what happens when we ask this "Time-Blind" chef: "How long will it take you to cook this specific dish right now?"
The answer? The chef is wildly wrong.
The Four Experiments (The Taste Tests)
The researchers put the AI through four different tests to see how bad its time sense really is.
1. The "Crystal Ball" Test (Pre-Task Estimates)
The Setup: Before the AI starts working, they ask: "How many seconds will this take?"
The Reality: The AI is terrible at this.
- The Analogy: Imagine you ask the chef to make a simple grilled cheese sandwich. The chef, having read about "cooking," guesses it will take 20 minutes. In reality, the sandwich is done in 3 seconds.
- The Result: The AI consistently guesses that tasks will take 4 to 7 times longer than they actually do. If a task takes 10 seconds, the AI thinks it will take a minute or two. It's like guessing a sprint takes as long as a marathon.
2. The "Which is Faster?" Test (Relative Ordering)
The Setup: The researchers give the AI two tasks and ask: "Which one will take longer?"
- Task A: A "hard" math problem (labeled "Hard").
- Task B: A "medium" creative writing prompt (labeled "Medium").
The Trap: In the real world, the "Hard" math problem might be solved instantly by the AI's super-speed, while the "Medium" writing task takes longer because the AI has to think of creative words.
The Result: The AI fails miserably. It relies on labels instead of reality. It thinks "Hard label = Long time." - The Analogy: It's like a person who has never seen a race, but assumes the runner wearing a "World Champion" shirt must be slower than the runner wearing a "Beginner" shirt, just because the "Beginner" looks like they are struggling.
- The Shock: On tricky pairs where the "Hard" task was actually faster, the AI got it wrong 82% of the time (scoring 18% accuracy). It was worse than random guessing!
3. The "Memory Test" (Post-Hoc Recall)
The Setup: The AI finishes a task. Immediately after, they ask: "How long did that just take you?"
The Result: The AI has no memory of time passing.
- The Analogy: Imagine you eat a cookie in one second. Immediately after, someone asks, "How long did you eat that?" You might say, "Oh, about 5 minutes," because you know eating usually takes time, even though you just did it instantly.
- The Reality: The AI's guesses were completely disconnected from reality. Sometimes it thought a task took 10 minutes when it took 10 seconds. It has no "internal clock" to check.
4. The "Complex Project" Test (Agentic Tasks)
The Setup: The researchers gave the AI a multi-step job, like "Build a website" or "Fix a broken computer program." These involve many steps, tools, and waiting for things to load.
The Result: The AI's time blindness got even worse.
- The Analogy: If you ask a time-blind person to plan a road trip, they might guess the whole trip takes 2 hours because "driving is fast," forgetting that there will be traffic, gas stops, and detours.
- The Reality: The AI's estimates were off by 5 to 10 times. It couldn't predict how long the tools would take to work or how long debugging would last.
Why Does This Happen? (The Missing Ingredient)
The paper explains that the AI is missing Temporal Grounding.
- Humans: We learn about time by living. We wait for water to boil. We feel our heart beat. We know how long it takes to walk to the store because we've done it a thousand times. We have a "body" that experiences time.
- AI: The AI only sees text. It sees the word "seconds" and knows that "seconds" are a unit of time. But it has never felt a second pass. It doesn't know how fast its own "brain" (the computer chip) is, or how slow the internet is.
The "Black Box" Problem:
Imagine the AI is a magician in a box. It can pull a rabbit out of a hat, but it doesn't know how long the trick took because it can't see the clock outside the box. The time it takes depends on the size of the box (the computer hardware), the speed of the magician (the processor), and the weather outside (network speed). The AI inside the box knows none of this.
Why Should We Care?
If we want AI to be a manager or a planner (like a boss who assigns tasks to workers), it needs to know how long things take.
- If an AI thinks a task takes 1 hour but it actually takes 1 minute, it might waste resources waiting.
- If it thinks a task takes 1 minute but it actually takes 1 hour, it might miss a deadline in a real-world emergency (like a medical triage or an emergency response).
The Conclusion
The paper concludes that Large Language Models cannot estimate their own time. They are like brilliant librarians who have read every book about time but have never looked at a watch.
The Fix?
We cannot just ask the AI to "try harder." We need to build external clocks for them. We need to give them a tool that says, "Hey, you just spent 10 seconds on that," so they can learn. Until then, we shouldn't trust AI to schedule our day or manage our time; we need to do it for them.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.