Think Fast: Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models
This paper evaluates the ability of frontier AI models to perform complex reasoning without explicit chain-of-thought tokens by measuring their task-completion time horizons against human baselines, revealing a trend of doubling performance annually and projecting that these models could soon solve tasks requiring over 25 minutes of human effort in a single pass by 2030.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are watching a magician perform a trick. Usually, the magician talks you through their process: "I'm going to shuffle the cards, then I'll peek at the top one, and now I'm going to make it disappear." This talking part is like an AI's "Chain of Thought" (CoT). It's the step-by-step reasoning the model shows us before giving an answer. Safety experts love this because they can listen to the "thinking" to make sure the AI isn't planning anything dangerous, like lying or trying to trick its human supervisors.
But what if the magician stopped talking entirely? What if they could do all that complex shuffling, peeking, and planning inside their head, in total silence, and then just snap their fingers to reveal the answer? If an AI can do this, the safety monitors are left in the dark, unable to see the "thinking" that might be hiding a bad plan. This paper asks a big, slightly scary question: How much can these super-smart AI models actually do in their heads without showing us a single word of their work?
The researchers decided to find out by putting frontier AI models through a massive gauntlet of over 30,000 puzzles, math problems, coding challenges, and logic games. But there was a catch: they forced the models to answer immediately, without any "thinking tokens" or internal monologue. They wanted to measure the "Time Horizon"—a fancy way of asking, "How long would it take a human to solve the same problem that the AI just solved in a split second?"
Here is what they discovered. Over the last six years, these AI models have gotten incredibly good at silent thinking. In 2019, an AI could only solve problems that a human could do in a fraction of a second without help. But by the time of this study's projections for 2026, the latest models (like GPT-5.5) are estimated to be able to solve problems that would take a human over 3 minutes to figure out, all without saying a word of their reasoning.
The growth is happening fast. The researchers found that the complexity of these silent tasks is doubling roughly every year (specifically, every 373 days). It's like the AI is learning to do a year's worth of mental gymnastics in just one calendar year. If this trend continues, by 2028, these models might be able to silently solve tasks that take humans 7 minutes, and by 2030, that could jump to 25 minutes of pure, hidden mental work.
To double-check their math, they also looked at how many "reasoning tokens" (the digital equivalent of thought steps) a reference model would need to solve these same problems. They found that the AI's silent capacity is also doubling every year in terms of these tokens, reaching over 1,500 tokens of hidden processing power for the latest models.
The paper suggests that while we can't be 100% certain about the future (the numbers have a wide range of possibilities), the trend is clear: AI models are rapidly gaining the ability to do complex, multi-step reasoning entirely inside their "minds" without showing us. This means that relying on listening to an AI's thoughts to keep it safe might become much harder in the near future, because the AI might just be doing all that thinking in silence. The authors urge developers to keep tracking this "silent" ability closely, because it's growing predictably and could change how we monitor these powerful systems.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.