When the Canonical Completion Is Wrong: Formalizing and Measuring the Jump in Large Language Models
This paper formally defines and measures the "jump" capability of large language models—specifically their ability to abandon a default canonical completion in favor of a unique, correct axiom system—demonstrating that frontier models successfully perform this step in all trials, thereby suggesting that any alleged incapacity lies in generating constraints or frameworks rather than in the jump itself.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the study of artificial intelligence, a central question has long divided researchers: can a machine truly think beyond the patterns it has already seen? Large language models are experts at induction, learning to predict the next word in a sentence by spotting statistical trends in vast amounts of text. They are also becoming increasingly skilled at deduction, following strict logical rules to reach a conclusion. However, a third form of reasoning, known as abduction, remains a mystery. This is the ability to leap from a set of observations to a completely new system of rules or axioms that explains them, a mental jump that often requires discarding the most obvious answer in favor of a surprising, unseen possibility. For years, skeptics have argued that these models are structurally incapable of such leaps, suggesting they are trapped in a loop of repeating what they already know. Others have challenged this, pointing to machine-generated mathematical discoveries, but without a precise way to define or measure this "jump," the debate has remained stuck in theory.
A team of researchers has now cut through this uncertainty by building a rigorous test to see if a model can abandon its default instinct when the facts demand it. They did not ask the models to invent new theories from scratch, which would be impossible to verify. Instead, they created a controlled laboratory environment where the rules of the game were set in stone. Imagine a puzzle where you are given a partial picture and a list of strict constraints that rule out the most obvious way to finish it. The researchers defined a "jump" as the moment a model realizes the standard, most logical completion is forbidden and must instead construct a new, hidden structure that fits the constraints perfectly. To do this, they used a mathematical framework that identifies the "canonical completion"—the single, most natural way to extend data based on what is already known. They then designed problems where this natural answer was explicitly forbidden, forcing the model to find a unique, non-obvious solution that the data itself did not show.
The researchers tested this on nine different certified problems, ranging from simple puzzles to complex scenarios with millions of possible answers, using four of the most advanced language models available. They measured how often the models stuck to the forbidden, default answer versus how often they successfully made the jump to the correct, constrained solution. The results were striking. In every single trial where the models were forced to choose, not one of them fell back to the default answer they were trained to prefer. Across 248 attempts, the models successfully abandoned the excluded default every time, landing on the correct, certified solution. When they failed, it was not because they refused to jump; it was because they ran out of mental energy or made a calculation error, but they never reverted to the old, wrong answer.
This finding suggests that the ability to select a new path when the old one is blocked is not the bottleneck for these models. The models proved they can override their own instincts when the constraints are clear. The researchers conclude that if the models truly lack the capacity for the kind of creative leap seen in human history, it is not in the act of choosing between options. Instead, the difficulty likely lies in the earlier, more chaotic steps: generating the constraints that force the jump in the first place, or inventing the entirely new framework of rules that makes the jump possible. The study does not claim to have solved the mystery of machine creativity, but it has firmly established that when the path is clearly blocked, these models are capable of finding a new way forward.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.