Predictable Emergence: An Empirical Analysis of Whether Sharp Capability Jumps Follow from Smooth Per-Token Scaling Laws
This paper demonstrates that the sharp "emergent" jumps in exact-match accuracy observed in multi-digit integer addition are not genuine discontinuities in model capability but are instead fully predictable artifacts resulting from smooth, power-law improvements in per-token accuracy compounded by the nonlinearity of requiring all digits to be correct.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are watching a magic show where the magician gets better at pulling rabbits out of hats as the hat gets bigger. For years, scientists have noticed that when they make AI brains (called "transformers") bigger and feed them more data, they get smarter in a very predictable, smooth way. It's like a car engine that gets slightly more powerful every time you add a cylinder; the improvement is steady and follows a clear math rule. But then, people started noticing something weird: sometimes, these AI models seem to suddenly "wake up" and learn a brand-new skill overnight. One day they can't do simple math, and the next day, with just a tiny bit more size, they can solve complex problems perfectly. This sudden jump is called "emergence," and it's scary because if skills appear out of nowhere, we can't predict when an AI might suddenly become dangerous or incredibly smart.
The big question is: Are these skills really appearing out of thin air, or is it just an illusion caused by how we measure them? Think of it like a student taking a test. If you grade them on "Did they get every single question right?" (a pass/fail grade), they might look like they failed completely for a long time, and then suddenly pass. But if you look at "How many questions did they get right on average?" (a smooth score), you might see they were actually improving slowly the whole time. This paper asks: Is the "sudden jump" real, or is it just a trick of the test we're using?
The Great Math Magic Trick
A researcher named Alexander Memming decided to investigate this mystery using a very specific, controlled experiment. He didn't train any new AI models; instead, he used a bunch of existing, public AI models (called Pythia and OPT) that ranged from very small to quite large. He asked them a simple task: adding two big numbers together. The catch was that he could control exactly how long the answer needed to be (from 1 digit up to 6 digits).
When he looked at the AI's performance using the "smooth" way of measuring (checking how often it got each individual digit right), the results were boringly predictable. As the models got bigger, they got slightly better at guessing each digit, following a smooth, gentle curve. There was no sudden jump. It was like watching a runner slowly get faster mile after mile.
However, when he switched to the "sharp" way of measuring (checking if the entire answer was perfect), the story changed completely. For short answers, the AI was okay. But for long answers, the AI looked like it was failing miserably for a long time, and then—poof—it suddenly started getting perfect scores. This looked exactly like the famous "emergence" jump everyone was talking about.
The Secret Ingredient: A Simple Formula
Memming realized that this "jump" wasn't magic at all; it was just math. He built a simple model to explain it. Imagine you are trying to guess a 6-digit password. If you have a 90% chance of guessing any single digit correctly, that sounds great. But if you need to get all six digits right to win, your chances drop to about 53%. If you only have a 50% chance per digit, your chance of getting the whole password right is almost zero.
The paper shows that as AI models get bigger, their "per-digit" accuracy slowly creeps up from 50% to 90%. Because the test requires getting every digit right, the final score stays near zero for a long time, then shoots up to 100% very quickly once the per-digit accuracy gets high enough. It's not that the AI suddenly learned a new superpower; it's just that the "all-or-nothing" test makes a slow, steady improvement look like a sudden explosion.
Predicting the Future
The coolest part of the study is that Memming proved you can predict these "jumps" before they happen. He took the smallest AI models (which were still failing the long math problems) and measured their slow, smooth improvement on single digits. Using a simple formula that accounts for how many digits are in the answer, he was able to predict exactly when the larger models would suddenly start passing the test. His predictions were incredibly accurate, guessing the "jump" point within a tiny margin of error.
He also checked if there was any hidden "super-power" behavior that the math couldn't explain. He looked for signs that the AI was doing something extra-special at the moment of the jump, but found nothing. The "emergence" was entirely explained by the smooth improvement of small parts adding up to a big result.
The Verdict
So, what does this mean? For the specific task of adding big numbers, the paper suggests that the scary "sudden jumps" in AI ability are likely an illusion created by the way we test them. The AI isn't switching on a new lightbulb; it's just slowly turning up a dimmer switch, and our "pass/fail" test makes it look like the light just snapped on.
This doesn't mean AI will never have sudden surprises, but it gives us a powerful tool: if we can measure the smooth, steady progress of the small parts, we can predict when the big, scary jumps will happen. It turns a mystery into a math problem, showing that for now, at least, the future of AI is more predictable than we thought.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.