CurveShift: Is Agent Progress Scalar? Separating Level from Shape
This paper argues that while most apparent shifts in AI progress toward harder tasks are artifacts of ceiling effects and rising overall ability, a distinct, genuine improvement in solving the hardest competitive programming problems emerged in models released after September 2024, a finding isolated by controlling for scaffold confounds using the LiveCodeBench benchmark.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are watching a video game leaderboard where players are trying to solve puzzles of increasing difficulty. For years, the story has been simple: "The players are getting better." We measure this with a single number, like an average score or a "time to beat" metric. It's like saying a runner is faster because their average race time dropped. But what if the runner didn't get faster at everything? What if they only got faster at the really hard, mountainous trails, while their speed on flat roads stayed exactly the same? Or worse, what if the "faster" time was just an illusion caused by the fact that they were already so good at the easy trails that they couldn't get much faster there anyway, making the hard trails look like they were improving by comparison?
This is the puzzle scientists are trying to solve with Artificial Intelligence (AI). We have these massive AI models that are supposed to be getting smarter every day. Researchers often summarize this progress with a single "scalar" number—a simple score that goes up over time. But a single number can hide a lot of secrets. It can't tell us if the AI is getting uniformly smarter, or if it's just getting a little bit better at the easy stuff while suddenly mastering the hardest, most complex problems. This paper asks a crucial question: Is the AI's progress just a general "level up," or is there a special, weird change happening specifically when the tasks get really, really hard?
The authors of this paper, a team of researchers from top universities, decided to investigate this using a clever trick called "CurveShift." They wanted to separate the "Level" (how smart the AI is in general) from the "Shape" (how its performance changes as tasks get harder). Think of it like a video game character. If you give the character more health and strength (a "Level Shift"), they will beat both easy and hard enemies better. But if you give them a special "Hard-Mode Skill" that only works on bosses (a "Shape Change"), they will crush the bosses while barely improving against the weak goblins. The big question is: Are our AI models just getting stronger overall, or have they unlocked a secret boss-killing skill?
The researchers started by looking at a lot of data from previous AI tests. They found that the popular story—that AI is suddenly getting much better at the hardest tasks—might be mostly an illusion. When they used a standard mathematical model (called a Rasch model) that assumes AI just gets generally smarter over time, the model actually predicted the "improvement on hard tasks" perfectly well. It turns out, when you are already 99% perfect at easy tasks, you can't get much better there. So, any small improvement you see on hard tasks looks huge in comparison, even if it's just the result of general improvement. This is like a ceiling effect: you can't hit the ceiling harder, so the only place you can see progress is on the floor. The paper argues that most of the "migration" of gains toward harder tasks is just this statistical artifact, not a magical new ability.
However, the story doesn't end there. After stripping away that "general improvement" illusion, the researchers found a tiny, but real, leftover effect. They focused on a specific type of test called LiveCodeBench, which is like a competitive programming contest where humans write problems and AI tries to solve them. Crucially, this test doesn't use any "scaffolding" or extra tools; it's just the raw AI trying to write code. This is important because in other tests, newer AI models are often paired with newer, better tools, making it impossible to tell if the AI got smarter or if the tools just got better. LiveCodeBench isolates the AI's raw brain power.
What they found was surprising. Even after accounting for the fact that newer models are generally smarter, models released after September 2024 were solving the absolute hardest problems better than their performance on easy and medium problems would predict. It wasn't a massive jump, but it was a real one. The authors estimate this "hard-item gain" to be about +0.40 logits (a specific unit of measurement for probability) under their most conservative assumptions. This means the solve rate for the hardest problems jumped from roughly 18% to 25%.
But here is the catch: this effect is very specific. It only shows up clearly when the AI is working alone without a "harness" or a team of tools helping it. In tests where AI agents use tools (like browsing the web or running code step-by-step), the researchers couldn't tell if the improvement was from the AI or the tools, because the new AI models always came with new tools. The "hard-item" boost was led by the strongest "reasoning" models (those designed to think step-by-step) and seemed to help most with tasks that require short bursts of deep thinking, not long, autonomous missions.
So, what's the takeaway? The paper suggests that the narrative of AI suddenly becoming a "super-intelligence" that can handle any hard task is mostly a statistical mirage caused by how we measure progress. The AI is indeed getting generally smarter, which explains most of the hype. But, there is a small, real spark of something new: a specific ability to tackle the hardest, most complex logic puzzles that goes beyond just being "generally smarter." It's not a magic bullet that solves everything, but it is a genuine, measurable shift in how these models handle the toughest challenges, provided they are allowed to work without extra help. The authors are careful to say this is a specific finding for coding puzzles and doesn't necessarily mean AI is ready to run a whole company or manage a complex, long-term project on its own yet. It's a small, real step forward in the "hard mode" of thinking, hidden underneath a mountain of general improvement.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.