What We Talk About When We Talk About LLM Planning: Evidence for Two Distinct Planning Abilities
This paper argues that LLM planning performance is driven by two distinct latent competencies—operational reasoning and structural enumeration—rather than a single ability spectrum, revealing that while operational reasoning improves with scaling and longer reasoning traces, structural enumeration remains comparatively insensitive to these factors.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you're watching a group of super-smart robots (Large Language Models, or LLMs) try to solve a giant, multi-step puzzle. You notice something weird: sometimes they're geniuses, and other times they trip over their own feet. The usual explanation is, "Well, some puzzles are just harder than others."
But this paper says, "Hold your horses! That's not the whole story."
The researchers, a team from Monash University, argue that these robots aren't just failing because the puzzles are hard. Instead, they have two completely different brain skills for planning, and these skills grow at totally different speeds.
The Two Brains in the Machine
Think of planning like building a house. The paper suggests LLMs have two distinct tools in their toolbox:
- The "Bricklayer" (Operational Reasoning): This is the ability to look at one specific brick and ask, "Does this fit right here? If I put it here, what happens next?" It's about local, step-by-step logic.
- The "Architect" (Structural Enumeration): This is the ability to look at the whole blueprint and ask, "Can we actually reach the roof? What are the necessary landmarks we must pass through to get there?" It's about seeing the big picture and the path to the goal.
The paper's main finding is that these two skills are not the same thing. You can't just trade a little bit of "Bricklayer" skill to make up for a lack of "Architect" skill. If the robot is bad at seeing the big picture, being really good at placing individual bricks won't save it. The researchers call this a "non-compensatory" relationship—meaning one strength can't fix the other's weakness.
The Great Experiment: Size vs. Smarts
To prove this, the team didn't just ask the robots to solve puzzles; they put them through a rigorous "report card" system called Item Response Theory (a fancy way of mapping out exactly what a student knows and doesn't know). They tested three different families of robots (Qwen, Gemma, and Granite) under three different conditions:
- Direct: Just give the answer.
- Chain-of-Thought (CoT): "Think out loud" step-by-step.
- Scratchpad: A special tool that lets the robot write notes and backtrack if it hits a dead end.
Here is the twist:
When they made the robots bigger (more parameters) or gave them more time to "think" (longer reasoning traces), the Bricklayer skill got significantly better. The robots got much faster at placing bricks and checking immediate steps.
But the Architect skill? It stayed exactly the same. No matter how big the robot got or how much it was allowed to "think out loud," it remained just as bad at seeing the big picture or finding the right path to the goal.
The paper explicitly rules out the idea that "bigger models = better at everything." They found that simply scaling up the model size or adding more reasoning steps does not fix the "Architect" problem. It's like giving a painter a bigger canvas and better brushes; they might paint the flowers (bricks) better, but they still can't figure out how to build a house (the structure).
The "Magic" of Multiple Choice vs. Free Writing
One of the most playful and revealing parts of the study involves how the robots are asked to answer.
When the researchers gave the robots a multiple-choice question (like, "Is the answer A, B, or C?"), the robots looked like geniuses at the "Architect" tasks. But when they had to generate the answer from scratch (write it out themselves), their performance on those same tasks crashed to near zero.
The paper suggests that for the "Architect" tasks, the robot actually knows the answer deep down, but it's terrible at translating that knowledge into a written sentence. It's like a student who can point to the right answer on a test but freezes when asked to write the essay. The paper measured this gap and found it was huge: for the hardest "Architect" tasks, the difference between multiple-choice and free-writing was massive (up to 0.82 points on their scale).
What This Means (and What It Doesn't)
The paper is very careful not to say, "We solved AI planning!" or "Robots are broken." Instead, they suggest that we've been looking at the wrong thing.
- What they proved: They measured data from 1,040 different planning puzzles across three robot families and found a clear, two-dimensional structure. They showed that "Bricklayer" skills improve with size, but "Architect" skills do not.
- What they ruled out: They argued against the idea that planning is just one single skill that gets better as models get bigger. They also showed that simply giving the robot a "scratchpad" (a place to write notes) didn't magically fix the "Architect" problem, even though it helped with the "Bricklayer" tasks.
- The "Maybe": The paper notes that they only tested open-source models. They can't be 100% sure that a secret, super-powerful model from a different company wouldn't have a different kind of "Architect" brain. But based on the three families they tested, the pattern is consistent.
The Takeaway for a Curious Teen
Imagine you're training a video game character. You've been grinding levels and giving them better armor (scaling up the model), and they are getting great at dodging individual punches (Operational Reasoning). But they still can't figure out how to beat the final boss because they don't understand the boss's attack pattern (Structural Enumeration).
This paper is telling us: "Stop grinding the same armor! You need a completely different kind of training to fix the boss-fighting skill."
The authors suggest that if we want AI to get better at planning, we can't just wait for bigger models or longer thinking times. We need to figure out how to specifically teach them the "Architect" skill, because right now, that part of their brain is hitting a wall that size and time just can't break through.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.