Scaffolding the Strategist: Architecture-Dependent Reasoning Interventions in Hotelling Spatial Markets
This paper demonstrates that structured reasoning interventions in Hotelling spatial markets have architecture-dependent effects, where commitment scaffolding benefits standard models but hinders reasoning-optimized ones, while principled separation shows the opposite pattern, ultimately revealing that adversarial stress degrades reasoning models more severely and that a persistent gap between identifying and executing strategies can only be closed for reasoning models through specific architectural alignment.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine two different kinds of brainy robots trying to solve a tricky game of "Hotelling's Hot Dog Stand." In this game, two vendors have to pick the perfect spot on a beach to sell hot dogs, knowing that customers will walk to the closest one, but they hate walking too far. It's a puzzle that requires thinking ahead, guessing what the other guy will do, and doing some serious math.
The researchers at Caltech wanted to see if giving these robots a "cheat sheet" or a special set of instructions (called scaffolding) would help them play better. They tested two specific robots:
- The Standard Robot (GPT-4.1-mini): A smart, fast worker who follows instructions well but doesn't have a built-in "thinking" mode.
- The Reasoning Robot (GPT-5-mini): A super-smart worker who has a built-in "thinking" mode, constantly chatting with itself to solve problems before answering.
Here is the big surprise they found: What helps one robot actually hurts the other. It's like trying to teach a chess grandmaster to play by forcing them to write down every single move on a piece of paper before they make it, while telling a beginner to just "think out loud" to stay focused.
The Great Crossover
The researchers tried four different ways to help the robots, but two of them created a perfect "crossover" effect:
The "Commitment" Trick: This asked the robots to write down their rules and promise to stick to them before solving the problem.
- Result: The Standard Robot got better! It scored +0.21 points higher. It needed that extra structure to stay on track.
- Result: The Reasoning Robot got worse! It dropped -0.63 points. Why? Because this robot was already thinking deeply on its own. Forcing it to write down a "commitment" first was like putting a speed bump in front of a race car—it just got in the way and confused the robot.
The "Separation" Trick: This forced the robots to break the problem into three strict steps: 1) State the rules, 2) Make predictions, 3) Give the answer.
- Result: The Reasoning Robot loved it! It scored +0.31 points higher. The structure helped organize its powerful internal thinking.
- Result: The Standard Robot hated it! It dropped -0.40 points. It didn't have the brainpower to use this fancy three-step process, so the extra steps just slowed it down and caused mistakes.
This wasn't a fluke. The researchers ran 720 different tests (8 questions, 3 different ways of asking, 3 times each, for 2 robots). The difference was so clear that the math says there's only a 0.2% chance this happened by accident.
The "Talk Too Much" Trap
The researchers also tried a "Stress Test." They asked the robots to argue against their own answers and try to prove themselves wrong.
- The Result: Both robots got worse, but the Reasoning Robot crashed harder, dropping -1.47 points compared to the Standard Robot's -0.57.
- The Lesson: The more the Reasoning Robot thought about its own correct answers, the more it started to doubt them. It's like a person who knows the answer to a riddle but starts overthinking it until they get it wrong. The researchers found that the easier the question was, the more the robot messed up when forced to doubt itself.
Knowing vs. Doing
There was another weird thing the researchers noticed called the "Declarative–Procedural Gap."
- The Standard Robot could often say the right strategy (like "I should move to the center of the beach"), but when it tried to actually do the math to prove it, it failed. It knew the "what" but couldn't do the "how."
- The Reasoning Robot was better at doing the math, but it still struggled to connect its knowledge to action—unless they used the "Separation" trick. When they forced the Reasoning Robot to separate "stating the rule" from "doing the math," it finally closed the gap, achieving a 25.9% success rate where its ability to state the right strategy perfectly matched its ability to execute it.
What This Means
The main takeaway isn't that one robot is "better" than the other. It's that one size does not fit all.
- If you have a robot that doesn't think much on its own, you need to give it structure (like "Commitment").
- If you have a robot that thinks too much on its own, you need to give it a clear process to organize those thoughts (like "Separation").
The researchers are careful to say this is based on simulations with two specific robots from one company. They haven't proven this works for every robot in the world, but it strongly suggests that the way we talk to AI needs to change depending on how the AI is built. If you treat a super-thinker like a simple worker, or a simple worker like a super-thinker, you might just make things worse.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.