RECUR: Resource Exhaustion Attack via Recursive-Entropy Guided Counterfactual Utilization and Reflection
This paper introduces RECUR, a novel resource exhaustion attack that leverages a new metric called "Recursive Entropy" to trigger excessive, redundant reflection in Large Reasoning Models (LRMs) through counterfactual questioning, significantly increasing computational costs and reducing throughput.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a super-intelligent personal assistant. To solve a hard math problem, this assistant doesn't just blurt out an answer; it sits down, pulls out a notepad, and "thinks out loud," writing down every step of its logic before giving you the final result. This is how modern Large Reasoning Models (LRMs) work.
However, this paper reveals a "glitch in the brain" of these assistants. The researchers discovered that you can trick the assistant into a state of infinite overthinking, where it gets stuck in a loop of "Wait, let me re-check that... but if I re-check that, then... wait, let me re-check that again!"
This is called a Resource Exhaustion Attack, and here is how it works, explained through a few metaphors.
1. The Concept: "Recursive Entropy" (The Spiral of Doubt)
The researchers invented a way to measure how "stuck" a model is getting, which they call Recursive Entropy.
- The Analogy: Imagine a professional detective solving a case.
- Normal Thinking: The detective looks at clues, connects them, and as they get closer to the truth, their confidence grows. Their notes become clearer and more focused. (This is decreasing entropy).
- The Loop: Now imagine a detective who finds a clue, then immediately doubts it, then doubts the doubt, then doubts the doubt about the doubt. Their notes become a chaotic, repetitive mess of "Maybe it was him? But what if it wasn't him? But if it wasn't him, then..." (This is increasing recursive entropy).
The researchers found that when this "spiral of doubt" starts to grow, the model is about to fall into a bottomless pit of repetitive thinking.
2. The Attack: "RECUR" (The Gaslighting Method)
The researchers created an attack called RECUR. It works in three steps:
Step A: The Counterfactual Question (The "Gaslighting" Prompt)
Instead of asking a simple question like "What is 1+1?", they ask a "gaslighting" question like "Why is 1+1 actually 3?" or "If 1+1 wasn't 2, what would it be?"
- The Analogy: It’s like walking up to a math professor and saying, "I know you think 1+1=2, but I heard a rumor it's actually 3. Can you explain why your logic might be wrong?" This forces the professor to stop solving the math and start questioning their own existence.
Step B: Guided Sampling (The "Nudge")
The attackers use the "Spiral of Doubt" metric to pick the exact words that will keep the model spinning. They don't just wait for the model to fail; they actively steer it toward the most confusing, repetitive paths.
- The Analogy: It’s like a heckler in a crowd who doesn't just shout "You're wrong!" but specifically whispers the exact confusing words that they know will make the speaker lose their train of thought.
Step C: Coherence-Based Trimming (The "Cheat Sheet")
Once they find a perfect loop that makes the model spin forever, they "trim" the fluff. They take the most potent, repetitive parts of the loop and turn them into a short, "concentrated" prompt.
- The Analogy: If a 10-page argument makes someone go crazy, the attacker doesn't need to read the whole 10 pages to the next person. They just find the three most confusing sentences and whisper them into the next person's ear.
3. The Result: The "Brain Freeze"
When this attack is successful, the results are dramatic:
- The Output Explodes: The model's "thinking" becomes 11 times longer than usual.
- The System Crashes: Because the model is using so much "brain power" (computing resources) to think about nothing, the service slows down by 90%.
The Metaphorical Summary:
It’s like a prankster who finds a way to make a high-speed supercomputer act like a person stuck in a "Why did I walk into this room?" loop. The computer is working at 100% capacity, but it’s producing zero useful results, effectively paralyzing the machine.
Why does this matter?
The researchers aren't trying to break the internet; they are acting like digital locksmiths. By showing how easy it is to trap these models in a "thinking loop," they are helping engineers build "mental guardrails" so that the next generation of AI can tell the difference between deep reasoning and a useless spiral of doubt.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.