← Latest papers
🤖 AI

What Drives Interactive Improvement from Feedback?

This paper introduces a controlled student-teacher evaluation framework across diverse reasoning tasks to demonstrate that multi-turn improvements often stem from additional computation rather than feedback utility, revealing that the student's ability to effectively act on high-quality external guidance is the primary bottleneck for genuine interactive improvement.

Original authors: Bartłomiej Cupiał, Jan Łojek, Mikołaj Garstecki, Szymon Pobłocki, Alicja Ziarko, Piotr Miłoś

Published 2026-07-01
📖 5 min read🧠 Deep dive

Original authors: Bartłomiej Cupiał, Jan Łojek, Mikołaj Garstecki, Szymon Pobłocki, Alicja Ziarko, Piotr Miłoś

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to solve a very difficult puzzle, like a complex math problem or a tricky coding challenge. You try once, fail, and then try again. Sometimes, you get better just because you kept trying (resampling). Other times, you get better because someone gave you a hint about what you did wrong (feedback).

This paper asks a simple but tricky question: When we see an AI getting better after multiple tries, is it actually learning from the hints, or is it just getting better because it has more time to think and try again?

To find out, the researchers set up a "classroom" experiment with AI models playing two roles: the Student (who tries to solve the problem) and the Teacher (who gives feedback). They tested 13 different AI models in both roles across four different types of hard tasks: math, coding, rule-learning, and pattern puzzles.

Here is what they discovered, explained through everyday analogies:

1. The "Just Keep Trying" Trap

The Finding: Often, when an AI improves after a second or third try, it's not because the feedback was helpful. It's just because the AI got to "roll the dice" again with more computer power.
The Analogy: Imagine you are trying to guess a combination lock. If you guess wrong, and then you just guess again without any new information, you might eventually get lucky. The paper found that for many AI models, giving them a "hint" (feedback) didn't help much more than just letting them guess again. The improvement came from the extra attempts, not the advice.

2. The "Self-Talk" vs. "Real Teacher" Difference

The Finding: When an AI gives itself feedback (talking to itself), it barely improves compared to just trying again. However, when a stronger AI acts as a teacher, the student improves significantly.
The Analogy:

  • Self-Feedback: Imagine you are stuck on a riddle and you say to yourself, "I think I messed up the first part." You try again, but you're still stuck in the same loop. You aren't really learning anything new.
  • External Teacher: Now imagine a master puzzle-solver looks at your attempt and says, "You're looking at the wrong shape; try rotating it." This specific, high-quality advice actually helps you solve the puzzle. The paper shows that a "good teacher" is essential; a weak teacher or talking to yourself doesn't cut it.

3. The Student is the Star, Not the Teacher

The Finding: The most important factor in whether an AI improves is who the student is, not who the teacher is. A smart student can learn from almost anyone, but a confused student won't learn even from a genius teacher.
The Analogy: Think of a music lesson. If you have a world-class violin teacher (the Teacher), but the student has never held a violin before and doesn't know how to read music (the Student), they won't suddenly become a virtuoso just because the teacher is famous. The student's own ability to understand and use the advice is the bottleneck. The paper found that the "student's identity" explained way more of the success than the "teacher's identity."

4. More History Doesn't Always Mean More Help

The Finding: Giving the teacher and student a longer history of their past conversation (more context) helps, but only if the models are smart enough to use it. It's not a magic fix.
The Analogy: Imagine a detective trying to solve a crime. If the detective is very sharp, giving them a 50-page file of past clues helps them find the culprit. But if the detective is easily overwhelmed, a 50-page file might just confuse them, and a short summary works better. The paper found that longer conversation histories only helped the "smarter" AI models.

5. Knowing the Answer Doesn't Always Help the Teacher

The Finding: Sometimes, letting the teacher see the correct answer (privileged information) helps them give better feedback. But not always. In some tasks, knowing the answer didn't make the teacher much better at explaining how to fix the mistake.
The Analogy: Imagine a math teacher who knows the answer is "42." If the student wrote "40," the teacher might just say "Wrong." But if the teacher knows why the answer is 42, they can say, "You forgot to carry the one." The paper found that for some tasks, knowing the answer helped the teacher give that "carry the one" advice. For other tasks, knowing the answer didn't help the teacher explain the error any better.

The Bottom Line

The paper concludes that feedback is only as good as the student's ability to use it.

If you want to build an AI that learns from feedback, you shouldn't just focus on finding the smartest "Teacher" AI. You need to focus on building a "Student" AI that is actually capable of listening, understanding the mistake, and changing its strategy. Without a student who can act on the advice, even the best feedback is just noise.

In short: Don't just give an AI more hints; make sure the AI is smart enough to actually listen to them.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →