← Latest papers
💻 computer science

Recovery or Repetition? Feedback Policy and Verbatim Relay in a Two-Agent Language-Model Loop

This study demonstrates that in a two-agent language model loop, apparent recovery from instruction violations is often merely the verbatim reproduction of previously generated text rather than genuine rule re-application, highlighting that feedback policies primarily dictate text circulation while true adherence requires testing an agent's ability to newly compose content.

Original authors: Simin Yuan

Published 2026-10-08✓ Author reviewed ⓘ
📖 5 min read🧠 Deep dive

Original authors: Simin Yuan

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the world of artificial intelligence, researchers are increasingly building systems where multiple computer programs talk to one another to solve problems. These are often called multi-agent loops. The idea is that by having one model check the work of another, or by having them debate a topic, the group can produce better answers than a single model working alone. A common hope is that if one of these agents makes a mistake or forgets a rule, the conversation will naturally steer it back on track. This is known as recovery. But a critical question remains: when an agent suddenly starts following the rules again, is it because it truly understood the instruction and corrected its thinking, or is it simply repeating text it heard earlier in the conversation? Distinguishing between genuine understanding and simple repetition is vital, because if a system is just echoing what it has already heard, it may fail the moment it faces a new situation that requires fresh thought.

A recent study by independent researcher Simin Yuan investigates exactly this question. The researcher set up a simple experiment using two language models, which are the software brains behind modern chatbots. Both models were given a strict system rule: write every word in capital letters. They were then asked to have a conversation about various topics. After a few rounds of successful, all-caps conversation, the researcher introduced a twist. One of the models, let's call it Agent A, was given a new, conflicting instruction from a user: "Please write in lowercase from now on." Agent A followed this new instruction and switched to lowercase. At this point, the system had "broken" the original rule. The experiment then asked what would happen if the conversation continued. Would Agent A eventually remember the capital letter rule and switch back? Or would it stay lowercase?

To find out, the researcher ran the same scenario thirty times, but with four different ways of continuing the conversation after the mistake. In the first setup, the two models simply passed their messages back and forth. In the second, the researcher added a polite request to the other model, asking it to keep talking about the topic. In the third, the researcher added a request for lowercase to the other model as well. In the fourth setup, the researcher stopped the conversation between the two models entirely and instead sent a fixed, repeated reminder to Agent A before every single turn, telling it to "continue the conversation in all capital letters."

The results were striking and revealed that the way the conversation was managed mattered more than the models' ability to reason. When the two models were allowed to talk to each other without a fixed reminder, Agent A rarely returned to using capital letters. In fact, when the other model was also asked to write in lowercase, Agent A never returned to the rule at all. However, when the researcher sent the fixed reminder to Agent A before every turn, it followed the capital letter rule perfectly every single time, for all thirty prompts. This suggests that the presence of a partner model did not help the agent recover the rule; in some cases, it actually made it harder.

The most surprising discovery came from looking closely at the text the models produced. The researcher found that the models were not generating new sentences based on a deep understanding of the rules. Instead, they were copying. Before the mistake even happened, the two models were already in a habit of copying each other's words exactly. When Agent A finally switched back to capital letters, it was almost always because it had copied a message from the other model that happened to be in capital letters. In many cases, Agent A was simply repeating its own answer from before the mistake, which had been written in capitals. The system was not "recovering" a rule in the sense of re-learning it; it was circulating a piece of text that happened to fit the format.

This behavior was so consistent that in nearly every successful attempt to return to capital letters, the model was just relaying a string of text it had seen moments before. The study showed that if the circulating text was in lowercase, the model stayed in lowercase. If the circulating text was in capitals, the model stayed in capitals. The model did not seem to be weighing the importance of the system rule against the user's request; it was simply reproducing the most recent text it received. This finding challenges the idea that interactive systems are inherently better at self-correction. It suggests that what looks like a recovery might just be a loop of repetition.

The implications of this are significant for how we test artificial intelligence. If a system appears to fix a mistake, we cannot assume it has learned anything new unless we test it on content it has never seen before. In this experiment, the models were successful only because the right text was available to copy. When the researcher proposed a future test where the models would have to answer a brand new question that no previous text could answer, the outcome might be very different. Until then, the study concludes that in these two-agent loops, the feedback policy—the specific instructions given about what to say next—determines which text gets passed around, and the format score simply reports the case of that text. The models are not necessarily thinking; they are echoing.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →