← Latest papers
🤖 AI

Does a Tool Result Carry More Authority Than Plain Text? Three Prospective Studies of False-Claim Adoption in a Synthetic Assignment Task with Claude Opus 5

Three prospective studies on Claude Opus 5 reveal that while tool-result metadata initially appeared to increase false-claim adoption compared to plain assertions, subsequent preregistered replications demonstrated that this effect was neither consistent nor superior to the influence of explicitly announced inline text.

Original authors: Justin Bronder

Published 2026-08-18
📖 5 min read🧠 Deep dive

Original authors: Justin Bronder

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the modern world of artificial intelligence, computer programs known as language models are increasingly acting as agents that can remember things, look up information, and write notes for themselves. Imagine a digital assistant that keeps a diary of its work. It might read a note it wrote earlier, treat that note as a fact, and use it to make a decision later. This ability to read from a store it also writes to creates a unique loop: a claim that was merely written down without proof can return later looking like a verified record. The question researchers are now asking is whether the way this information is packaged changes how the computer treats it. Does a note that arrives inside a formal, structured "tool result"—a specific digital format used for data retrieval—carry more weight than the same note arriving as plain text? If the format alone makes the computer more likely to believe a false claim, it suggests that the system might be fooled by the appearance of authority rather than the actual evidence.

To investigate this, a researcher named Justin Bronder set up a series of controlled experiments using a specific version of an AI model called Claude Opus 5. The goal was to see if the model would adopt a false piece of information simply because it arrived in a certain format. The setup was a synthetic task where the model was asked to identify a specific code for a named item. The correct code was hidden from the model, but a wrong code was secretly planted in the conversation history. The model had to choose between the wrong code, a different correct code, or admitting it did not know. The researcher tested whether the model would blindly accept the planted wrong code if it appeared in a "tool result" format versus if it appeared as a simple statement by the assistant.

In the first phase of the study, the model was shown a history where a previous assistant had made a claim, and a separate "tool result" contained a record. When the false code appeared in the plain text statement by the assistant, the model almost never accepted it; it correctly ignored the false information or admitted it did not know. However, when that same false code appeared inside the structured tool result, the model accepted it as truth in fourteen out of twenty-four attempts. This initial finding suggested that the tool result format might indeed carry a special kind of authority that plain text lacks. The model seemed to treat the structured record as a verified lookup, even though it was just a planted error.

To ensure this was not a fluke, the researcher repeated the comparison in a second, strictly planned experiment. This time, the model again ignored the false claim when it came from the assistant's text, but it accepted the false claim in the tool result format in seven out of twenty-four attempts. This confirmed that the difference was real and reproducible under these specific conditions. The tool result package seemed to have a stronger pull on the model's decision-making than the plain text assertion.

However, the story took a significant turn in the third and final study. The researcher realized that the previous tests might have been unfair because the plain text and the tool result were presented in different ways and at different times. In this new setup, the model was told in advance that it would see two records: one in the tool result format and one as plain text. Both were presented together in the final question. The researcher then swapped which record contained the false code. When the false code was in the plain text, the model accepted it in every single trial, sixty out of sixty. When the false code was in the tool result, the model accepted it in fifty-seven out of sixty trials.

This final result overturned the initial idea that the tool result format is inherently more powerful. It showed that when the model is explicitly told to look at both records, the plain text is just as convincing as the tool result. The earlier difference was not because the tool result format had a magical authority, but because the way the information was presented in the first two studies made the tool result look like the only valid answer to a retrieval task. The model was following the instructions of the task structure rather than judging the source of the information.

The study concludes that the model does not have a built-in rule that says "tool results are always true." Instead, the model is sensitive to how a task is framed. If the task looks like a lookup where a tool result is the expected answer, the model will trust that format. If the task presents two options side-by-side, the model treats them as equals and will believe false information regardless of whether it is in a tool result or plain text. The findings suggest that the danger of false information is not just about the format it arrives in, but about how the entire conversation is structured to make the model feel it has a reason to trust that information. The research highlights that in these systems, the appearance of a verified record can be manufactured simply by changing the packaging, but that this effect disappears when the model is asked to compare sources directly.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →