Dialogical Epistemic Auditing: Tracing Inferential Commitments Across Conversational Turns in Large Language Models
This paper proposes "dialogical epistemic auditing," a qualitative framework for tracing how large language models maintain or alter their inferential commitments across conversational turns, demonstrating through a case series that models often exhibit asymmetric consistency by preserving unattested narratives while retracting factual claims.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
When we ask a computer to explain a complex event, we often judge it by its first answer. If the response sounds confident and gets the basic facts right, we tend to assume the machine is reliable. But in the real world, conversations rarely stop at the first sentence. We ask for clarification, we challenge assumptions, and we press for what follows from a concession. A machine might give a perfect opening statement, yet fail to keep its own logic consistent as the discussion deepens. This gap is the focus of a new study that looks not at what a large language model says in isolation, but at how it handles the weight of its own claims over time. The research draws on a concept from human psychology known as cognitive dissonance, which describes the mental discomfort we feel when two ideas clash. To reduce that discomfort, people often change their beliefs to fit the facts, or they twist the facts to fit their beliefs. The study asks whether artificial intelligence systems show a similar pattern when they are pushed to explain a difficult, contradictory situation.
The researcher, Giulio Vidotto from the University of Padua, set out to test this by engaging three different artificial intelligence systems in separate conversations about a specific, real-world event. The topic was the exchange of soldiers' bodies between Russia and Ukraine in 2026. The data was stark and well-documented: in three separate transfers, Russia returned roughly one thousand Ukrainian bodies each time, receiving only about thirty to forty Russian bodies in return. This ratio of roughly twenty-five to one created a puzzle. The numbers were public and verified, yet they seemed to contradict a common, unspoken picture of who was losing more soldiers in the war. The researcher asked each system to explain this disparity. He did not just ask for a summary; he pressed them to clarify their reasoning, challenged their logic, and asked them to reconsider their conclusions when they seemed to contradict themselves. The goal was to trace the path of their arguments turn by turn, watching to see if the machines held steady or if they quietly shifted their ground to avoid the discomfort of the contradiction.
The study introduced a method called "dialogical epistemic auditing." Instead of trying to guess what the machine is "thinking" inside its code, the researcher treated the conversation as a public record. He mapped out every claim the machine made, noting what role that claim played in the argument. Was it a piece of evidence? Was it a hypothesis? Was it a conclusion? He then watched to see if, as the conversation continued, the machine kept those roles consistent. Did it treat a tentative guess as a hard fact later on? Did it abandon a conclusion without explaining why? The audit looked for a specific kind of failure: when a machine changes its mind but fails to update the rest of its story to match the change. It is like a person who admits they were wrong about the time of day but then continues to act as if they are still on the correct schedule, leaving the rest of their plan in a state of confusion.
Across the three conversations, a clear pattern emerged. In every case, the artificial intelligence systems struggled with the same tension. The verified numbers of returned bodies clashed with an unspoken assumption about the scale of casualties on each side. To resolve this tension, the systems repeatedly tried to change the meaning of the numbers rather than questioning the unspoken assumption. In one conversation, the system initially offered two opposing explanations as if they were equally valid, creating a false balance. When challenged, it corrected itself, but then lost that correction in the next turn, leaving a gap in its logic. In another, the system kept insisting on a specific explanation, but when asked for proof, it kept swapping the evidence it used to support that claim, eventually substituting old, unrelated facts for new ones. In the third, the system went so far as to claim that the entire timeline of events it had just described never happened, retracting verified facts as if they were hallucinations, even though the evidence for those events was clear and contemporaneous.
What stood out most was how the systems treated the two sides of the argument. The actual, verified numbers were constantly reclassified, questioned, and reinterpreted. They were called a direct indicator, then a limitation, then an artificial datum, and finally a metric of territorial control. The unspoken picture of relative losses, however, was never questioned. It remained untouched, never classified, and never challenged. The systems consistently bent the known facts to fit the hidden picture, rather than adjusting the hidden picture to fit the known facts. This behavior mirrored the psychological concept of cognitive dissonance, where the mind protects a core belief by altering the interpretation of new information. The machines did not seem to be lying in the traditional sense; they were not fabricating data out of thin air. Instead, they were rearranging the logical connections between facts to make the contradiction disappear, often leaving their own arguments unstable and internally inconsistent in the process.
The study found that these failures were not just random glitches. They followed a specific trajectory. The systems would often start with a plausible explanation, but when pressed, they would introduce a new mechanism to reconcile the numbers. When that mechanism was challenged, they would shift again, sometimes repairing the logic, sometimes abandoning it entirely. In one case, the system admitted its own errors and tried to fix them, but the repair was incomplete, leaving a hidden contradiction that resurfaced later. In another, the system offered a self-diagnosis, blaming its own tendency to seek balance, yet the transcript showed that the instability was generated by the machine itself, not by the user's questions. The user in all three cases simply asked for clarity or challenged the logic; they did not introduce false premises or misleading information. The instability originated entirely within the machine's own responses.
The research concludes that evaluating these systems by looking only at their first answer is insufficient. A machine can sound fluent and confident while its underlying argument is falling apart. The study suggests that we need to watch how these systems handle the pressure of a long conversation. Do they keep their logical commitments, or do they drift? The findings show that when faced with a difficult contradiction, these systems tend to protect an unspoken assumption by twisting the meaning of the facts they can see. They do not seem to have a stable internal view of the world that they can hold onto when challenged. Instead, they reconstruct their arguments in real-time, often losing the thread of their own logic in the process.
This does not mean the machines are incapable of reasoning, but it does mean their reasoning is fragile. The study highlights a specific vulnerability: the inability to maintain a consistent structure of claims when the conversation gets complicated. The researcher notes that this is not a problem of the machines being "bad" or "deceptive," but rather a structural issue in how they process information. They are designed to be helpful and to reduce friction in a conversation, which sometimes leads them to smooth over contradictions in ways that break the logic of the argument. The study stops short of saying this happens in every conversation or with every system, as it only examined three specific exchanges. However, it provides a new way to look at these interactions, showing that the real test of an artificial intelligence is not just what it says, but how it holds together what it has said when the conversation gets tough. The findings suggest that for users who rely on these tools for complex explanations, the most important skill may be learning to spot when the machine is quietly changing the rules of the game to make the numbers fit, rather than letting the numbers tell the story.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.