← Latest papers
💬 NLP

Semantic Variability of Replies Across LLMs: Implications for Designing Conversation-Based Assessment

This study demonstrates that the choice of LLM and conversational context significantly impact the semantic consistency of generated replies, suggesting that current prompting strategies are insufficient for ensuring stable, comparable responses across different models in conversation-based assessments.

Original authors: Jiangang Hao

Published 2026-08-27
📖 4 min read☕ Coffee break read

Original authors: Jiangang Hao

Original paper dedicated to the public domain under CC0 1.0 (http://creativecommons.org/publicdomain/zero/1.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the modern world of education, measuring how well people communicate and work together has long been a stubborn problem. Unlike a math test where a single number proves a student knows the answer, skills like collaboration happen in the messy, flowing back-and-forth of conversation. To measure this fairly, educators need to ensure that every student faces a similar challenge, even if the conversation takes a different path. Recently, artificial intelligence has offered a new way to solve this. By using computer programs that can talk and listen, schools can create virtual partners for students to practice with. These programs can adapt to what a student says, creating a natural conversation that reveals how well the student is thinking and working. However, a quiet worry has grown among experts: if the computer program changes, does the conversation change too? If a student takes a test today with one version of the software and another student takes it next year with a newer version, will they be answering the same questions in the same way? This question strikes at the heart of fairness. If the computer's replies shift wildly depending on which model runs it, the test might stop measuring the student's ability and start measuring the software's quirks instead.

A researcher at the Educational Testing Service set out to investigate this very issue. They wanted to know if the meaning of a computer's reply stays the same when the underlying software changes. To find out, they did not invent fake conversations. Instead, they took real messages from past online science projects where two people worked together to solve problems. They selected 61 specific moments from these chats where one person had said something important, and the other person had given a clear, helpful reply. The researcher then asked four different versions of a large language model—a type of artificial intelligence designed to understand and generate human language—to reply to those same 61 messages. They ran this experiment twice for each message. In the first run, the computer saw only the single message it needed to answer. In the second run, it saw that message plus the entire history of the conversation that came before it. For every single message, each computer model generated 100 different replies to see how much the answers varied.

The results showed that both the choice of computer model and the conversational context affect response similarity. When the researcher compared the 100 replies generated by a single model, they found the answers were quite similar to each other, with similarity scores ranging from 0.715 to 0.795. But when they compared the replies from one model against the replies from a different model, the similarity dropped significantly. It was as if each model had its own distinct voice and way of thinking. The effect of adding conversation history depended on the specific model used; in some cases, including chat history slightly improved the alignment of replies with human responses, but it did not make the models agree more with one another in all cases. The study also found that newer models from the same family of software tended to produce more similar answers to each other than they did to older, different models. This suggests that the "family tree" of the software plays a significant role in how it behaves alongside the specific instructions given to it.

Perhaps most importantly, the researcher found that relying on prompting and conversational context alone may not be sufficient to ensure highly similar response behavior when the underlying LLM changes. Even when the models were given the exact same prompt and the same conversation history, the choice of model still led to significant differences in the semantic content of the replies. The study suggests that relying on the natural flexibility of these artificial intelligence systems is risky for high-stakes testing. If a test relies on the computer to keep the conversation on track, and the computer changes its style every time it is updated, the test results could become unreliable. The researcher concluded that to build a fair testing system that lasts, they cannot just rely on the software to do the right thing on its own. Instead, they need to build extra layers of control into the system—like rules or templates—that ensure the computer's replies stay within a safe, consistent range, no matter which version of the software is running underneath. Without these safeguards, the rapid evolution of artificial intelligence could undermine the very fairness that educational testing is meant to protect.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →