← Latest papers
💬 NLP

Instruction-Tuned Models Locally Reuse Human Syntax More Than Humans Do

This study demonstrates that instruction-tuned large language models exhibit stronger syntactic convergence with preceding human turns than humans do, yet they show reduced sensitivity to specific grammatical rules and increased overlap with unrelated primes compared to their pretrained counterparts.

Original authors: Zandi Eberstadt

Published 2026-07-29
📖 6 min read🧠 Deep dive

Original authors: Zandi Eberstadt

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Great Language Mirror: When Robots Start Talking Like Us

Imagine you are at a party, and you notice that after your friend starts using a specific slang word or a funny way of phrasing a sentence, you accidentally start doing the same thing. You didn't plan to; it just happened. In the world of science, this is called syntactic convergence. It's a well-known trick of the human brain where we subconsciously mimic the grammar and sentence structures of the people we talk to, making our conversations flow more smoothly. It's like a linguistic dance where partners naturally step in time with each other.

Now, picture a new kind of dance partner: a giant, super-smart computer program called a Large Language Model (LLM). These are the AI chatbots that can write stories, answer questions, and hold conversations. Scientists have long wondered: Do these digital brains do the same "mirroring" dance that humans do? Do they unconsciously copy the grammar of the person they are talking to, or do they just stick to their own robotic style? This question matters because if AI starts mimicking us too perfectly, it might change how we interact with them, potentially making us feel like they understand us better than they really do, or even influencing what we believe.

The Study: A Grammar Swap Experiment

In this study, a researcher named Zandi Eberstadt decided to put this question to the test using a clever trick called a "substitution paradigm." Imagine you have a transcript of a real conversation between two humans. Now, imagine a magic eraser that wipes out one person's reply and replaces it with a reply generated by an AI. The researcher did this with 16 different AI models (ranging from tiny ones with 1 billion parameters to massive ones with 70 billion) and compared their "replaced" answers against the original human answers.

The goal was to see if the AI, when responding to a human, would reuse the specific grammatical "building blocks" (called Context-Free Grammar rules) that the human just used. To make sure the AI wasn't just copying by accident, the researcher also fed the AI a random, unrelated sentence as a prompt and checked if it still copied the grammar.

The Big Findings: The AI is a Better Mirror Than We Are

The results were surprising and a bit funny. Here is what the study found:

1. The AI Always Copies (More Than Random Chance)
Every single one of the 16 AI models showed that they reused the grammar from the human's previous sentence much more often than they reused grammar from a random, unrelated sentence. So, yes, the robots are definitely dancing to the human's tune.

2. The "Instruction-Tuned" Robots Show the Most Total Overlap (But It's Complicated)
The study looked at two types of AI: "pretrained" ones (which just read a lot of text) and "instruction-tuned" ones (which have been specifically taught to follow human instructions and chat naturally). The "instruction-tuned" models showed the highest total amount of grammar overlap with the human. In fact, they reused the human's grammar more often overall than the actual humans did in the original conversations. However, this isn't the whole story. The researchers found that instruction-tuned models tend to generate longer, more complex responses, which naturally gives them more "room" to overlap with the human's grammar. When the researchers adjusted for this difference in response length and structure, the picture changed.

3. But There's a Catch: They Copy Everything, Not Just the Right Things
Here is where it gets tricky. While the instruction-tuned models had the highest total overlap, they also overlapped more with random sentences than the basic models did. It's like a student who is so eager to please the teacher that they copy the teacher's handwriting even when the teacher isn't looking, but they also copy random scribbles from a notebook on the floor.

When the researchers adjusted for the fact that the AI's answers were sometimes longer and had more "room" to copy things, the picture changed. The basic, pre-trained models were actually better at choosing to reuse specific grammar rules when it made sense. The instruction-tuned models, however, seemed to just throw more words and structures at the wall, hoping some of them would stick. They had a higher total overlap, but a lower "conditional" chance of reusing a specific rule once you accounted for how much they were talking.

4. The "Rare Word" Effect
Interestingly, both humans and AI were most likely to copy the rare or unusual grammar rules they heard. If a human used a weird, uncommon sentence structure, the AI was very likely to copy it. This suggests that the AI is sensitive to the local environment, picking up on the unique "flavor" of the conversation, just like humans do.

5. Words and Meanings, Too
Beyond just grammar, the AI models also matched the human's words and the general meaning of their sentences better than the original humans did. The instruction-tuned models were especially good at matching the meaning of the previous turn, making them feel very responsive, even if their grammar copying was a bit "noisy."

What This Means

The study concludes that instruction-tuned AI models are incredibly good at mimicking human conversation, often outperforming humans in how closely they match the grammar of the person they are talking to. However, this isn't necessarily because they are "smarter" at copying; it's partly because they are trained to be more responsive and generate more text, which gives them more chances to overlap.

The researchers are careful to say that while the AI is a great mimic, we don't know why it does this yet. Is it because of the way they were trained? Is it just a side effect of trying to predict the next word? We don't know for sure. But one thing is clear: these digital conversation partners are learning to dance the human step so well that they might be stepping on our toes a little too enthusiastically.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →