← Latest papers
💬 NLP

More Aligned, Less Diverse? Analyzing the Grammar and Lexicon of Two Generations of LLMs

This study compares two generations of LLMs against human-authored news text using HPSG formalism and diversity metrics, revealing that while human writing remains stable, newer instruction-tuned models exhibit significantly reduced syntactic and lexical diversity, suggesting a potential narrowing of expressive range despite improved coherence.

Original authors: Adrián Gude, Roi Santos-Ríos, Francis Bond, Dan Flickinger, Carlos Gómez-Rodríguez, Olga Zamaraeva

Published 2026-05-08
📖 4 min read☕ Coffee break read

Original authors: Adrián Gude, Roi Santos-Ríos, Francis Bond, Dan Flickinger, Carlos Gómez-Rodríguez, Olga Zamaraeva

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a music critic comparing two bands: The Humans (professional journalists) and The Robots (Artificial Intelligence). You want to see how their music has changed over time.

This paper is like a deep dive into the sheet music of these two groups. Instead of just listening to the melody (the words), the researchers looked at the grammar (the chords and structure) and the vocabulary (the specific instruments used) to see if the robots are sounding more like humans or if they are becoming something else entirely.

Here is the breakdown of their findings in simple terms:

1. The Setup: Two Eras, Two Bands

The researchers compared two different time periods:

  • The Humans: They looked at news articles from The New York Times written in 2023 and 2025.
  • The Robots: They compared two generations of AI models:
    • Older Robots (2023): These were the "base" models. Think of them as raw, untrained musicians who just play whatever comes to mind.
    • Newer Robots (2025): These are the "instruction-tuned" models. Think of them as musicians who have been given a strict rulebook and a coach telling them exactly how to behave, what to say, and how to be "helpful."

2. The Big Discovery: "More Aligned, Less Diverse"

The main finding is a bit surprising. You might expect that as AI gets smarter and more "aligned" with human instructions, it would start sounding more like a human writer.

The Reality: The newer AI models actually sound less like humans and less like their own older versions.

  • The Humans: Professional journalists haven't changed much. Their writing style, sentence structure, and word choices have remained stable and diverse over the last two years. They are like a jazz band that keeps improvising with new ideas.
  • The Older Robots: These models were actually quite diverse. They used a wide variety of sentence structures and words, sometimes even more varied than the humans in terms of vocabulary.
  • The Newer Robots: These models have become boringly consistent. They have lost their "flavor."
    • Syntactic Diversity (Structure): They use fewer types of sentence structures.
    • Lexical Diversity (Words): They use a much narrower range of words. They avoid specific names, dates, and unique details.

3. The "Safety Filter" Effect

Why did the robots change? The paper suggests it's because of Instruction Tuning.

Imagine the older robots were wild painters who threw paint everywhere. The newer robots are painters who have been told, "Don't make mistakes, don't hallucinate, and stick to the facts."

To follow these rules, the newer AI models started playing it safe.

  • They stopped using specific names (like "Washington" or "Senator Ben Cardin") because they didn't want to get the facts wrong.
  • They stopped using complex sentence structures that might be risky.
  • Instead, they started using long, safe, repetitive sentences that are easy to understand but lack the "spark" of human writing.

4. The "Easy to Read" Paradox

Here is a strange twist: The newer AI models write longer sentences than humans, but those sentences are easier for a computer to analyze.

  • The Analogy: Imagine a human writer building a complex, winding castle with hidden doors and secret passages. It's beautiful, but hard to map out.
  • The newer AI builds a very long, straight hallway with no doors. It's longer than the human's castle, but it's incredibly easy to walk down because it's so predictable.

The researchers found that the newer AI text is so regular and formulaic that a computer grammar checker can process it almost perfectly (99% success), whereas human writing is slightly more chaotic and harder to parse (94-95% success).

5. The Conclusion

The paper concludes that while newer AI models are better at following instructions and sounding "helpful," they have traded creativity and variety for safety and predictability.

They are no longer "average" writers; they are becoming "safe" writers. They have narrowed their expressive range so much that they are now less diverse than both human journalists and the older, less-trained AI models.

In short: The robots got better at following the rules, but in doing so, they forgot how to be interesting.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →