← Latest papers
💬 NLP

How Human-Like Are Large Language Models? A Register-Aware Linguistic Evaluation Framework

This paper introduces a register-aware evaluation framework using Maximum Mean Discrepancy and Biber's lexico-grammatical features to demonstrate that while large language models consistently deviate from human linguistic patterns, their degree of human-likeness varies by communicative register and is not determined by model size.

Original authors: Björn Nieth (Department Artificial Intelligence in Biomedical Engineering, Chair of AI-supported Therapy Decisions LMU München Munich Germany), Marianna Gracheva (Department of Digital Humanities and
Published 2026-05-25
📖 4 min read☕ Coffee break read

Original authors: Björn Nieth (Department Artificial Intelligence in Biomedical Engineering, Chair of AI-supported Therapy Decisions LMU München Munich Germany), Marianna Gracheva (Department of Digital Humanities and Social Studies), Michaela Mahlberg (Department of Digital Humanities and Social Studies, University of Birmingham United Kingdom), Bjoern Eskofier (Department Artificial Intelligence in Biomedical Engineering, University of Birmingham United Kingdom, Munich Center for Machine Learning, Institute of AI for Health Helmholtz Zentrum München Neuherberg Germany), Emmanuelle Salin (Department Artificial Intelligence in Biomedical Engineering)

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot to write a story, send a text message, or write a news report. You might ask, "Is what the robot wrote true?" or "Did it finish the task?" But this paper asks a different, more subtle question: "Does what the robot wrote feel like it was written by a human?"

The authors argue that just because a robot's text is factually correct doesn't mean it sounds natural. To humans, language changes depending on the situation. A text message to a friend sounds very different from a scientific paper or a news article. The authors call these different situations "registers."

Here is a simple breakdown of their study:

1. The Problem: The Robot's "Uncanny Valley"

Think of language like a costume party.

  • Humans are great at changing costumes. If you walk into a library, you whisper and use formal words. If you walk into a bar, you shout and use slang. You instinctively know which "costume" (register) fits the room.
  • Large Language Models (LLMs) are like actors who are very good at memorizing lines but sometimes forget to change their costume. They might write a news report that sounds like a casual chat, or a story that sounds like a textbook. Even if the facts are right, the "vibe" is off, making it feel unnatural to a human reader.

2. The Solution: A "Style Detector"

The researchers built a new way to test these robots. Instead of asking a human to read and guess if a text is real (which is slow and subjective), they created a mathematical "style detector."

  • The 67 Features: They used a classic linguistic tool (from a researcher named Biber) that breaks language down into 67 specific "ingredients." These aren't just words; they are patterns like "how often do people use past tense?" "Do they use short words or long words?" "Do they use passive voice?"
  • The Human Baseline: They gathered a huge pile of real human writing for five different "rooms" (registers):
    1. Casual Chat: Real conversations between people.
    2. Academic Papers: Introductions to computer science research.
    3. How-To Guides: Step-by-step instructions (like WikiHow).
    4. Creative Stories: Fiction written by users.
    5. News Reports: BBC news articles.
  • The Test: They asked seven different AI models to write in these same five styles. Then, they used a mathematical formula (called MMD) to measure the "distance" between the AI's writing style and the Human writing style.

The Analogy: Imagine you have a jar of real blueberries (Human writing) and a jar of artificial blueberry candies (AI writing). The researchers don't just taste them; they measure the exact shade of blue, the texture, and the sugar content. They calculate exactly how far the candy is from the real fruit. If the distance is huge, the candy tastes fake.

3. What They Found

The results were clear and surprising:

  • No Robot is Perfect: In every single test, the AI writing was statistically "far away" from how humans actually write. None of the robots could perfectly mimic the human style.
  • Size Doesn't Matter: You might think a bigger, more powerful robot would sound more human. The study found that bigger models are not necessarily better. A smaller model (Qwen 8B) sometimes sounded more human than a giant one (Llama 70B), depending on the topic.
  • Context is King: A robot might be great at writing a news report but terrible at writing a casual chat. You cannot judge a robot's "human-likeness" with a single score; you have to test it in every specific situation.
  • The "Family" Trait: Robots from the same "family" (like the Llama family or the Gemma family) tended to sound similar to each other, regardless of their size. They seemed to have a shared "accent" or "personality" that stuck with them.

4. Why This Matters (According to the Paper)

The paper concludes that we need to stop looking at AI just as a "task-completer" and start looking at it as a "style-matcher."

If an AI is going to be used in the real world (like in chatbots or content creation), it needs to fit the specific "room" it is in. The authors suggest that in the future, we could use this "distance" measurement to train robots to sound more natural, essentially teaching them to put on the right costume for the right party.

In short: The paper proves that while AI is getting smarter, it still hasn't mastered the art of sounding like a human in every situation. It's like a very well-read student who knows the facts but still sounds a bit stiff and unnatural when trying to chat with friends.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →