← Latest papers
💬 NLP

Em-ergence of the em-dash: a population-level rise in em-dash frequency in medRxiv preprints at the dawn of the large-language-model era

This pre-registered study analyzes over 69,000 medRxiv preprints to reveal a significant, gradual population-level increase in em-dash usage within discussion sections following the advent of large language models, suggesting a distinct stylistic shift in scientific writing during the early 2020s.

Original authors: Przemysław Czuma

Published 2026-06-30
📖 4 min read☕ Coffee break read

Original authors: Przemysław Czuma

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to figure out if a massive library of medical notes has been touched by a new, invisible hand. You aren't looking for fingerprints or DNA; you are looking for a specific, tiny punctuation mark: the em-dash (—).

Here is the story of what the paper found, told simply.

The Mystery: A "Typo" That Isn't

For years, scientists have whispered a rumor: "Large Language Models" (like ChatGPT) love to use em-dashes. Humans, on the other hand, usually find them annoying or hard to type, so they use them sparingly. It's like how a human might write a sentence with a comma, but a robot might write it with a dramatic dash to make it sound more "professional."

But until now, this was just gossip. No one had actually counted the dashes in thousands of real scientific papers to see if the rumor was true.

The Investigation: Counting Dashes in a Sea of Text

The researcher, Przemysław Czuma, decided to play detective on a massive scale. He didn't look at published books (which often get "cleaned up" by editors and lose their original punctuation). Instead, he looked at medRxiv, a server where scientists post their raw, unedited drafts (preprints) before they are published.

  • The Target: He focused on the "Discussion" section of these papers. This is the part where scientists tell a story, explain their feelings about the results, and try to persuade the reader. It's the most "human" part of a paper.
  • The Tool: He used a computer program to scan over 69,000 of these drafts, looking specifically for the em-dash character (—).
  • The Timeline: He split the data into two eras:
    1. Before ChatGPT (before November 2022).
    2. After ChatGPT (after November 2022).

The Findings: A Slow Burn, Not a Flash

The results were dramatic, but they didn't happen overnight like a light switch flipping on.

  • Before the AI boom: Only about 4% of the papers had even one em-dash in their Discussion section. It was rare.
  • After the AI boom: That number jumped to 11.5%.
  • The Trend: It didn't happen immediately in 2023. The number stayed low, then started creeping up in 2024, and by 2025, it had skyrocketed to 20%.

The Analogy: Imagine a quiet party where almost no one is wearing red hats. Then, a new fashion trend starts. At first, only a few people try it. Then, slowly, more people join in. By the next year, almost one in five people at the party is wearing a red hat. That's what happened with the em-dash.

The Proof: Ruling Out Other Suspects

A good detective checks if there are other explanations. The researcher ran several "lie detector" tests to make sure the em-dash increase wasn't just a random trend or a change in how the website formats text.

  1. The "Boring Section" Test: He checked the "Acknowledgments" and "Data Availability" sections (the boring, standard parts of a paper). If the whole website changed its formatting, these sections would have more dashes too. They didn't. The increase was only in the storytelling parts.
  2. The "Fake Date" Test: He looked at the time before ChatGPT existed and pretended a random date was the "start." Nothing changed. This proved the increase wasn't just a slow, natural drift over time.
  3. The "Word Choice" Test: He also looked for specific words that AI models love to use (like "delve," "groundbreaking," or "leverage"). These words rose at the exact same time as the dashes.

What This Means (and What It Doesn't)

The paper is very careful about what it claims.

  • It DOES show: Something big changed in how scientific papers are written between 2022 and 2025. The style of writing shifted in a way that matches the rise of AI tools. The em-dash is a "population-level" signal, like noticing that the average height of a crowd increased, not that every single person grew taller.
  • It DOES NOT show: You cannot look at a single paper with an em-dash and say, "This was written by a robot." A human might just love dashes. You also cannot look at a paper without a dash and say, "This is 100% human."
  • The Conclusion: The study confirms that the "fingerprint" of AI writing is becoming visible in the scientific literature. It didn't happen instantly; it happened as scientists slowly started trusting and using these new tools to polish and write their stories.

In short: The em-dash is the "smoke" that suggests a fire (AI usage) has started in the kitchen of scientific writing, even if we can't point to exactly which stove was used for every single pot.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →