← Latest papers
💬 NLP

Self-Training Doesn't Flatten Language -- It Restructures It: Surface Markers Amplify While Deep Syntax Dies

This paper challenges the notion that self-training merely flattens language by demonstrating that it instead restructures it through an asymmetric collapse where deep syntactic features decay while surface markers amplify, a phenomenon formalized as the Structural Depth Hypothesis.

Original authors: Ming Liu

Published 2026-05-21
📖 5 min read🧠 Deep dive

Original authors: Ming Liu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Misconception: The "Flattening" Myth

Imagine you have a talented storyteller (a Language Model). If you ask them to tell a story, then take that story, give it back to them, and ask them to tell a new story based on the first one, and repeat this process over and over, what happens?

Most researchers believed the storyteller would eventually get boring. They thought the stories would become "flatter"—less diverse, repetitive, and simple, like a photocopy of a photocopy that keeps losing detail. This is called "model collapse."

This paper says that's not quite right. The stories don't just get simpler; they get weirdly restructured. The storyteller doesn't lose all complexity; they lose the hard complexity while accidentally getting better at the easy complexity.

The Two Types of "Complexity"

To understand what's happening, the author splits language features into two buckets:

  1. Surface Markers (The "Decorations"): These are the flashy, easy-to-add things like "However," "Moreover," "Maybe," or using an em-dash (—). They sit on top of the sentence and don't require much brainpower to attach.
  2. Deep Syntax (The "Skeleton"): These are the hard structural bones of language. Things like asking a question ("Why did you do that?"), using the passive voice ("The ball was thrown"), or using the subjunctive mood ("If I were you..."). These require the sentence to hold multiple ideas together in a specific, nested order.

The Experiment: The "Copy-Paste" Loop

The author took five different AI models (including GPT-2) and ran them through an 11-round "copy-paste" loop:

  1. The AI writes text.
  2. The AI is trained on that text.
  3. The AI writes new text based on the training.
  4. Repeat 11 times.

The Result: The "Superficial Complexity Paradox"

Here is the surprising twist the paper discovered:

1. The Decorations Explode (They get "Richer")
The AI started using way more "However," "Maybe," and em-dashes. By the 11th generation, the text looked more "essay-like," more formal, and more connected. If you just counted these words, you would think the AI was getting smarter and more complex.

  • Analogy: Imagine a house that keeps adding more porch columns, fancy paint, and decorative bushes. From the outside, it looks grander and more elaborate.

2. The Skeleton Collapses (The "Deep" Stuff Dies)
While the decorations grew, the hard structural bones vanished.

  • Questions disappeared by 92%.
  • Parenthetical thoughts (side comments in the middle of a sentence) dropped by 57%.
  • Passive voice and complex verb tenses dropped by over 50%.
  • Analogy: While the house added more bushes, the foundation and the load-bearing walls started crumbling. The house looks fancy, but it can't actually support a second floor anymore.

The Paradox: The text looks more complex (longer words, more "fancy" connectors), but the actual sentence structure is becoming shallower and simpler.

Why Does This Happen? The "Depth Hypothesis"

The author proposes a rule called the Structural Depth Hypothesis.

Think of language features as having a "depth score":

  • Depth 0 (Surface): Just adding a word like "However." Easy.
  • Depth 2 (Deep): Building a question or a passive sentence. Hard.

The Rule:

  • Easy things get amplified: Because the AI is trained on its own output, it sees "However" a lot. It thinks, "Oh, I should use 'However' more!" It becomes a feedback loop where the AI loves these easy decorations.
  • Hard things get punished: To build a complex sentence (like a question), the AI has to make a chain of correct decisions. If it misses one step, the sentence breaks. In the "copy-paste" loop, the AI gets worse at making those long chains of decisions. It stops trying to build them because it's easier to just add a decoration.

The "Frequency" Myth:
Old theories said rare things die because the AI doesn't see them often enough. This paper says no. It's not about how rare the word is; it's about how structurally deep it is. Even if a complex sentence is common, if it's "deep," it will die. If a simple word is rare, it might still survive if it's "shallow."

The "Exclamation Mark" Exception

There is one weird case: Exclamation marks (!). They are "shallow" (Depth 0), so you'd expect them to grow. But they died (dropped by 99%).

  • Why? The AI's "default mode" (when it's not being creative) rarely uses exclamation marks. Even though they are easy to add, the AI never really saw them in its own "dreams" to begin with. So, they never got the "rich-get-richer" boost. This proves the rule: it's not just about being shallow; you also need to be common enough to start the loop.

The Takeaway for the Real World

The paper warns us about how we measure AI.

  • Current detectors look at "aggregate fingerprints" (like: "Is the text long? Does it have many different words?").
  • The Problem: Under self-training, these numbers go up. So, a detector might look at a broken, shallow AI text and say, "Wow, this looks very complex and human-like!"
  • The Reality: The text is actually a hollow shell. It has the decorations of a complex essay but none of the structural integrity.

In short: Self-training doesn't make AI dumb; it makes it superficially fancy but structurally hollow. It learns to decorate the house while forgetting how to build the walls.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →