← Latest papers
💬 NLP

Who Wrote This Line? Evaluating the Detection of LLM-Generated Classical Chinese Poetry

This paper introduces ChangAn, a comprehensive benchmark comprising over 30,000 classical Chinese poems, to evaluate the effectiveness of current AI detectors in distinguishing human-written poetry from LLM-generated works, revealing significant limitations in existing detection tools for this specific literary domain.

Original authors: Jiang Li, Tian Lan, Shanshan Wang, Dongxing Zhang, Dianqing Lin, Guanglai Gao, Derek F. Wong, Xiangdong Su

Published 2026-04-14
📖 5 min read🧠 Deep dive

Original authors: Jiang Li, Tian Lan, Shanshan Wang, Dongxing Zhang, Dianqing Lin, Guanglai Gao, Derek F. Wong, Xiangdong Su

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a world where a robot can write a poem so beautiful, so perfectly structured, and so full of ancient wisdom that it sounds exactly like it was written by a human master from 1,000 years ago. Now, imagine trying to tell the difference between a poem written by a tired human poet and one written by a super-smart computer in a split second.

That is the challenge tackled in this paper, "Who Wrote This Line?" by a team of researchers from China. They built a new "test track" called ChangAn to see if our current tools can catch AI poets in the act.

Here is the story of their discovery, explained simply:

1. The Problem: The "Uncanny Valley" of Poetry

For years, we've been getting better at spotting AI-written news articles or essays. But poetry is different. Think of classical Chinese poetry like a highly regulated dance.

  • The Rules: Every step (character) must fit a specific rhythm, rhyme, and tone pattern.
  • The Vocabulary: There is a limited set of "dance moves" (imagery) everyone uses, like "moon," "wine," "willow," and "sorrow."

Because the rules are so strict and the vocabulary is so shared, it's like trying to tell if a dancer is human or a robot when both are following the exact same choreography perfectly. Current AI detectors, which are like "security guards" trained on messy, everyday text, get confused. They can't tell if the strict rhythm is a human following tradition or a robot following code.

2. The Solution: Building the "ChangAn" Arena

The researchers realized they needed a specific gym to test these detectors. They built ChangAn, a massive dataset containing over 30,000 poems.

  • The Humans: 10,000+ poems written by modern-day poets who still love the old style.
  • The Robots: 20,000+ poems written by four famous AI models (like GPT-4, Kimi, and DeepSeek).

They didn't just ask the AI to "write a poem." They tested two scenarios:

  1. Direct Generation: The AI writes from scratch.
  2. Critique-Driven Refinement: The AI writes a poem, gets a list of errors (like "this rhyme is off"), and then fixes it. This is like a human editor polishing a draft.

3. The Big Reveal: The Detectors Are Mostly Blind

When they ran their tests, the results were shocking:

  • The "Security Guards" Failed: Most of the standard AI detectors (the ones we use for essays) were terrible at this. They guessed wrong almost as often as they guessed right. It was like asking a dog to identify a specific type of flower; the dog just smells the dirt.
  • The "Math Whizzes" Did Better (But Not Perfectly): Detectors that looked at the statistical patterns of words (like counting how often certain words appear together) did much better. One specific tool, a "fine-tuned" model named RoBERTa, was the star of the show, getting it right 95% of the time.
  • The "Self-Reflection" Trap: The researchers asked the AI models, "Did you write this poem?" Surprisingly, the AI models were terrible at recognizing their own work. They couldn't tell their own poetry from a human's. It's as if the AI was so good at mimicking the human style that it lost its own "robot fingerprint."

4. The "Volume" Effect: One Poem vs. A Book

Here is a fascinating twist: Context matters.

  • Single Poem: If you show a detector just one poem, it's very hard to tell if it's AI. The sample size is too small, like trying to guess a person's personality from a single sentence.
  • A Collection: If you show the detector a bundle of 6 or 12 poems, the AI becomes much easier to spot. The AI leaves a subtle "statistical trail" across many poems that a human doesn't. It's like a forger who can fake one signature perfectly, but if they forge 100 letters, their hand starts to shake in a pattern only a forensic expert can see.

5. The "Polishing" Problem

The researchers found that when they let the AI critique and fix its own poems, the detectors got even worse.

  • Analogy: Imagine a human writer writes a draft, then an AI editor fixes the grammar and flow. The result is so smooth and "perfect" that it actually hides the AI's tracks better than the raw draft did. This "Critique-Driven" method makes the AI poetry look even more human.

The Bottom Line

The paper concludes that we are not ready. Our current tools are not reliable enough to catch AI-generated classical poetry. The AI has become so good at following the strict rules of this ancient art form that it has blurred the line between human and machine.

However, the ChangAn dataset is now open for everyone. It's like giving the world a new set of magnifying glasses and a training manual, hoping that future researchers can build better "poetry detectives" before the robots take over the literary world entirely.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →