← Latest papers
💬 NLP

Measuring AI "Slop" in Text

This paper addresses the lack of a standard definition for AI "slop" by developing a taxonomy and interpretable dimensions through expert interviews, demonstrating that while binary judgments are subjective, they correlate with latent factors like coherence and relevance to enable better evaluation of AI-generated text.

Original authors: Chantal Shaib, Tuhin Chakrabarty, Diego Garcia-Olano, Byron C. Wallace

Published 2026-01-27
📖 5 min read🧠 Deep dive

Original authors: Chantal Shaib, Tuhin Chakrabarty, Diego Garcia-Olano, Byron C. Wallace

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the internet is a giant, bustling marketplace. For a long time, vendors (humans) sold fresh, handcrafted goods. But recently, a flood of cheap, mass-produced items has arrived. They look okay from a distance, but up close, they feel flimsy, repetitive, and hollow. In the world of AI, people have started calling this low-quality, AI-generated junk "slop."

This paper is like a team of experts trying to build a quality control manual for this slop. They want to answer two big questions: What exactly makes text feel like "slop"? and Can we build a machine to spot it automatically?

Here is the breakdown of their findings, using simple analogies:

1. Defining the "Slop" (The Taxonomy)

The researchers didn't just guess what slop is. They interviewed 19 experts—writers, journalists, philosophers, and AI scientists—to get a clear definition. They realized "slop" isn't just one thing; it's a mix of three main problems, like a bad meal that is too salty, too bland, and served on a dirty plate.

They organized these problems into a checklist called a Taxonomy:

  • Information Utility (The "Empty Bowl"):
    • Density: The text is full of words but has very little actual meat. It's like a soup that is 90% water and 10% broth.
    • Relevance: The text talks about things that don't matter to the question asked. It's like ordering a burger and getting a lecture on the history of cows.
  • Information Quality (The "Spoiled Ingredients"):
    • Factuality: The text makes things up or gets facts wrong (hallucinations).
    • Bias: The text sounds too robotic or neutral when it should have a human voice, or it takes a weird, one-sided stance.
  • Style Quality (The "Bad Presentation"):
    • Repetition & Templatedness: The AI keeps saying the same thing or uses the exact same sentence structure over and over, like a broken record.
    • Verbosity: It uses 100 words to say what could be said in 10. It's "wordy" without being "wise."
    • Coherence & Tone: The ideas don't flow logically, or the tone feels weird (like a robot trying to tell a joke).

2. The Human Test (Annotating the Slop)

To see if their checklist worked, they hired professional copy-editors to read 150 news articles and 100 Q&A passages. They asked the editors to highlight specific "sloppy" parts of the text.

The Surprise:
Even the experts didn't always agree on what was "slop."

  • The Subjectivity Problem: One editor might think a text is "slop" because it's too wordy, while another thinks it's fine.
  • The Connection: However, when the editors did agree a text was slop, it was usually because the text was missing relevance (didn't answer the prompt) or density (was too empty).
  • The Domain Difference:
    • In News, slop was mostly about bad style and tone (too formal, too repetitive).
    • In Q&A, slop was mostly about getting facts wrong or being structurally messy.

3. Can Machines Spot It? (The Automatic Check)

The researchers tried to see if current AI tools could automatically flag this slop without human help. They tested three things:

  1. Standard Math Metrics: They used old-school math tools (like counting word length or checking for repeated words).
    • Result: These tools were okay at spotting "wordy" text, but they failed to catch the subtle stuff like "is this relevant?" or "does this make sense?"
  2. AI Reward Models: They used advanced AI models trained to grade writing quality.
    • Result: These models could tell the difference between "good" and "bad" writing generally, but they couldn't specifically identify "slop" as defined by the humans. They missed the nuance.
  3. AI Judges (LLMs as Judges): They asked powerful AI models (like GPT-5 and others) to read the text and say, "Is this slop?" and "Highlight the bad parts."
    • Result: They failed miserably. The AI judges were terrible at this. They often missed the slop entirely or highlighted the wrong parts. They tended to focus only on "wordiness" and ignored everything else.

The Bottom Line

The paper concludes that while we have a good human-defined checklist for what "slop" looks like (it's mostly about being irrelevant, empty, or repetitive), we do not yet have a reliable robot that can automatically find it.

Currently, if you want to know if text is "slop," you still need a human to read it and use their judgment. The automatic tools are like a metal detector that only finds big rocks but misses the tiny, valuable gems (or the tiny, annoying pieces of trash) hidden in the sand.

What the paper does NOT claim:

  • It does not say we should ban AI writing.
  • It does not claim this method will fix AI in the future (it just says it's an open challenge).
  • It does not suggest using this for medical or legal decisions.
  • It strictly focuses on defining the problem and showing that current automated solutions aren't ready to solve it yet.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →