← Latest papers
💬 NLP

The signal is the ceiling: Measurement limits of LLM-predicted experience ratings from open-ended survey text

This paper demonstrates that while prompt customization can slightly improve LLM accuracy in predicting fan experience ratings from survey text, the fundamental performance ceiling is determined by the inherent linguistic content of the input and the gap between written feedback and actual decisions, rather than by model selection or extensive engineering.

Original authors: Andrew Hong, Jason Potteiger, Luis E. Zapata

Published 2026-04-22
📖 6 min read🧠 Deep dive

Original authors: Andrew Hong, Jason Potteiger, Luis E. Zapata

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a baseball team trying to understand how much your fans loved a game. You ask them two questions:

  1. The Number: "On a scale of 0 to 10, how was the game?" (This is the verdict).
  2. The Story: "Tell us what happened." (This is the evidence).

Usually, these two match up. If someone says "It was amazing!" and gives a 10, that's easy. But sometimes, a fan writes a story full of complaints about long lines and bad parking, yet still gives the game a 9 or 10. Why? Because they loved the feeling of the win, even if the logistics were annoying.

This paper is about teaching an AI (a Large Language Model, or LLM) to read those stories and guess the number the fan would have given. The researchers wanted to know: How good can the AI get, and what stops it from being perfect?

Here is the breakdown of their findings, using some simple analogies.

1. The "Ceiling" Analogy

The main idea of the paper is that there is a ceiling on how accurate the AI can be. You can't paint a picture that is better than the canvas allows.

  • The Canvas (The Text): The fan's story.
  • The Painter (The AI): The computer model.
  • The Ceiling: The limit of accuracy.

The researchers found that the text itself is the biggest factor. If the fan writes a clear, happy story, the AI is great (90%+ accurate). If the fan writes a messy story full of mixed feelings (happy about the win, angry about the hot dog), the AI struggles. No amount of "tweaking" the AI can fix this because the missing information (how much the fan really cared about the win vs. the hot dog) isn't in the text. It's in the fan's head.

2. The Two Parts of the Ceiling

The paper breaks this ceiling into two distinct parts, like a wall with two different types of bricks:

Brick A: The "Missing Information" Wall (The Cognitive Gap)

  • What it is: Humans are weird. We remember the peak moment (the home run) and the end (the final cheer) more than the boring stuff (waiting in line). We write down the boring stuff because it's easy to describe, but we give a high score because of the feeling.
  • Can AI fix it? No. The AI only sees the text. It doesn't have the fan's feelings. It sees "long lines" and thinks "bad experience." It can't know the fan thought, "The lines were long, but the game was worth it."
  • Analogy: It's like trying to guess a person's mood by reading their grocery list. If they bought ice cream and a movie ticket, you might guess they are happy. But if they also bought a heavy coat and a thermometer, you might guess they are sick. You don't know if they are having a "sick day at home" or a "cozy movie night." The list (text) doesn't tell you the whole story.

Brick B: The "Bad Habit" Wall (The Measurement Bias)

  • What it is: AI models are trained on millions of reviews from the internet. They learn a bad habit: "Negative words = Low Score." If a fan says "The parking was terrible," the AI automatically thinks, "Oh, they must hate the whole game," and gives a low score.
  • Can AI fix it? Yes. This is a bad habit the AI can unlearn with a little coaching.
  • The Fix: The researchers wrote a special "prompt" (a set of instructions) that told the AI: "Hey, stop just counting negative words. Look at the big picture. If they say 'Great game' but mention 'long lines,' don't punish them too hard."
  • Result: This "coaching" improved the AI's accuracy by about 2%. It didn't break the ceiling, but it helped the AI climb a little higher on the part of the wall it could reach.

3. The "Model Swap" Experiment

The researchers tried swapping the AI brains. They used a super-smart AI, a tiny AI, and a newer AI.

  • The Result: It didn't matter much. The super-smart AI didn't get much better, and the tiny AI got much worse.
  • The Lesson: You can't just buy a more expensive car to fix a bad road. If the "road" (the fan's text) is confusing, even the smartest AI will get lost. The instructions (the prompt) matter more than the engine (the model).

4. What Should Teams Actually Do?

Since the AI can't be perfect, how should teams use it?

  • Don't treat it as a replacement: Don't throw away your survey numbers and just use the AI's guess. They measure different things.
  • Use it as a "Red Flag" detector: The AI is actually really good at spotting two specific groups of fans:
    1. The "Fragile Fans": People who gave a high score (9 or 10) but wrote a story full of complaints. The AI might give them a lower score, which is a signal: "Hey, this fan is happy, but they are annoyed. Fix the lines before they leave!"
    2. The "Silent Critics": People who gave a low score but wrote a nice story.
  • Trust the "Confidence" meter: The AI can tell you how sure it is. If the AI says, "I'm 90% sure this is a 9," trust it. If it says, "I'm only 50% sure," put that one in a pile for a human to read.

Summary

The paper concludes that you can't engineer your way out of human psychology.

  • The Text is the limit. If the story is confusing, the guess will be off.
  • The Prompt is a tool. It helps the AI stop making stupid mistakes (like over-punishing negative words), but it can't read minds.
  • The Value isn't in getting the exact number right; it's in spotting the fans whose words and numbers don't match. Those are the fans who need your attention the most.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →