LLM Predictive Scoring and Validation: Inferring Experience Ratings from Unstructured Text
This study demonstrates that GPT-4.1 can accurately infer baseball fans' overall experience ratings from unstructured text with high consistency, revealing that while the AI captures salient, emotionally intense moments, its systematic underestimation compared to self-reported scores reflects a fundamental difference between the model's focus on memorable events and the fans' holistic evaluative judgment.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a baseball team manager. You want to know if your fans had a good day. You ask them two questions after the game:
- The Verdict: "On a scale of 0 to 10, how was your day?"
- The Story: "Tell us why you gave that number."
Usually, you look at the number (the Verdict) to decide if you're doing well. But what if you could read the Story and guess the number before you even saw it? That's exactly what this paper did. They used a super-smart AI (a Large Language Model) to read thousands of fan stories and guess the score they would give.
Here is the simple breakdown of what they found, using some everyday analogies.
1. The "Crystal Ball" is Pretty Good
The AI read the text and guessed the score. It was surprisingly accurate.
- The Result: About two-thirds of the time, the AI's guess was within just one point of the actual score the fan gave.
- The Analogy: Imagine a weather forecaster trying to guess the temperature based on how people are dressed. If the fan says, "It was a bit chilly," the AI guesses 6/10. If the fan says, "It was perfect," the AI guesses 9/10. It gets the general vibe right most of the time.
2. The AI is a "Rock-Solid" Reader
The researchers ran the AI three times on the exact same stories to see if it would change its mind.
- The Result: It almost never changed its mind. If it guessed a "7" the first time, it guessed a "7" the next two times.
- The Analogy: Think of the AI like a very consistent judge. If you show a judge a painting three times in a row, they will give it the same score every time. They aren't guessing; they are reading the text with laser focus.
3. The Big Twist: The "Gap" is Real (and Useful)
Here is the most interesting part. The AI's guesses were systematically lower than the fans' actual scores.
- The Scenario: A fan gives a 10/10 (Perfect!). But in the text, they write: "The game was great, but the parking cost $40 and the line for hot dogs was 40 minutes long."
- The AI's Reaction: The AI reads the complaints about parking and hot dogs and guesses a 6 or 7.
- The Reality: The fan did give a 10. But the AI didn't make a mistake. It just heard a different part of the story.
4. Why the Gap Exists: The "Verdict" vs. The "Highlight Reel"
The paper explains that the fan and the AI are doing two different mental jobs.
- The Fan's Verdict (The Score): This is like a final grade for the whole day. It includes the excitement of the win, the love for the team, the memories of childhood, and the fact that you had a beer. It's a "gut feeling" that sums everything up, even the boring good parts that don't need mentioning.
- The Fan's Story (The Text): This is like a highlight reel of the drama. Humans are wired to talk about what went wrong or what was weird. We don't usually write, "The seats were comfortable and the sun was nice." We write, "The parking was a nightmare!"
- The AI's Job: The AI reads the "Highlight Reel of Drama." It sees the parking complaint and thinks, "Oh, this was a bad experience." It doesn't feel the fan's underlying love for the team that pushed the score to a 10.
The Metaphor:
Imagine you go on a date.
- The Verdict (Score): You give it a 9/10. You had a great time overall.
- The Story (Text): You tell your friend, "The waiter spilled soup on my shirt and the traffic was terrible."
- The AI: Reads your story and says, "Sounds like a 4/10."
The AI isn't wrong about the soup or the traffic. You aren't wrong about the date. You are just reporting on different things.
5. Why This Matters for Teams
The paper argues that teams shouldn't try to fix the AI to make it match the score perfectly. Instead, they should use the gap between the two as a warning system.
- If the Score is high (10) but the AI guess is low (6): This is "Fragile Satisfaction." The fans love the team, but they are annoyed by specific things (parking, prices, lines). They are forgiving now, but if things get worse, their love might turn into anger.
- If the Score and AI match: The experience is consistent.
The Takeaway:
The AI is a tool that tells you what fans are talking about (the friction points). The survey score tells you how much they still love you.
- The Score answers: "Are we doing okay?"
- The AI Guess answers: "What is hurting our fans right now?"
By looking at both, a team can see that while fans are happy, they are getting annoyed by parking costs. If they only looked at the score, they would think everything is fine and miss the problem until it's too late.
Summary
The AI is a reliable, consistent reader that can guess a fan's mood based on their complaints and compliments. It doesn't always match the fan's final score because the fan's score includes their love for the team, while the AI only sees the specific problems the fan wrote down. That difference isn't an error; it's a secret signal that helps teams fix problems before fans stop coming.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.