← Latest papers
🤖 AI

Business Utility of Large Language Models as Exploratory Data Analysis Agents

This paper evaluates the business utility of Large Language Models as exploratory data analysis agents using a supply chain simulation benchmark, revealing that most configurations lack the necessary repeatability for autonomous use despite acceptable average performance, while a specific high-effort GPT-5.4 configuration demonstrates the strongest balance of quality and reliability.

Original authors: Rafał Łabędzki, Patryk Miziuła, Hubert Rutkowski, Szymon Betlewski, Cezary Depta, Szymon Janowski, Jarosław Kochanowicz, Jan Kanty Milczek

Published 2026-06-02
📖 4 min read☕ Coffee break read

Original authors: Rafał Łabędzki, Patryk Miziuła, Hubert Rutkowski, Szymon Betlewski, Cezary Depta, Szymon Janowski, Jarosław Kochanowicz, Jan Kanty Milczek

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you hire a team of highly intelligent, super-fast detectives to solve a mystery: Which supplier sent the bad apples that caused a grocery store to lose money?

This paper is a report card on how well these "AI detectives" (Large Language Models) actually perform this specific job. The researchers didn't just ask, "Did they get the right answer?" They asked, "Can we trust them to get the right answer every single time we ask?"

Here is the breakdown of their findings using simple analogies.

1. The Job: Finding the "Bad Apples"

In the real world, businesses don't always have a label that says "This batch is rotten." Instead, they have to look at clues: maybe customers stopped buying a product, or stores returned items.

  • The Task: The AI had to look at messy data trails (like a detective looking at footprints) to figure out which specific supplier-product combination was the culprit.
  • The Challenge: This isn't a simple math problem with one right answer. It requires guessing, connecting dots, and forming a theory. This is called Exploratory Data Analysis (EDA).

2. The Test: The "Five-Run" Rule

The researchers didn't just let the AI try once. They made 15 different AI models (like GPT-5, Gemini, Claude, etc.) try to solve this mystery five times each, under four slightly different conditions (like giving them a clearer map, or a messier map).

They measured two things:

  1. Accuracy: Did they find the bad apples?
  2. Consistency: If you asked them the same question five times, did they give the same answer every time?

3. The Big Problem: The "Rollercoaster" Effect

The paper found a surprising truth: Most AI models are like rollercoasters.

  • Sometimes they zoom to the top and solve the mystery perfectly.
  • Other times, they crash and give a wrong answer.
  • Even if their average score looks okay, the fact that they are so unpredictable makes them dangerous to use in a real business.

The Analogy: Imagine hiring a chef. If they make a perfect steak 50% of the time and a burnt piece of charcoal the other 50%, their "average" meal might look okay on paper. But you wouldn't trust them to run your restaurant because you can't predict what you're getting. That is the problem with these AI agents right now.

4. The New Scorecard: "Business Utility"

The researchers invented a new way to grade the AI called Business Utility. Think of it as a "Trust Score."

  • It takes the AI's average accuracy and punishes it heavily if it is inconsistent.
  • The Formula: If an AI is smart but erratic, its score drops. If an AI is slightly less smart but very reliable, its score stays higher.

The Result:

  • The Winner: One model, GPT-5.4 (with extra "thinking" effort), was the clear champion. It was both smart and consistent. Its "Trust Score" was high.
  • The Losers: Most other models, even some that are usually very popular, were too unstable. They got good scores on average, but their "Trust Score" was low because they were too unpredictable.

5. The "Signal" Test

The researchers also tested if the AI got confused when the clues were harder to find (a "weaker signal").

  • Some models changed their behavior drastically when the clues got harder.
  • However, the paper found that internal inconsistency (the rollercoaster effect) was a bigger problem than getting confused by hard clues. Even when the clues were easy, the AI still couldn't agree with itself on the answer.

6. The Bottom Line

The paper concludes that while AI is getting better at solving these puzzles, we are not ready to let them work alone yet.

  • Current State: AI is like a brilliant intern who has great ideas but forgets to write them down half the time. You can't rely on them to run the show without a human double-checking every single step.
  • The Lesson: When companies want to use AI for data analysis, they shouldn't just look at the "average" success rate. They need to look at reliability. If the AI can't do the job consistently, it's not ready for the real world, no matter how smart it looks on a good day.

In short: The AI detectives are smart, but they are too moody to be trusted with the keys to the business yet. We need them to be consistent before we let them work without supervision.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →