← Latest papers
💬 NLP

DRIP-R: A Benchmark for Decision-Making and Reasoning Under Real-World Policy Ambiguity in the Retail Domain

This paper introduces DRIP-R, a novel benchmark designed to evaluate LLM-based agents' decision-making and reasoning capabilities in the retail domain by systematically testing their performance on real-world scenarios characterized by inherent policy ambiguities where no single correct resolution exists.

Original authors: Hsuvas Borkakoty, Sebastian Pohl, Cheng Wang, Bei Chen, Yufang Hou

Published 2026-05-11
📖 4 min read☕ Coffee break read

Original authors: Hsuvas Borkakoty, Sebastian Pohl, Cheng Wang, Bei Chen, Yufang Hou

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a customer service robot working for a giant online store. Your job is to handle returns. You have a rulebook (the company policy) that tells you what to do.

In most robot tests, the rulebook is like a math textbook: clear, precise, and with only one right answer. "If the box is open, you can't return it." Simple.

But in the real world, rulebooks are more like a messy, handwritten note from a manager. They say things like, "Items must be in unused condition."

Here is the problem: What does "unused" actually mean?

  • Does it mean the plastic wrap is still on?
  • Does it mean you can't have taken it out of the box?
  • Does it mean you can't have plugged it in for 10 seconds to see if it works?

This is ambiguity. And this is exactly what the paper DRIP-R is all about.

The Big Idea: The "Gray Area" Test

The authors created a new test called DRIP-R (Decision-making and Reasoning In ambiguous Policy for Retail). Instead of giving robots clear rules, they gave them real-world Amazon return policies full of gray areas.

Think of it like a driving test where the road signs are blurry.

  • Old Tests: "Stop at the red light." (Easy. The robot stops.)
  • DRIP-R Test: "Stop if the light is kind of red or if it looks like it might turn red soon." (Hard. Does the robot stop? Does it slow down? Does it keep going?)

How They Tested the Robots

They set up a simulation with two characters:

  1. The Customer: A simulated person with a specific personality (some are angry, some are polite, some are nervous).
  2. The Agent: The AI customer service bot trying to solve the return problem.

They gave the agents a list of 40 tricky return scenarios (like "I opened the headphones but only listened to one song") and 10 different customer personalities. The agents had to chat with the customers, check their internal tools (like looking up order details), and decide: Do I give a full refund? A partial one? Or do I say no?

The "Multi-Judge" Panel

How do you grade a robot when there is no single right answer? If the robot says "Yes, refund," and another says "No, deny," who is right?

The authors didn't just pick one winner. They built a panel of AI judges (like a jury) to evaluate the robots on four different things:

  1. Did they stick to the rules? (Even if the rules were fuzzy, did they try to follow the spirit of the policy?)
  2. Did they talk well? (Was the conversation polite and logical?)
  3. Did they act like their character? (If the robot was supposed to be "Very Helpful," did it sound eager? If it was "Direct," was it blunt?)
  4. Did they solve the problem? (Did the customer get what they needed, and did the company stay safe?)

What They Found (The Surprises)

The results were eye-opening. When the rules were clear, all the smartest AI models agreed. But when the rules were ambiguous:

  • They all disagreed: For the exact same return request, one AI might say "Full Refund," another might say "Partial Refund," and a third might say "Deny." There was almost no agreement between them.
  • Personality matters: If the customer was "Neurotic" (anxious) or "Extroverted" (loud), the AI was more likely to give them a refund. If the customer was "Agreeable" (nice), the AI was stricter. This suggests the robots might be unfair, favoring certain types of people over others.
  • The "Say-Do" Gap: Sometimes, the robot would say, "I am following the policy strictly," but then give a refund anyway. They were good at sounding like they were following rules, but bad at actually doing it consistently.

The Takeaway

The paper concludes that current AI agents are great at following clear instructions, but they are terrible at navigating the messy, gray areas of real life.

Imagine a robot that is a perfect lawyer but a terrible judge. It can recite the law perfectly, but when the law is vague, it doesn't know how to make a fair decision. The authors built DRIP-R to show us exactly where these robots stumble, so we can build better ones that can handle the messy reality of human rules.

In short: Real life isn't a math problem with one answer. It's a conversation with many possible outcomes. DRIP-R proves that our current AI robots are still struggling to have that conversation fairly.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →