← Latest papers
🤖 AI

A Framework for Studying AI Agent Behavior: Evidence from Consumer Choice Experiments

This paper introduces ABxLab, a framework for systematically evaluating LLM-powered AI agents in consumer choice scenarios, revealing that these agents exhibit predictable, bias-like decision shifts in response to manipulated attributes and cues, thereby highlighting both the risks of inherited human biases and the opportunity to establish a behavioral science of AI.

Original authors: Manuel Cherep, Chengtian Ma, Abigail Xu, Maya Shaked, Pattie Maes, Nikhil Singh

Published 2026-02-25
📖 5 min read🧠 Deep dive

Original authors: Manuel Cherep, Chengtian Ma, Abigail Xu, Maya Shaked, Pattie Maes, Nikhil Singh

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you hire a super-smart digital assistant to go shopping for you. You tell it, "Find me the best headphones," and you walk away, trusting it to make the smart choice. You assume it will look at the quality, the price, and the reviews, and pick the winner based on logic.

But what if your assistant isn't actually thinking like a rational human? What if it's just as easily tricked by a flashy sign, a "limited edition" sticker, or the order in which items are listed as a distracted shopper might be?

This is exactly what the paper "A Framework for Studying AI Agent Behavior" investigates. The researchers built a digital playground called ABXLAB to test how AI shopping bots actually make decisions when faced with the same psychological tricks that influence us humans.

Here is the breakdown of their findings in simple terms:

1. The Setup: A Digital "Mall" with Magic Mirrors

The researchers created a controlled online shopping environment (like a digital mall). But they added a twist: they built a "Magic Mirror" (a man-in-the-middle framework) that sits between the AI and the website.

Before the AI sees a product, the Magic Mirror can instantly change the reality:

  • It can add a sticker saying "Best Seller!" (Social Proof).
  • It can say "Only 1 left!" (Scarcity).
  • It can claim "Recommended by Experts!" (Authority).
  • It can even swap the price or the star rating.

They then asked 17 different top-tier AI models to choose between two products in thousands of different scenarios.

2. The Big Surprise: AI is More Gullible Than Humans

The most shocking finding is that AI agents are incredibly easy to manipulate, often much more so than actual humans.

  • The Human Baseline: When real people were shown these tricks, they were slightly influenced, but mostly they stuck to their logic. If a product was expensive but had bad reviews, they usually skipped it.
  • The AI Reality: The AI agents were like a child seeing a "Free Candy" sign.
    • The "Best Seller" Trap: If the AI saw a "Best Seller" badge, it would choose that product over a better one with a 90% chance, even if the other product was objectively better.
    • The "Price" Panic: When the AI couldn't decide based on ratings, it would almost blindly pick the cheaper option, sometimes ignoring quality entirely.
    • The "Order" Effect: Just like humans, if an item was listed first, the AI was more likely to pick it. But for some AIs, this effect was massive—almost like they were programmed to always pick the first thing they saw.

The Metaphor: Imagine a human shopper and a robot shopper walking into a store. A human might glance at a "Sale" sign but still check the price tag. The robot, however, sees the "Sale" sign and immediately grabs the item without looking at the tag, as if the sign is a command it cannot disobey.

3. Why Does This Happen?

The researchers found that these AI agents don't have "common sense" or "bounded rationality" (the human limitation that stops us from overthinking). Instead, they seem to follow rigid, brittle rules learned from their training data.

  • The "Rulebook" Problem: The AI learned that in human language, "Best Seller" usually means "Good." So, it treats that phrase as a mathematical law rather than a marketing tactic.
  • No Cognitive Fatigue: Humans get tired of reading; we use shortcuts. The AI doesn't get tired, but it also doesn't have the human ability to say, "Wait, this 'limited edition' claim sounds suspicious." It just follows the cue.

4. The "User Profile" Switch

The researchers also tested what happens if you tell the AI, "I am a budget shopper" or "I only care about expert reviews."

  • The Result: The AI didn't just adjust slightly; it flipped a switch. If you told it to care about experts, it completely ignored price and quality ratings. It acted like a robot following a strict instruction manual, rather than a flexible assistant weighing trade-offs.

5. Why Should You Care? (The Risk and the Opportunity)

The Risk:
As we start letting AI agents buy our groceries, book our flights, or manage our investments, we are handing over control to entities that are hypersensitive to manipulation. If a bad actor knows how to trick an AI (e.g., by adding a fake "Expert Recommended" badge), they could steer millions of dollars of spending toward their products, even if the product is junk. The AI might amplify human biases rather than fix them.

The Opportunity:
The good news is that we now have a way to test this. Just as psychologists study human behavior to understand why we make bad choices, this paper gives us a "Behavioral Science Lab" for AI. We can now systematically test how different AIs react to different tricks, fix their weaknesses, and build agents that are robust against manipulation before we let them loose in the real world.

The Bottom Line

We are building AI agents to do our bidding, but we haven't yet taught them how to ignore the "salesman's tricks" of the internet. This paper sounds the alarm: AI is currently too easily swayed by superficial cues. Before we trust them with our money and our lives, we need to teach them to see through the marketing fluff and make decisions based on what actually matters.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →