← Latest papers
📈 economics

Agent-Facing Information Design in LLM Tool Registries

This paper reveals that legal puffery, rather than fabricated claims, is the primary driver of bias in LLM tool selection, demonstrating that current disclosure methods fail and proposing a registry-layer design that separates structured capability descriptions from marketing copy to ensure agent accountability.

Original authors: Haochuan Kevin Wang

Published 2026-05-26
📖 5 min read🧠 Deep dive

Original authors: Haochuan Kevin Wang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Picture: The "Wild West" of AI Tool Stores

Imagine a massive digital marketplace where AI assistants (the "agents") go to pick tools to do jobs for you. Maybe you ask your AI to "find the best weather forecast," and it has to choose between five different weather apps.

Right now, this marketplace is like a Wild West flea market with no rules.

  • The Problem: The tool owners write their own descriptions. They can use fancy words, superlatives ("The #1 Best Tool!"), and big claims to trick the AI into picking them.
  • The Missing Piece: Unlike human shopping, where we have ratings, verified reviews, or "Sponsored" labels that actually work, this AI marketplace has no measurement system. There is no way to check if a tool is actually good or if the description is just a lie.

The researchers ran over 17,700 experiments to see how easily these AI agents can be swayed by marketing fluff.


Key Findings: What They Discovered

1. The "Hype" Wins Every Time

The researchers found that if two tools do the exact same job, but one has a boring description and the other has a "hyped" description (using words like "fastest," "most accurate," or "trusted by millions"), the AI picks the hyped one 100% of the time.

  • The Analogy: Imagine two identical cars parked side-by-side. One has a plain sign saying "Car." The other has a giant neon sign saying "The World's Fastest, Most Reliable Car!" The AI doesn't care that the cars are identical; it blindly drives toward the neon sign.
  • The Result: Just changing the words (copywriting) captures all the traffic. The AI is essentially "clickbait-prone."

2. Lying Doesn't Help (But Hype Does)

The team tested if making up fake numbers (like "Used by 12 million developers" when it's not true) made the AI pick the tool more often than just using legal exaggeration (puffery).

  • The Finding: No. The AI was just as likely to pick the tool with legal exaggeration as the one with fake lies.
  • The Takeaway: The "lie" didn't add any extra power. The problem isn't that the AI is being tricked by lies; it's that the AI is reacting to the style of the language (superlatives) rather than the facts. This means regulators can't just ban "fake ads" to fix the problem, because the "legal" hype is already doing 100% of the damage.

3. Warning Labels Don't Work

The researchers tried to "fix" the problem by adding warnings, similar to how we see "Sponsored" labels on Google or star ratings on Amazon.

  • The Result: It failed completely.
    • For most AIs: The warning was ignored. The AI still picked the hyped tool. It's like putting a "Caution: This ad is paid for" sticker on a neon sign, and the AI still runs toward the light.
    • For one AI (Claude): It overreacted. When it saw a "Sponsored" label, it actually avoided the tool, even if it was the best one.
  • The Conclusion: You can't fix this by just telling the AI "be careful." The system is broken, not the AI's brain.

4. The "Prisoner's Dilemma" (The Race to the Bottom)

Because the hyped tool wins every time, every tool owner is forced to write hyped descriptions just to survive.

  • The Analogy: Imagine a job interview where everyone is qualified. If one candidate wears a suit and says "I'm the best," they get the job. So, everyone else starts wearing a suit and saying "I'm the best." Eventually, everyone is wearing a suit and shouting, so the interviewer can't tell who is actually good.
  • The Result: The marketplace becomes useless. The descriptions stop telling you anything about quality because everyone is just shouting the same marketing slogans.

The Proposed Solution: The "Menu" vs. The "Flyer"

The paper suggests a complete redesign of how these tool stores work. They propose splitting the information into two separate buckets:

  1. The "Selection Menu" (For the AI):

    • This is what the AI sees before it picks a tool.
    • Rule: No marketing words allowed. No "best," no "fastest," no "trusted by."
    • Format: Just cold, hard facts in a structured list (e.g., "Input: Text. Output: 10 links. Speed: 200ms").
    • Why: The researchers found that AIs actually prefer this boring, structured format over flowery prose. It helps them pick the right tool without being distracted by hype.
  2. The "Marketing Flyer" (For You, the Human):

    • This is what the human sees after the AI has already picked the tool.
    • Rule: The tool owner can write all the fancy marketing copy they want here.
    • Why: The AI has already done its job based on the facts. Now, you (the human) can read the "Sponsored" pitch to decide if you like the tool.

Summary in One Sentence

AI agents are currently being fooled by marketing hype because tool registries are unregulated advertising platforms; the solution isn't to teach the AI to ignore lies, but to force the registries to show the AI a boring, fact-only menu while saving the fancy sales pitch for the human user.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →