← Latest papers
💻 computer science

Divergent Recommendations, Convergent Diagnoses: Cross-Provider Failure-Mode Convergence in AI Commercial Recommendation

Despite significant divergence in specific product recommendations between ChatGPT and Claude, the study reveals that these AI providers exhibit near-perfect agreement (95.1%) in diagnosing the underlying reasons for joint recommendation failures, suggesting that while optimization strategies may need to be provider-specific for top brands, a unified approach to addressing root causes is highly effective for long-tail brands.

Original authors: Will Jack, Noah Lehman, Keller Maloney, Sarah Xu

Published 2026-06-26
📖 5 min read🧠 Deep dive

Original authors: Will Jack, Noah Lehman, Keller Maloney, Sarah Xu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Question: Do AI Assistants Agree?

Imagine you are a famous brand (like Nike or a local bakery). You want your products to be recommended when people ask AI assistants like ChatGPT (OpenAI) or Claude (Anthropic) for advice.

The big question the authors asked is: If I fix my marketing for ChatGPT, will it automatically work for Claude? Or do I need two completely different strategies?

To find out, they ran a massive experiment. They asked both AIs the same 215 questions about products and watched what they recommended.

The First Finding: They Pick Different Winners

The Analogy: Imagine two different food critics (Chef A and Chef B) reviewing a menu of 500 dishes.

  • The Result: When asked to pick their top 10 dishes, Chef A and Chef B only agreed on about 3 or 4 dishes. The other 6 or 7 dishes were totally different.
  • The Paper's Data: The two AIs agreed on which brands to recommend only about 35% of the time. If you run the same question twice on the same AI, it agrees with itself about 50–60% of the time. But if you switch to a different AI, the recommendations change significantly.

Takeaway: If you are a big brand, you cannot assume that what works for one AI will work for the other. They have different "tastes."


The Second Finding: They Agree on Why They Fail

Here is where it gets interesting. The researchers didn't just look at what the AIs did recommend; they looked at what they missed.

They classified "missing" brands into three types of failures:

  1. The "Never Heard Of It" (Discoverability): The AI never found the brand in its search results. It's like a librarian who doesn't even know the book exists.
  2. The "Saw It, Didn't Like It" (Compellingness): The AI found the brand but didn't mention it in the final answer.
  3. The "Mentioned but Not Picked" (Positioning): The AI mentioned the brand but didn't recommend it as a top choice.

The Analogy: Imagine the two critics are asked to recommend a restaurant.

  • If they recommend different places, they are arguing about taste.
  • But if they both fail to recommend a specific local spot, why did they fail?

The Result: When both AIs failed to recommend a brand, they almost always failed for the exact same reason (95% of the time).

  • If they both missed a brand, it was usually because neither of them could find it (Discoverability).
  • They didn't disagree on why they missed it; they just both missed it for the same reason.

The "Long Tail" Effect:
This agreement gets even stronger for smaller, less famous brands (the "long tail").

  • Famous Brands (Top Tier): The AIs disagreed more often on why they missed them (maybe one saw them but didn't like them, the other never saw them).
  • Obscure Brands (Long Tail): The AIs agreed 99.6% of the time that they simply couldn't find the brand.

The "How" They Think: Different Roads to the Same Place

The paper also looked at how the AIs made their choices.

  • OpenAI (ChatGPT): Is like a researcher. It almost always goes out and searches the web first, reads the results, and then makes a recommendation. It relies heavily on what it finds in the search.
  • Anthropic (Claude): Is like a knowledgeable professor. It often relies on what it already knows from its training data (its "memory"). It recommends brands without even needing to search the web first about 43–52% of the time.

The Analogy:

  • ChatGPT is like a student who opens a textbook before answering a question.
  • Claude is like a student who answers from memory, only opening the textbook if they aren't sure.

Even though they use different methods (search vs. memory), they end up missing the same obscure brands because those brands aren't in the textbooks and aren't on the web pages they search.


The Practical Advice for Brands

Based on these findings, the paper suggests a specific strategy for brands:

  1. For Small, Obscure Brands (The Long Tail):

    • The Problem: The AIs can't find you.
    • The Solution: You only need one strategy. If you make your brand easier to find on the web (better SEO, more articles), both AIs will start recommending you. You don't need to tailor your message for ChatGPT vs. Claude; just fix the "discoverability" issue, and it fixes both.
  2. For Big, Famous Brands (The Top Tier):

    • The Problem: The AIs can find you, but they might not recommend you for different reasons (maybe one thinks your content is boring, the other thinks your positioning is weak).
    • The Solution: You need two different strategies. What works to get ChatGPT to recommend you might not work for Claude. You have to treat them as different audiences.

Summary

  • Recommendations: The AIs pick different winners (they disagree).
  • Failures: When they both miss a brand, they usually miss it for the same reason (they agree).
  • The Fix: If you are a small brand, fix your "visibility" once, and it helps everywhere. If you are a big brand, you might need to tweak your message for each specific AI.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →