ShareChat: A Dataset of Chatbot Conversations in the Wild
This paper introduces ShareChat, a large-scale, multi-platform dataset of 142,808 real-world chatbot conversations that preserves native interface affordances across 95 languages to enable more authentic evaluation of Large Language Models beyond uniform, text-only benchmarks.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you've been studying how people talk to robots, but you've only ever watched them through a single, plain glass window. You see the words they type, but you miss the unique features of the room they are in: the specific tools on the wall, the way the robot thinks out loud, or the different ways different robots are designed to help.
That's the problem the authors of this paper, SHARECHAT, are trying to fix. They built a massive new library of real conversations between humans and five different AI chatbots (ChatGPT, Perplexity, Grok, Gemini, and Claude) to show that where you talk to an AI changes how you talk to it.
Here is a simple breakdown of what they did and what they found:
1. The "Glass Window" Problem
Most previous studies looked at AI conversations through a "one-size-fits-all" lens. They stripped away all the unique features of each app to make the data look the same.
- The Analogy: Imagine studying how people order food by only looking at a list of ingredients, ignoring whether they are at a fast-food drive-thru, a fancy sit-down restaurant, or a food truck. You miss the context!
- The Fix: SHARECHAT is like a high-definition video recording of people ordering at all five of those places. It keeps the "drive-thru speaker" (Grok's social media links), the "chef's special notes" (Claude's code tools), and the "menu citations" (Perplexity's source links).
2. What's in the Library?
The researchers collected 142,808 conversations (over 660,000 back-and-forth messages) from people who voluntarily shared their chat links online.
- The Size: It's huge. While other libraries had short, choppy chats (like a quick text message), SHARECHAT has long, deep conversations where users ask follow-up questions, change their minds, and solve complex problems over many turns.
- The Variety: It covers 95 different languages, not just English.
- The Privacy: They scrubbed all names, emails, and phone numbers from the chats so no one can be identified, like blurring faces in a crowd photo.
3. Three Big Discoveries (The Case Studies)
The authors used this library to run three specific tests to see how different the robots really are:
A. The "Did You Get It?" Test (Conversation Completeness)
They checked if the AI actually finished what the user asked for.
- The Finding: Some robots are like reliable librarians who always find the whole book (ChatGPT and Claude). Others are more like search engines that give you a great first page but might stop before you find the answer (Perplexity and Grok). In the real world, many conversations end with the user only partially satisfied, a nuance previous studies missed because they only looked at single questions.
B. The "Where Did You Look?" Test (Source Grounding)
They looked at how two search-focused robots (Grok and Perplexity) found their answers.
- The Finding: They are like two different detectives.
- Grok is obsessed with social media. It mostly cites posts from X (Twitter) and news sites, acting like a real-time news ticker.
- Perplexity is a researcher. It cites Wikipedia, scientific journals, and forums like Reddit, acting like a librarian digging through archives.
- The Lesson: They aren't just "AI"; they are built with different tools for different jobs.
C. The "Speed of Thought" Test (Timing)
They measured how long it took for users to type a reply after the AI spoke.
- The Finding: The rhythm of the conversation changes depending on the robot.
- With ChatGPT, the robot gets faster as the chat goes on (maybe it's learning your style or caching answers).
- With Grok, the robot gets slower as the chat gets longer (maybe it's getting overwhelmed by all the history).
- This shows that the "personality" of the chat isn't just about the words; it's about the timing and flow.
4. Why This Matters
The paper argues that if we only study AI through a generic, stripped-down lens, we are missing the most important part: the user's reality.
- Real Life is Messy: People don't just ask one question and leave. They have long, winding conversations.
- Tools Matter: Users treat different AIs as different tools. They use one for coding, another for news, and another for creative writing.
- Better Training: By giving researchers this "real-world" data with all its unique features intact, we can build better AI that understands not just what to say, but how to say it in the specific context of the app it lives in.
In short: SHARECHAT is a giant, diverse collection of real-world AI chats that proves we can't understand these robots by looking at them in a vacuum. We have to see them in their natural habitats to understand how they really work.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.