Tracking Conversations: Measuring Content and Identity Exposure on AI Chatbots
This paper systematically measures web tracking on 20 popular AI chatbots and reveals that the majority share sensitive user content and identity information with third-party analytics, advertising, and session replay services, even in private chat modes.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine AI chatbots as high-tech, friendly librarians. You walk up to them, whisper a question, and they pull out a book to help you. In the past, you might have thought, "This is just a conversation between me and the librarian."
However, this research paper reveals that while you're talking to the librarian, a whole crowd of strangers is standing behind a one-way glass wall, taking notes, snapping photos, and even recording audio.
Here is the breakdown of what the researchers found, using simple analogies:
1. The Setup: The "Glass House" Librarians
The researchers went into 20 of the most popular AI chatbot websites (like ChatGPT, Gemini, Claude, etc.). They didn't just ask random questions; they asked a very specific, sensitive question: "Where can I get a pregnancy test near me?"
This is like walking into a library and asking for medical advice in a very hushed tone. The researchers wanted to see what happens to that specific piece of information once it leaves your mouth.
2. The Big Reveal: The "Third-Party Peeping Toms"
The study found that 17 out of the 20 chatbots are sharing your conversation details with outside companies (third parties).
Think of it this way: You are having a private conversation in a room, but the walls are made of glass, and the chatbot owner has invited 17 different neighbors to come over and listen in.
- The "Session Replay" Cameras: Three of the chatbots were using a tool called "Microsoft Clarity." Imagine this as a security camera that doesn't just record the room; it records exactly what you said and what the librarian said back, in plain text. The researchers found that these cameras were sending the full text of the "pregnancy test" conversation to Microsoft, allowing them to read your private medical query.
- The "Chat ID" Postcards: Even if the chatbot doesn't send the full text, 15 of them sent a "Postcard" to advertisers and analytics companies. This postcard didn't say what you asked, but it had a unique Chat ID or a link to your specific conversation. It's like sending a note that says, "I am in Room 404," to a marketing company. They might not know you asked about pregnancy tests yet, but they know you are in that specific room, and they can link that room to your identity later.
3. The "Name Tag" Problem
The researchers also found that many chatbots were handing over your Identity (your name, email, or account ID) along with the conversation.
- The "Support Widget" Leak: Some chatbots have little "Help" buttons on the screen. When you load the page, these buttons automatically shout out your name and email address to the company that built the button (like Intercom), even if you never clicked the button.
- The "Hashed" Secret: Some chatbots sent your email address, but "scrambled" it (hashed). Think of this like sending a secret code that only the recipient can decode. If you use the same scrambled code on a shopping site and a chatbot site, the advertisers can link the two together and realize, "Oh, the person who bought diapers is the same person asking about pregnancy tests."
4. The "Incognito" Mode: Does it Work?
Many chatbots offer a "Private" or "Temporary" chat mode, promising that nothing is saved. The researchers tested this.
- The Result: It worked! In private mode, the "crowd of neighbors" largely disappeared. The number of outside companies listening dropped significantly, and in the private chats they tested, no one saw the conversation text or your identity.
- The Catch: The chatbot's privacy policy didn't explicitly say, "Private mode stops the advertisers." It just talked about not saving your history. The researchers found that the design of the private mode stopped the tracking, even if the rules didn't clearly explain it.
5. The "Privacy Policy" vs. Reality
The researchers compared what the chatbots said they were doing (in their long, boring privacy policies) with what they were actually doing.
- The Mismatch: Some policies were vague, saying "we share data with partners." Others were specific but missed the biggest leakers. For example, three chatbots listed specific partners in their policies but forgot to mention Microsoft Clarity, even though Clarity was the one recording their private conversations in plain text.
- The Exception: One chatbot, Duck.ai, was the "honest librarian." They didn't share data with anyone and even stripped out your location before sending the question to the AI.
Summary
The paper concludes that while AI chatbots are great tools, they are currently built like glass houses.
- 17 out of 20 chatbots let outside companies peek in.
- 3 chatbots let those companies read your private medical or personal questions out loud.
- 15 chatbots send a unique ID to advertisers that links your conversation to your identity.
- Private mode is the only "curtain" that effectively blocks the neighbors, but the chatbots don't always tell you that's what it does.
The researchers suggest that chatbot owners need to build better walls (design changes) and be more honest about who is listening (clearer policies).
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.