Synthetic Sources?: Auditing Generative Search Engine Citations for Evidence of AI-Generated Sources
This paper audits four major generative search engines using 712 real-world queries and finds that approximately 16% of their cited sources are AI-generated, highlighting a significant risk of users receiving unverified synthetic information presented as authoritative.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you ask a group of four very smart, well-traveled librarians (ChatGPT, Copilot, Gemini, and Perplexity) to answer a question about politics, health, or the environment. Instead of just giving you a direct answer, these librarians say, "Here is the answer, and here are the 10 books we read to write it."
This paper is an audit of those "books" (the websites they cite) to see if any of them were actually written by a robot pretending to be a human.
Here is the breakdown of what the researchers found, using simple analogies:
1. The Setup: The "Library" of the Internet
The researchers treated these AI search engines like a new kind of library. They asked 712 real questions (about things like "Is this diet safe?" or "What is the latest political scandal?") to four different AI engines.
The engines responded with answers and provided a list of links (citations) to prove where they got their info. The researchers then went to those links, read the content, and ran it through a "lie detector" test (an AI-detection tool called Pangram) to see if the text was written by a human or generated by another AI.
2. The Big Discovery: The "Robot Ghost" in the Library
The main finding is that about 16% of the "books" these librarians cited were actually written by robots.
- The Metaphor: Imagine you go to a library to find a history book. You pick one off the shelf, open it, and realize the author is a machine that wrote the whole thing in a second. You didn't know that when you picked it up.
- The Reality: In this study, roughly 1 out of every 6 sources the AI engines cited was AI-generated content.
- The Worst Offender: The AI engine Copilot was the most likely to cite these robot-written sources (nearly 3 out of 10 of its citations were AI-generated). ChatGPT was the most careful, but still cited them about 7% of the time.
3. The "Long Tail" Problem: The Crowd vs. The VIPs
The researchers noticed a strange pattern in which websites the AI engines liked to cite.
- The VIPs (The Head): A small group of famous, trusted websites (like Wikipedia, government sites like
whitehouse.gov, and major news outlets) get cited over and over again. They are the "VIPs" of the library. - The Long Tail: However, the vast majority of the citations come from a "long tail" of thousands of obscure, rarely visited websites. Most of these obscure sites are cited only once or twice.
- The Finding: The "robot-written" content was mostly hiding in this Long Tail. The famous VIP sites were mostly human-written, but the obscure, one-hit-wonder websites were the ones filled with AI-generated text.
- The Analogy: It's like a party where the host (the AI) keeps introducing you to the same 5 famous people (Wikipedia, Gov sites), but then spends 80% of the time introducing you to 1,000 strangers in the back of the room, many of whom are actually mannequins (AI content) pretending to be people.
4. The Topic Breakdown
The researchers checked three specific areas:
- Environment: This topic had the highest amount of AI-generated sources cited.
- Health: A significant portion of health advice sources were also AI-generated.
- Politics: Also contained AI sources, though slightly less than the other two.
5. Why This Matters (According to the Paper)
The paper argues that this is risky because:
- The "Trust" Trap: When an AI gives you an answer and shows you a list of links, you assume those links are real, human-verified facts. If the links are actually AI-generated, the AI is essentially citing its own "hallucinations" or low-quality robot text as proof of its own answer.
- The Quality Gap: While not all AI text is bad, the paper notes that AI text can contain errors, biases, or made-up facts. If an AI search engine uses AI-written sources to build its answer, it creates a "house of cards" where the foundation is shaky.
- The "Black Box": The engines don't tell you, "Hey, this source was written by a bot." They just show you the link.
Summary
The paper concludes that these new "Generative Search Engines" are currently acting like librarians who sometimes accidentally cite books written by robots, without telling the reader. They rely heavily on a few famous websites, but they also pull a massive amount of information from a "long tail" of obscure websites, many of which are filled with AI-generated content.
The researchers say this is a warning sign: if we don't fix how these engines choose and check their sources, users might unknowingly trust information that is entirely synthetic.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.