Is Misinformation More Open? A Study of robots.txt Gatekeeping on the Web
This study reveals a growing asymmetry in web gatekeeping where reputable news sites are increasingly blocking AI crawlers via robots.txt and active measures, while misinformation sites remain largely open, potentially skewing the training data available to Large Language Models.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the internet as a giant, bustling library where every book, article, and blog post is a piece of information waiting to be read. For years, "robots" (special computer programs) have been the librarians of this digital world, quietly walking the aisles to copy books so search engines can help us find them. But recently, a new kind of librarian has arrived: the Artificial Intelligence (AI) robot. These aren't just looking for one book; they are trying to read everything at once to learn how to write, answer questions, and chat with us. To manage this, website owners have a simple, voluntary sign they can put on their door called robots.txt. Think of it like a "Do Not Enter" or "Please Come In" sign for these digital librarians. If a website owner writes "No AI allowed" on this sign, polite AI robots are supposed to listen and stay away. But here's the big question: Do all websites put up these signs? And if they do, do they put them up for the same reasons? As AI gets smarter and more powerful, understanding who is opening their doors and who is locking them up becomes crucial, because it decides what kind of information the AI learns from.
This paper, titled "Is Misinformation More Open?", acts like a digital detective story, investigating how two very different groups of websites handle these "Do Not Enter" signs for AI. The researchers looked at two groups: "reputable" news sites (the serious, fact-checking journalists) and "misinformation" sites (the places spreading false or misleading stories). They wanted to see if the serious news organizations were better at telling AI robots to stay away compared to the fake news sites.
The findings reveal a surprising and stark contrast. The researchers found that reputable news websites are much more likely to put up a "Do Not Enter" sign for AI robots. In fact, 60.0% of these trustworthy sites explicitly told at least one AI crawler to stay out of their robots.txt files. In sharp contrast, only 9.1% of misinformation sites did the same. It's as if the serious libraries are locking their doors to the new AI students, while the fake news pamphleteers are leaving their doors wide open. On average, a reputable news site listed 15.5 different AI robots it didn't want to visit, while misinformation sites listed fewer than one.
The study also looked at whether these websites actually enforced their signs or just wrote them down. They found that reputable news sites not only wrote the rules but also actively blocked AI robots that tried to sneak in, matching their written rules with their actions. Misinformation sites were less consistent; while some did block AI, many didn't bother writing the rules down at all. The researchers also watched how things changed over time, looking at snapshots from September 2023 to May 2025. They saw that reputable news sites were quickly learning to lock their doors, with the number of sites blocking AI jumping from 23% to nearly 60% in just a couple of years. Misinformation sites, however, remained largely passive, rarely updating their signs.
The authors suggest that this growing gap creates a strange situation: as the "good" sources start locking their doors to protect their content, the AI robots might end up spending more time reading the "bad" sources that are still wide open. This could mean that the AI learns more from misinformation than from accurate news, simply because the misinformation sites are easier to access. The paper doesn't claim this has already ruined AI, but it highlights a real risk: if we don't pay attention to who is opening their doors, the future of AI might be built on a foundation of falsehoods.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.