From Wikipedia to AI: Measuring 25 years of synthesis of human genetics research in the public-facing information ecosystem
This study analyzes 25 years of Wikipedia revisions and emerging AI outputs to reveal that human genetics research is frequently synthesized into public information about ethnicity and nationality, with AI-generated encyclopedias exhibiting a higher frequency of genetic references and a greater tendency to hallucinate or misrepresent this research compared to traditional sources.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Imagine the human genome as the ultimate instruction manual for building a person, a massive library of code that scientists spent decades decoding. But here's the tricky part: this manual doesn't just talk about blue eyes or curly hair; it often gets tangled up with how we define groups of people, like nations, ethnicities, or races. Think of it like a recipe book that accidentally gets mixed in with a history book. Because these topics are so personal and political, it's very easy for the science to get twisted or misunderstood when regular people read about it. We all rely on public information—like the websites we visit or the chatbots we ask for homework help—to make sense of this complex science. But if the information out there is a bit jumbled, or if it invents facts that don't exist, it could lead us to believe things about our own biology that simply aren't true. That's why it matters to check what's actually written in these public spaces and see if the science is being told correctly.
This paper acts like a giant time-traveling detective story, looking back over 25 years to see how human genetics research has been mixed into the public information ecosystem. The researchers decided to take a massive snapshot of Wikipedia, the world's most popular online encyclopedia, which they treat like a giant, living mural that gets repainted millions of times. They didn't just look at a few pages; they analyzed a staggering 3,050,422 historical revisions from 6,738 different pages about ethnicity, nationality, and race.
What they found is that genetics terminology has quietly slipped into 14.8% of these pages, but if you zoom in on the top 1,000 most popular pages, that number jumps to 55.5%. It's as if more than half of the most famous stories about national groups now include a biological chapter. Specifically, 67.8% of the pages about nationalities mention genetics, suggesting that research is being synthesized to present a biological element to the idea of ethnicity and nationality. The authors also peeked behind the scenes at the "talk pages," where editors argue about what should be written, and found that 10.1% of the 56,908 discussions there contained genetics terminology.
The story doesn't stop at Wikipedia. The team also asked three popular chatbots questions about nationalities, and they discovered that these AI assistants often point back to both genetics and Wikipedia as their sources. But the most surprising twist comes from a new, AI-generated encyclopedia called Grokipedia. When the researchers looked at 133 pages from this new source, they found it mentioned genetics even more frequently than Wikipedia. However, unlike the human editors who try to be careful, this AI-generated encyclopedia was found to hallucinate or misrepresent human genetics research, essentially making up facts or getting the science wrong. The paper suggests that while we are synthesizing more science into our public stories, we need to be careful that the new AI storytellers aren't inventing a biological reality that doesn't exist.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.