Convergence in Science, Divergence in Religion: Calibrated Framing Differences Across Wikipedia's Language Editions
This paper introduces a calibrated framing distance metric to demonstrate that while scientific concepts are described with high consistency across Wikipedia's language editions, religious concepts exhibit the greatest divergence, even after accounting for language-family alignment differences and encoder biases.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a library where every book about the same subject is written by a different group of people, in a different language, without ever consulting one another. This is Wikipedia. With editions in over 300 languages, it stands as one of humanity's largest collaborative projects, yet its chapters are not translations of a single master text. Instead, a Turkish editor, a Korean editor, and a Polish editor might all write about "pilgrimage" or "censorship," drawing from their own local sources, cultural histories, and community values. This creates a unique opportunity for researchers: if we look at how these different communities describe the exact same idea, can we see where their perspectives align and where they drift apart? This question sits at the intersection of language, culture, and technology. It matters because the text we write today often becomes the training data for the artificial intelligence systems of tomorrow. If those systems learn from a library where the definition of "democracy" or "sacrifice" shifts depending on the language, the machines they build might end up telling different versions of history to different people.
A researcher set out to measure this drift, not by counting how many facts are missing from one language versus another, but by measuring how differently the same facts are framed. They gathered 2,799 articles from twenty different languages, covering a wide range of topics. To ensure they were comparing apples to apples, they anchored their study to 150 specific concepts, such as "karma," "democracy," "quantum mechanics," and "refugee," using a global database to guarantee that every language edition was discussing the same core idea. They also included a set of neutral concepts, like chemical elements and basic colors, which serve as a control group. The goal was to see if the way people write about religion or politics differs more from one language to another than the way they write about settled science.
The researcher faced a significant hurdle: the computer tools used to compare text are not perfect. These tools, known as sentence encoders, translate words into mathematical points in a vast space to measure similarity. However, these tools sometimes work better for some pairs of languages than others. For instance, the tool might naturally think French and Spanish are closer than Hindi and Chinese, simply because of how it was trained, not because of any real cultural difference. If the researcher had just measured the raw distance between the articles, they might have mistaken these technical quirks for cultural gaps. To fix this, they created a "calibrated distance." They first measured how far apart the neutral, uncontroversial articles were for each pair of languages. This gave them a baseline, a kind of zero point that accounted for the tool's natural bias. Then, they subtracted this baseline from the distance of the actual topics they were studying. This adjustment stripped away the technical noise, leaving only the genuine differences in how the communities described their subjects.
The results revealed a clear and striking pattern. When it came to religion, the different language editions were significantly more divergent than the neutral baseline. Concepts like "sacrifice," "clergy," and "martyr" showed the widest gaps in description. Similarly, politically charged terms like "censorship" and "refugee" were described very differently across languages. In contrast, the articles on science and technology were the most aligned. Topics such as "evolution," "DNA," and "quantum mechanics" were described with remarkable consistency across all twenty languages, often even more closely than the neutral control concepts. This suggests that while scientific knowledge tends to converge into a shared global understanding, cultural and religious narratives remain deeply rooted in local contexts.
The study also found that this divergence is not spread evenly across all political topics. While "censorship" caused a rift, concepts like "democracy" and "human rights" were described quite similarly across languages. This indicates that the split is not about politics as a whole, but about specific, sensitive ideas that touch on local power dynamics and identity. The researcher tested their findings using three different computer models to ensure the results were not just an artifact of one specific tool. While the exact size of the differences changed slightly depending on the model, the ranking remained the same: religion was the most divergent domain, and science was the most aligned.
One of the most important takeaways is that language families do not tell the whole story. One might expect that languages from the same family, like French and Spanish, would frame concepts more similarly than distant languages like Hindi and Chinese. While this was true for the raw, uncalibrated data, the effect largely disappeared once the researcher adjusted for the technical biases of the computer tools. This means that the differences in framing are not simply a matter of linguistic heritage; they are genuine cultural choices made by the editors. The study also highlighted that the Wikipedia community for Chinese is unique. Because the mainland version of the site is blocked in China, the Chinese edition analyzed in this study reflects the perspectives of editors in Taiwan, Hong Kong, and overseas communities, rather than a single, unified national viewpoint.
Ultimately, this research provides a new way to look at the global knowledge base. It shows that while humanity agrees on the facts of the physical world, we still tell very different stories about our beliefs and our struggles. The data suggests that if artificial intelligence is trained on this material, it will likely inherit these divergences, potentially offering different historical or cultural narratives depending on the language a user speaks. By calibrating their measurements to remove technical noise, the researcher has provided a clearer map of where these differences lie, separating the artifacts of the machine from the genuine variations of human culture. The code and data from this study are now available for others to use, ensuring that future investigations into how we frame our world can be built on a foundation of clarity rather than confusion.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.