Agreement is not corroboration: bounding the independence of cultural-heritage authority links
This study demonstrates that agreement between Getty ULAN and Library of Congress authority records in Wikidata does not guarantee independent corroboration, as a significant portion of links are propagated or untraceable, resulting in an independence rate that remains statistically indeterminate and highlights the critical need to explicitly model provenance dependence in linked-data workflows.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the vast, interconnected world of digital libraries, a quiet revolution is taking place to make information easier to find and connect. Imagine a global network where every museum, archive, and library speaks the same language, allowing a researcher in one country to instantly find a painting in another or a historical figure mentioned in a third. To make this happen, these institutions rely on "authority files"—carefully curated lists that ensure everyone uses the same name for the same person or place. When different systems agree on a name, they link their records together, creating a web of knowledge that is far more powerful than any single collection could be alone.
However, a subtle but critical question arises when these links are formed: does agreement mean truth? If two different databases list the same identifier for a famous artist, it is tempting to assume they have both independently confirmed the artist's identity, doubling the reliability of the information. But in the digital realm, one system might have simply copied the identifier from the other, or both might have inherited it from a third, common source. In this scenario, the agreement is real, but the evidence is not independent. Distinguishing between a genuine, double-checked confirmation and a simple chain of copying is essential for knowing how much trust to place in the data.
A recent study by Emin Talip Demirkiran of Eskisehir Technical University tackles this exact problem within Wikidata, a massive, collaborative knowledge graph that acts as a central hub for connecting these various library systems. The research focuses on the links between Getty's Union List of Artist Names (ULAN) and the Library of Congress (LC) name-authority files, two of the most important standards in the cultural heritage world. The goal was not to see if the links existed, but to determine how many of those links were truly independent discoveries versus how many were merely copies of one another.
To answer this, the researcher examined a frozen snapshot of 3,000 specific entities from the Wikidata database. For each entity, the study looked at the "provenance," or the history of how the link was created. The team classified every instance of agreement into one of three categories. Some links were clearly "propagated," meaning the data had been imported or copied from one source to another. Others were "independent," showing clear signs that the two sources had established the link on their own. A third group, however, was "untraceable," where the digital history was missing or unclear, leaving the origin of the agreement a mystery.
The results revealed a landscape far more complex than a simple checkmark of agreement. Out of the 1,546 specific statements analyzed, only about 256 could be confirmed as independently established. A much larger portion, 728 statements, were clearly propagated, meaning they were copies. The most significant finding, however, concerned the 562 statements that fell into the untraceable category. Because the digital trail was broken for these items, the researchers could not say for certain whether they were independent or copied.
This missing information created a range of possibilities rather than a single answer. If every untraceable statement were actually a copy, the overall rate of independent agreement would be very low, around 16 percent. If every untraceable statement were actually independent, the rate would rise to about 56 percent. The true answer lies somewhere in between, but the data simply cannot pinpoint it. Crucially, this entire range of uncertainty crosses the halfway mark, meaning the evidence does not support a definitive claim that most links are independent, nor does it prove they are mostly copies.
When the researchers looked only at the links they could trace, the picture was stark: approximately 71 percent of the traceable agreements were propagated. This suggests that when we see two systems agreeing, it is often because one is following the lead of the other, not because they have both done their own homework. Furthermore, the study found that this phenomenon was almost entirely concentrated in records about people, specifically artists and historical figures linked through the Getty system. The mechanisms for linking other types of cultural heritage, such as places or objects, were far less active in this specific context.
The study concludes that while agreement between authority files is valuable for connecting data and making it usable, it should not be mistaken for independent verification. In the world of digital libraries, a copied identifier is still useful for navigation and discovery, but it does not provide the same evidential weight as a double-checked fact. The research highlights that missing history is not just a gap in documentation; it fundamentally changes what we can know. Without a clear record of where a link came from, we cannot assume it is a fresh confirmation.
This work serves as a reminder that in our increasingly linked digital world, the path a piece of information takes matters as much as the information itself. Just as a historian would not accept a fact simply because two books say it, without checking if one book copied the other, digital systems must account for the source of their connections. The study does not suggest that these links are useless, but it argues that we must be careful not to count the same piece of evidence twice. By explicitly modeling where data comes from and acknowledging where the trail goes cold, digital libraries can build a more honest and reliable foundation for the future of knowledge.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.