← Latest papers
⚡ electrical engineering

Measuring Latent Documentary Quality in Scholarly Communication: An Indicator Framework Based on PDF Validation and Technical Debt

This paper proposes a technical-debt-based indicator framework that transforms PDF validation failures into a nine-family typology to measure the latent documentary quality of scholarly communication infrastructures, revealing widespread noncompliance and disciplinary variations in archival stability and machine operability.

Original authors: Kemal Yayla

Published 2026-09-08
📖 7 min read🧠 Deep dive

Original authors: Kemal Yayla

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a library where every book is open to the public, listed in a giant catalog, and cited by scholars around the world. On paper, this seems perfect. But what if the pages of those books were made of a material that crumbles after a few years, or if the text was written in a code that no machine could read? In the world of academic research, we have spent decades building systems to measure how open, visible, and influential a study is. We count how many times a paper is cited or check if it is available for free. These are useful measures, but they only tell us if a paper can be found. They do not tell us if the paper itself is built to last, or if it is structured in a way that allows computers, screen readers for the blind, and future archives to actually understand and use the information inside.

This gap is the focus of a new study by researcher Kemal Yayla. The study looks at the "documentary quality" of scholarly articles. This is a fancy way of asking: is the digital file that carries the research stable enough to survive for decades, and is it organized well enough for machines to process? The researcher treats the digital file not just as a container for words, but as a fragile object that needs specific technical features to remain useful. To find out if these features are present, the study examined 1,340 final versions of research papers published by library-run journals in the United States and Canada. These papers were checked against strict technical standards designed to ensure they can be preserved forever and read by assistive technologies. The results reveal a startling reality: while these papers are easy to find, the files themselves are often broken in ways that threaten their long-term survival and accessibility.

The study began by gathering a specific set of documents. The researcher selected journals that are part of the Library Publishing Coalition, a group of libraries in the US and Canada that run their own academic journals. These journals are known for being open and transparent. The researcher then pulled the final published PDF files for articles published between 2024 and 2025. In total, 1,340 files were collected. These files were then run through automated testing tools that check for compliance with three different technical standards. The first standard, known as PDF/A-1b, is a basic rulebook for making sure a document can be opened and read by any computer in the future. The second, PDF/A-2u, is a stricter version that requires the file to carry its own digital ID card so archives can identify it automatically. The third standard, PDF/UA-1, is a rulebook for accessibility, ensuring the document has a logical structure that screen readers can navigate and that text can be copied and searched by machines.

When the testing began, the results were stark. Out of the 1,340 documents, only 18 passed the basic test for long-term preservation. None of them passed the stricter test for digital identification. Only 10 passed the test for accessibility. This means that for almost every single paper checked, the file failed to meet the minimum requirements for being a reliable, permanent record. However, the study argues that simply saying "these files failed" misses the most important part of the story. The researcher developed a new way of looking at these failures, grouping them into nine different categories based on what exactly was broken and whether it could be fixed.

The analysis showed that the failures were not all the same. Some problems were like missing labels on a box. For example, nearly every single file failed because it was missing a specific digital declaration that tells an archive what kind of file it is. This is a small, patchable error. It does not mean the content of the paper is wrong; it just means a piece of metadata was left out during the final export. This type of error is classified as "post-hoc patchable," meaning it can be fixed later without changing the original work. In contrast, other failures were much more serious. The study found that 98.4% of the files failed the accessibility test because they lacked a logical structure. This means the text was not tagged in a way that tells a computer where the title ends and the body begins, or where a table starts and stops. These are "upstream-only" problems. They happen when the document is first created. Once the file is saved as a PDF, this structural information cannot be reliably added or fixed. The damage is done before the file even reaches the public.

The study also looked at whether these problems varied by subject or by country. The results showed that the lack of structural tagging was a universal problem. It appeared in almost every document, regardless of whether the paper was about biology, history, or engineering. This suggests the issue is not with the author or the specific fields of study, but with the software and systems used to create the files. However, some other types of errors did vary. For instance, problems with text that could be copied or searched were much more common in science and math papers than in art or law papers. This makes sense, as science papers often use complex symbols and special characters that are harder for computers to handle correctly. The study also compared journals from the United States and Canada. While both groups had very low pass rates overall, the specific types of errors differed slightly, with Canadian journals performing slightly better on basic preservation checks but worse on accessibility checks.

The most important takeaway from this work is that we cannot judge the quality of a research paper just by looking at its title, its abstract, or how many times it has been cited. A paper can be widely read and highly praised, yet its digital file might be unreadable by a blind person, unsearchable by a computer, or impossible to open in fifty years. The researcher's new framework proves that we can measure these hidden flaws. By breaking down technical errors into specific categories, we can see exactly what is broken and where the fix needs to happen. Some fixes are easy and can be done after the paper is published. Others require a fundamental change in how the papers are created in the first place.

This study does not claim that library publishing is failing. In fact, it uses library publishing as a test case because these institutions are supposed to be the most careful about preserving and sharing knowledge. If the files are this broken in a system designed to be perfect, the problem is likely even worse in other publishing environments. The research does not offer a magic solution, but it provides a clear map of the problem. It shows that the tools we use to evaluate science are blind to the physical and technical health of the documents themselves. Until we start measuring whether the files are stable and accessible, we are only measuring half of the scholarly record. The study concludes that we need to treat the digital file as a critical piece of evidence that requires its own quality checks, ensuring that the knowledge we produce today can actually be used by the world tomorrow.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →