← Latest papers
📊 statistics

MajinBook: An open catalogue of digitally mediated world literature

This paper introduces MajinBook, an open catalogue that links metadata from shadow libraries with Goodreads data to create a high-precision, machine-readable corpus of over 539,000 digitally mediated English-language books, while addressing traditional corpus biases and discussing the project's legal permissibility for research under EU and US frameworks.

Original authors: Antoine Mazières, Thierry Poibeau

Published 2026-05-13
📖 5 min read🧠 Deep dive

Original authors: Antoine Mazières, Thierry Poibeau

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Picture: Building a Better Library Map

Imagine you want to study how people read and what stories matter to the world. To do this, you need a massive map of books.

For a long time, researchers used HathiTrust, which is like a giant, official library built by universities. It has millions of books, but it has two big problems:

  1. It's locked: Many books are under copyright, so you can't read the full text; you can only look at them through a secure, restricted window.
  2. It's messy: Many books were just scanned from physical paper. The computer often misreads the letters (like confusing an 'o' for a '0'), making the text hard to use for advanced computer analysis.

The authors of this paper wanted a better map. They turned to "Shadow Libraries" (like Library Genesis and Z-Library). Think of these as massive, underground digital bazaars where millions of books are shared freely. They have way more books than the official library, but they are chaotic. The "shelves" are disorganized, the labels are often wrong or missing, and it's hard to tell which book is which.

MajinBook is the project that cleaned up this chaos. It created a high-quality, organized catalogue of over 539,000 English books (plus millions of others in French, German, and Spanish) by connecting the messy "shadow library" files with the organized data from Goodreads (the social network for book lovers).


The Ingredients: How They Built It

1. The "Digital-Only" Rule (The Sieve)

The shadow libraries have books in many formats. Some are PDFs (digital scans of paper pages), and some are EPUBs (native digital files, like a clean webpage).

The authors decided to throw away all the PDFs.

  • The Analogy: Imagine you are trying to study the taste of a soup. The PDFs are like taking a photo of the soup and trying to taste the picture. It's blurry and inconsistent. The EPUBs are like the actual soup in a bowl; you can taste every ingredient clearly.
  • The Result: By only keeping the "clean" digital files (EPUBs), they lost some very old books (because they weren't digitized yet), but they gained a dataset that is much cleaner, easier for computers to read, and surprisingly diverse in languages.

2. The "Work vs. Edition" Problem

In the world of books, there is a difference between a Work (the story itself, like The Odyssey) and an Edition (a specific version, like the 1990 translation or the 2020 audiobook).

  • The Problem: Shadow libraries might have 50 different files for The Odyssey. If you count them all as 50 different books, your data is skewed.
  • The Solution: The authors used Goodreads as a "skeleton" or a "filing system." Goodreads already knows that all 50 versions belong to the same "Work." They used this knowledge to group the messy shadow library files together, linking them to the correct story and its original publication date.

3. The Matching Game (Finding the Needle in the Haystack)

How do you connect a messy file from a shadow library to the correct entry on Goodreads?

  • The Strategy: They didn't try to match the exact cover art or page count (which often changes). Instead, they looked at identifiers (like ISBN numbers) and titles.
  • The Fuzzy Logic: Sometimes a title is slightly misspelled in the shadow library (e.g., "Harry Pottter"). The authors used a computer trick called "fuzzy matching" to realize, "Hey, that's probably Harry Potter."
  • The Safety Check: They tested this method by having humans check 200 random matches. They found that if the computer gave a "confidence score" of 80 or higher, it was almost always correct. They set their final rule at that 80-point threshold to ensure their catalogue is highly accurate.

The Results: What Did They Find?

  • A Massive, Clean Collection: They created a list of 539,000 English books that are ready for computers to analyze.
  • A Time Bias (But a Useful One): Because they only kept "native digital" files, their collection has fewer books from the 1800s and more from the 2000s.
    • The Paper's Take: The authors admit this isn't a perfect history of all literature. Instead, it's a perfect history of 21st-century digital culture. It shows us what books are popular right now in the digital age.
  • More Languages: Surprisingly, by filtering for digital files, they actually got a more diverse mix of languages. The "PDF" pile was mostly English, but the "EPUB" pile included a healthy amount of French, German, Spanish, and others.

The Legal "Fine Print"

The paper addresses a tricky question: Is it legal to use these shadow libraries for research?

  • The Argument: The authors argue that they aren't selling the books or distributing the stories. They are only publishing a list of facts (titles, authors, dates).
  • The Analogy: It's like a librarian making a card catalog. The card catalog doesn't contain the story; it just tells you where the story is and who wrote it. The authors claim that creating this "card catalog" for research purposes falls under legal exceptions for "Text and Data Mining" in both the US and EU.

Summary

MajinBook is a tool that takes the chaotic, massive, and often messy world of "shadow libraries" and organizes it using the social wisdom of Goodreads. The result is a super-clean, computer-friendly list of books that helps researchers study culture, literature, and reading habits in the digital age, all while navigating the complex legal landscape of copyright.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →