← Latest papers
💻 computer science

DLT-Corpus: A Large-Scale Text Collection for the Distributed Ledger Technology Domain

This paper introduces DLT-Corpus, a large-scale, domain-specific text collection of 2.98 billion tokens from scientific literature, patents, and social media, which is used to demonstrate that DLT research precedes market growth and to release enhanced NLP tools like LedgerBERT.

Original authors: Walter Hernandez Cruz, Peter Devine, Nikhil Vadgama, Paolo Tasca, Jiahua Xu

Published 2026-05-29
📖 5 min read🧠 Deep dive

Original authors: Walter Hernandez Cruz, Peter Devine, Nikhil Vadgama, Paolo Tasca, Jiahua Xu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the world of Distributed Ledger Technology (DLT)—the family of technologies that includes blockchain and cryptocurrencies—as a massive, bustling city. For years, researchers trying to understand this city had to navigate using only a few scattered street signs (mostly about price charts and trading apps). They were missing the blueprints, the city council meeting minutes, and the chatter in the local cafes.

This paper introduces DLT-Corpus, a massive new library that finally collects the "whole story" of this digital city. Here is what the authors built and what they discovered, explained simply:

1. The Big Library (The Corpus)

Think of the DLT-Corpus as a giant warehouse containing 2.98 billion words (tokens) from 22 million documents. The authors didn't just grab random text; they carefully organized it into three distinct sections, like three different wings of a library:

  • The Academic Wing (Scientific Literature): 37,440 research papers. These are the "blueprints" where scientists and engineers first sketch out new ideas.
  • The Patent Wing: 49,023 patent filings. These are the "legal deeds" where companies officially claim ownership of new inventions.
  • The Town Square (Social Media): 22 million posts from Twitter/X. These are the "chatter" where regular people talk about what they are buying, selling, and hoping for.

Why is this special?
Before this, most computer programs studying this field only looked at the "Town Square" (social media) or specific trading data. They missed the deep technical stuff. This new library is 8.7 times richer in specific technical keywords than general internet text. It's like comparing a dictionary of everyday slang to a specialized encyclopedia of engineering; the latter is much better for teaching a computer to understand the specific language of this industry.

2. The "Smart Student" (LedgerBERT)

To prove this library is useful, the authors taught a computer model (a type of AI called LedgerBERT) to read it.

  • The Test: They asked the AI to find specific technical terms (like "Proof of Stake" or "Merkle Tree") in the text, a task called Named Entity Recognition.
  • The Result: The AI trained on this new library got 23% better at finding these terms than standard AI models. It learned the "dialect" of the DLT world much faster because it had so much high-quality material to study.

3. What the Library Revealed (The Analysis)

The authors used this library to watch how ideas travel through the city. They found two fascinating patterns:

Pattern A: The "Idea Journey"
They tracked three big concepts: Stablecoins (digital money pegged to real currency), DEXs (decentralized exchanges), and AMMs (automated trading tools).

  • The Finding: These ideas almost always appear in the Academic Wing (research papers) first.
  • The Journey: Years later, they show up in the Patent Wing (companies trying to protect them). Finally, much later, they explode in the Town Square (social media) where everyone starts talking about them.
  • The Metaphor: It's like a scientific discovery in a lab, which eventually becomes a product in a factory, and finally becomes a trend on TikTok. The research leads the market, not the other way around.

Pattern B: The "Emotional Weather"
They looked at how people feel about the market (sentiment) versus how much work researchers are doing.

  • The Finding: Even when the market crashes (a "crypto winter" where prices drop), the people on social media remain overly optimistic (bullish). They keep talking up the technology regardless of the price.
  • The Contrast: However, the scientists and patent holders are more grounded. Their work doesn't spike and crash with the daily news. Instead, their activity grows steadily over time, tracking the long-term growth of the industry, not the short-term mood swings.
  • The Cycle: The authors describe a "virtuous cycle": Research creates new tools \rightarrow These tools help the market grow \rightarrow A growing market provides money to fund more research.

4. The Rules of the Game (Ethics & Limits)

The authors were very careful about how they built this library:

  • No Copyright Violations: They only used open-access papers, public government patents, and social media data collected before Twitter changed its rules in 2023.
  • Privacy: They removed all usernames from the social media posts to protect people's privacy, keeping only the text and the date.
  • Language: The library is currently only in English, as that is the dominant language for science and patents.

Summary

In short, the authors built the first massive, comprehensive library for the Distributed Ledger Technology world. They proved that if you want to understand where this technology is going, you shouldn't just look at the stock prices or Twitter trends. You need to look at the research papers, because that is where the future is being written years before the rest of the world catches up. They also released the library, the trained AI model, and the tools for free so others can continue this research.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →