← Latest papers
💬 NLP

Vidya: An AI-Driven Modular Pipeline for Archival Automation and Semantic Metadata Enrichment

Vidya is a modular, open-source pipeline developed at UEPG that leverages Large Language Models constrained by YAML ontologies and Pydantic validation to automate the semantic enrichment and archival ingestion of historical "dark data," significantly reducing processing time from decades to days while ensuring compliance with international standards like NOBRADE and ISAD(G).

Original authors: Cloter Migliorini Filho, Julia Graciela Machado, Edson Armando Silva, Marcella Scoczynski

Published 2026-05-19
📖 4 min read☕ Coffee break read

Original authors: Cloter Migliorini Filho, Julia Graciela Machado, Edson Armando Silva, Marcella Scoczynski

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a massive library of old, dusty boxes filled with historical documents, photos, and letters. You've spent years carefully scanning them all to save them from rotting away. But here's the problem: you now have a mountain of digital files that are essentially "dark data." They exist, but no one can find them because they lack labels, tags, or descriptions. It's like having a giant warehouse full of mystery boxes where you don't know what's inside without opening every single one by hand.

Traditionally, hiring people to read every document and write a summary for it is incredibly slow, expensive, and prone to mistakes. This is where Vidya comes in.

What is Vidya?

Think of Vidya as a super-smart, tireless robot librarian built specifically to solve this labeling problem. It's a "pipeline," which is just a fancy word for an assembly line. Instead of a human sitting at a desk, Vidya takes your digital files, reads them using Artificial Intelligence (AI), and automatically writes the necessary descriptions (metadata) so people can search for and find them later.

How Does It Work? (The "Digital Straitjacket")

You might be thinking, "Wait, AI can be unreliable. It sometimes makes things up or 'hallucinates' facts." The authors of the paper knew this was a huge risk for archives, where accuracy is everything.

To fix this, Vidya uses a clever trick they call a "digital straitjacket."

  • The Problem: Imagine asking a creative writer to describe a photo. They might invent details that aren't there.
  • The Solution: Vidya doesn't let the AI just "chat." Instead, it forces the AI to fill out a strict, pre-made form (defined by a file called YAML). It's like giving the AI a fill-in-the-blank worksheet with specific rules.
  • The Check: Before the AI's answer is accepted, a digital gatekeeper (called Pydantic) checks the work. If the AI tries to write something that doesn't fit the form or breaks the rules, the system rejects it. This ensures the output is always structured, accurate, and follows international library rules (like ISAD(G) and NOBRADE).

The "Maker" Spirit

Vidya wasn't built in a high-tech corporate lab with unlimited money. It was built by a team at a university in Brazil using Maker principles.

  • Low Cost: It's designed to run on modest, affordable computers (like a standard desktop or even a Raspberry Pi), not expensive supercomputers.
  • Open Source: It's free software that anyone can look at, fix, or improve.
  • Collaboration: It brings together historians (who know the history) and computer scientists (who know the code) to work together.

What Did They Find?

The team tested Vidya on a pile of 50 documents (newspapers and photos) and compared it to doing the work by hand. The results were dramatic:

  1. Speed: What would take a human team 18.5 years to process manually, Vidya could do in about 45 to 60 days.
  2. Cost: The AI-assisted approach cost only about 1.5% of what human labor would cost.
  3. Accuracy: Vidya got about 85% accuracy on key details like titles, creators, and dates. While not perfect, it's good enough to do the heavy lifting, leaving humans to just double-check the work rather than starting from scratch.
  4. Accessibility: By turning these "dark" images into text descriptions, Vidya makes these archives accessible to people who are blind or have low vision, who use screen readers to navigate the web.

The Big Picture

Vidya isn't just a tool; it's a way to unlock history. It takes the "dark data" of the past—files that were sitting in the dark, unseen and unsearchable—and turns them into a bright, organized, and searchable digital museum. It proves that you don't need a massive budget to use advanced AI; you just need a smart, structured approach to keep the technology honest and helpful.

In short: Vidya is the bridge that turns a chaotic pile of digital files into a usable, searchable library, saving decades of work and making history accessible to everyone.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →