← Latest papers
💻 computer science

On the missing data layer and a potential solution

This paper identifies the critical lack of a unified dataset layer in Latin American AI infrastructure, characterized by issues of discovery and insufficient supply, and proposes "DataHub," a task-first data infrastructure built on a specific ontology to enable dataset discovery, metadata management, contribution, licensing, and reuse.

Original authors: Francis F Daniel, Mauro Ibañez, Francis Perelman, Marian Basti

Published 2026-08-05
📖 6 min read🧠 Deep dive

Original authors: Francis F Daniel, Mauro Ibañez, Francis Perelman, Marian Basti

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Missing Ingredients for a Local Brain

Imagine you are trying to bake the perfect cake, but every recipe you find was written for a different kitchen, using ingredients you've never seen and ovens that run at different temperatures. This is the current situation for Artificial Intelligence (AI) in Latin America. To understand why this matters, we need to look at two simple ideas. First, AI models are like super-smart students; they learn by reading massive amounts of information (data) to figure out how to solve problems. Second, benchmarks are like standardized tests; they are the exams we give these students to see if they actually know their stuff or if they are just guessing.

Right now, most AI "students" in Latin America are being taught by teachers from other countries using textbooks written in other languages. While these students are great at general tasks like math or coding, they often stumble when asked to do things specific to their local neighborhood, like understanding local slang, recognizing local plants, or navigating local laws. The big question isn't just about getting smarter; it's about whether a region can build its own brain that truly understands its own people, or if it has to rely forever on a brain built elsewhere that might not care about its specific needs.


The Two Missing Layers of the AI House

According to a new paper by the SURUS team, Latin America is missing two critical floors in its AI house: a dataset layer and a benchmarks layer. Think of the dataset layer as the library of books the AI reads to learn, and the benchmark layer as the testing center where we check if the AI is actually smart. Without these, the region can't build its own AI, and it can't even properly check if the foreign AI it uses is doing a good job.

The authors argue that this absence creates two different kinds of trouble. For businesses, the problem is like trying to identify pests in a Brazilian farm using a guidebook made for European fields. The task (finding bugs) is the same, but the bugs are different. If you use the wrong data, your AI might miss the pests that actually destroy the harvest. For governments and public institutions, the problem is deeper: it's about identity. If an AI is trained on foreign stories and values, it might describe local citizens in ways that don't fit, or make decisions that don't reflect how those people actually live. It's like having a mayor who speaks a different language and doesn't understand the local culture—they might be polite, but they won't govern you well.

The Great Data Scavenger Hunt

The paper points out that Latin American data isn't gone; it's just lost. It's scattered across the internet like puzzle pieces hidden in different boxes—some in research papers, some on platforms like Hugging Face, and some stuck in private company archives. There is no single map to find them. Even if you could find them all, there simply isn't enough of it compared to what the US or Europe has.

This creates a vicious cycle: because there's no easy place to share data where you get credit, people don't share. Because people don't share, there's no big collection to motivate others to join in. The authors propose a solution called the DataHub, a central library designed to break this loop. It's not just a list; it's a system that rewards people for sharing. If you contribute a dataset, you get visibility and recognition, making it worth your while to open up your data.

A New Way to Organize the Library

One of the most creative parts of the paper is how they suggest organizing this new library. Instead of sorting data by "what it looks like" (like images or text) or "what model it's for," they suggest sorting by what the AI needs to do.

Imagine a library where books aren't sorted by color or size, but by the job they help you do. The authors propose a three-step ladder:

  1. Task: What is the job? (e.g., "Transcribe audio").
  2. Domain: What is the specific field? (e.g., "Medical" vs. "Legal").
  3. Language & Region: Who is speaking? (e.g., "Rioplatense Spanish" vs. "Colombian Spanish").

This matters because a doctor in Buenos Aires needs a different kind of medical transcription AI than a lawyer in Bogotá. A general AI might get the words right, but it might miss the specific medical terms or the local accent. By organizing data this way, the DataHub ensures that the AI you build is tuned exactly to your specific needs, making it perform better and feel more "local."

Open Doors and Incentives

The paper takes a strong stance on how this should be built: Open by design. The authors believe that for a region that is playing catch-up, collaboration is the only way to win. Hiding data behind closed doors won't help; sharing it openly allows everyone to learn faster. However, they also know that people won't share just out of the goodness of their hearts. The system must be "incentive-driven," meaning contributors need to see a benefit, like recognition for their work or their institution's name attached to the data.

The Road Ahead

The authors are clear that this is just the beginning—a "kickstart" rather than a finished product. They admit there are still big hurdles. Different countries in Latin America have different laws about privacy and data, and there are no agreed-upon rules for how to label or license this data yet. They don't claim to have solved these problems; instead, they are inviting everyone—companies, universities, and governments—to help figure it out.

The paper ends with a call to action: the dataset layer for Latin American AI must be built by the region, for the region, or it won't be built at all. The DataHub is the first step, a living invitation for anyone with data to come in, share, and help build a future where Latin America has its own voice in the world of Artificial Intelligence.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →