← Latest papers
💻 computer science

On the missing benchmarks layer and a potential solution

This paper identifies the absence of a regional benchmark layer as a critical barrier to Latin America's native AI development and proposes "EvalsHub" (with its initial instance, LatamBoard) as an open, incentive-driven infrastructure to enable independent auditing and optimization of AI systems against local social and economic requirements.

Original authors: Francis F Daniel, Mauro Ibañez, Francis Perelman, Marian Basti

Published 2026-08-05
📖 6 min read🧠 Deep dive

Original authors: Francis F Daniel, Mauro Ibañez, Francis Perelman, Marian Basti

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Invisible Scoreboard: Why AI Needs a Local Test

Imagine the world of Artificial Intelligence as a massive, high-tech kitchen where chefs (the AI models) are cooking up dishes for everyone to eat. For a long time, the most famous chefs have been working in just a few specific kitchens, using recipes and ingredients from their own neighborhoods. They are incredibly talented, but they don't always know how to cook a dish that tastes perfect for a family living in a different part of the world. This is the world of AI: powerful tools built in one place, but often used everywhere.

To know if a chef is actually good at cooking a specific meal, you need a taste test. In the world of AI, this taste test is called a benchmark. Think of a benchmark as a standardized exam or a scoreboard. It doesn't just ask, "Can the AI talk?" It asks, "Can the AI solve this specific problem in this specific language with these local rules?" Without these local scoreboards, countries and companies are like diners eating a meal without knowing if it's fresh, safe, or even what it's supposed to taste like. They have to trust the chef's word alone. This paper argues that Latin America is missing its own kitchen scoreboard, and without it, the region can't truly check if the AI it uses is working correctly or improve it to solve its own unique problems.


The Missing Scoreboard: A Story of Latin America's AI Gap

The authors of this paper, a team from SURUS and an independent researcher, are pointing out a huge hole in the map of Latin American technology. They call it the "missing benchmark layer." To understand why this matters, imagine you are trying to fix a car, but you only have tools designed for a different brand of car. You might get the job done, but it won't be perfect, and you won't know exactly where the engine is sputtering. That is what Latin America is facing with AI.

The Two Big Problems: Blind Audits and Stalled Engines

The paper explains that this missing layer causes two major headaches.

First, there's the Audit Problem. Imagine a government agency buying a new AI system to help manage public records. Because the AI was built in another country, the agency has no way to independently check if it understands local laws or speaks the local dialect correctly. They are forced to just trust the seller's claims. The authors suggest that without a local benchmark, these institutions are "blind." They can't see if the AI is making mistakes that only matter in their specific context, like misidentifying a local crop disease or misunderstanding a specific legal term.

Second, there's the Optimization Problem. This is for the companies trying to build cool AI tools. Modern AI development is like a video game where you tweak settings to get a higher score. But to get a higher score, you need a scoreboard to measure your progress. The paper argues that without a local benchmark, companies have no target to aim for. They can't tell if their AI is getting better at solving local problems because they have no "exam" to grade it against. It's like trying to run a race without a finish line; you might be running fast, but you don't know if you're actually winning.

The Proposed Solution: The EvalsHub and LatamBoard

So, what is the fix? The authors propose building a new infrastructure called an EvalsHub, with the first version being LatamBoard.

Think of LatamBoard as a giant, open-source "exam hall" for the region. It's a place where universities, governments, and companies can publish their own tests (benchmarks) and see how different AI models perform on them.

  • It's Task-First: Instead of just testing "how smart is the AI?", the tests are built around specific jobs. For example, a test might be designed specifically for "extracting medical info from Chilean health records" or "classifying pests in Colombian coffee crops."
  • It's Reusable: The paper uses a great phrase: "Built once, measured forever." Once a test is created, it can be run over and over again. Every time a new AI model comes out, or a company tweaks their software, they can run the same test to see if they improved. This turns the benchmark into a permanent piece of infrastructure, like a road or a bridge, rather than a one-time project.

The "Access Problem": Why It's Hard to Build

The paper also admits that building these scoreboards is really hard. It's not just about writing code. To create a valid test for a specific job (like diagnosing a disease), you need four different types of experts working together:

  1. A Domain Expert (like a doctor or lawyer) who knows what "correct" looks like.
  2. An ML Engineer who can turn that knowledge into a computer test.
  3. A Statistician to make sure the test questions are fair and representative.
  4. A Developer to make sure the test runs smoothly.

The authors call this the "Access Problem." Usually, the person who knows the most about the job (the doctor) doesn't know how to code the test, and the coder doesn't know the medical details. This mismatch means very few people can build these benchmarks on their own. The paper suggests that for LatamBoard to work, they need to find a way to lower the barrier so that experts can contribute their knowledge without needing to be coding wizards.

A Vision for the Future: Open and Multipolar

The authors have a strong stance on how this should work. They want a "multipolar" world, meaning they don't want one or two countries to control all the AI rules. They want Latin America to have its own voice and its own standards.

To make this happen, they propose that LatamBoard must be open by design (everyone can see the tests and results) and incentive-driven by construction (people get credit and recognition for helping). If a university creates a great test for local agriculture, they get to be named as a contributor, and everyone else benefits from that test. It's a system where sharing knowledge makes everyone smarter and more powerful.

What's Still Unknown?

The paper is honest that this is a starting point, not a finished product. They admit there are still big questions to answer, like:

  • How do we make sure the tests stay fair as AI gets smarter?
  • Who pays for the computers needed to run these tests?
  • How do we handle sensitive areas like healthcare where privacy is super important?

The authors aren't claiming to have solved all these problems yet. Instead, they are throwing out a challenge to the community. They are inviting universities, governments, and companies to come together and build this layer of infrastructure. The goal is to create a system where Latin America isn't just a consumer of AI, but a shaper of it, with its own tools to measure, improve, and trust the technology it uses.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →