← Latest papers
💻 computer science

Benchmarking Open-Weight Foundation Models for Global AI Technical Governance

This study addresses methodological limitations in existing research on geographic bias in AI by benchmarking four open-weight foundation models against a verified ground-truth dataset (GAID v2) using a refined five-category response classification scheme to accurately measure and analyze disparities in model performance across 227 countries.

Original authors: Jason Hung

Published 2026-06-26
📖 5 min read🧠 Deep dive

Original authors: Jason Hung

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have four very smart, well-read librarians (the AI models). You ask them a million questions about how different countries are doing with Artificial Intelligence (AI). You want to know if they are telling the truth or if they are just making things up to sound confident.

This paper is a report card on those librarians, specifically checking if they are biased against certain countries or if they just get confused about specific types of facts.

Here is the story of what they found, broken down simply:

1. The Setup: The "Truth" vs. The "Guess"

The researchers had a giant, verified "Answer Key" (called the Global AI Dataset) containing real numbers about 227 countries. They asked the four AI models to give them specific numbers from this key.

  • The Goal: See if the AIs get the numbers right.
  • The Problem: Sometimes, when an AI doesn't know an answer, it doesn't say "I don't know." Instead, it confidently makes up a number. This is called a Hallucination (or in the paper, "Confident Fabrication").

2. The Big Surprise: The "Global South" Didn't Get the Short End of the Stick

Usually, people think AI models are biased because they are trained mostly on data from rich, English-speaking countries (like the US and UK). The common belief is that the AIs would make up more lies about poorer countries (the "Global South") because they know less about them.

But the results were the opposite!

  • The AIs actually made up more lies about rich countries (Global North) than poor ones.
  • Why? Think of it like this: If you ask a librarian, "How many grains of sand are on this tiny beach?" and they guess "500," they might be close. But if you ask, "How many grains of sand are on this massive desert?" and they guess "500," they are wildly wrong.
  • Rich countries have huge numbers (millions of patents, billions of dollars in computing power). It is much harder for an AI to guess a huge number correctly than a small number. So, the AIs got "caught" making up big numbers for rich countries more often.
  • Note: This doesn't mean the AIs know everything about poor countries; it just means the "Answer Key" for rich countries had bigger, harder-to-guess numbers.

3. The "Safety" Trap: The AIs Are Terrible at Math

The researchers asked the AIs about specific, hard numbers like "How much computing power was used to train this model?" or "How many parameters does this model have?"

  • The Result: The AIs failed almost completely. They made up numbers 98% of the time for these "Safety" questions.
  • The Analogy: It's like asking a chef, "Exactly how many milligrams of salt are in this soup?" and the chef just guesses a number that sounds fancy but is totally made up.
  • The Lesson: If you need to know the exact size of a computer chip or the power of a supercomputer, do not trust these AIs. They are guessing.

4. The "Yes/No" Success: The AIs Are Good at Simple Facts

When the researchers asked simple "Yes or No" questions, like "Did this country publish a national AI strategy?"

  • The Result: The AIs did much better. They got the answer right about 42% of the time (which is high for them!).
  • Why: It's easier to remember a "Yes" or "No" than to invent a specific number like "14,502 patents."

5. The Librarian Comparison: Who is the Best?

The researchers tested four different "librarians" (AI models):

  1. Mistral (France): The most accurate overall.
  2. DeepSeek (China): Second best.
  3. Qwen (China): Third.
  4. Llama (US): Made up the most lies.

Interestingly, it didn't matter much if the librarian was from the US, Europe, or China. They all made up lies at roughly the same rate. The "nationality" of the AI didn't really change the outcome.

6. The "Don't Know" Problem

The researchers hoped that if the AIs didn't know an answer, they would politely say, "I don't know."

  • The Reality: The AIs almost never said they didn't know. Even when they were totally guessing, they sounded 100% confident.
  • The Danger: This is like a student taking a test who guesses on every question but writes the answers in big, bold letters. You can't tell if they are smart or just lucky.

7. How to Ask Better Questions

The researchers found that how you ask the question matters.

  • Bad Way: "What is the exact number of AI patents in Country X?" (The AI guesses a number).
  • Better Way: "How does Country X's patent count compare to the average for its region?"
  • Result: When you ask for a comparison instead of a specific number, the AI makes up fewer lies. It's like asking, "Is this beach bigger than that one?" instead of "How many grains of sand are here?"

The Bottom Line

If you are a government official or a researcher trying to use AI to understand the world:

  1. Don't trust the numbers: If the AI gives you a specific number about AI power or patents, check it with a human. It is likely a confident guess.
  2. They are bad at "Safety" math: They cannot tell you how much computing power is being used.
  3. They are okay at "Yes/No": They are better at telling you if a law exists than telling you the exact details of it.
  4. They never admit ignorance: They will always answer, even if they are lying.

The paper concludes that while these AI models are impressive, they are currently unreliable tools for gathering precise facts about global AI governance. You still need human experts to verify the data.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →