← Latest papers
🤖 AI

NeurIPS Should Require Reproducibility Standards for Frontier AI Safety Claims

This position paper argues that NeurIPS must mandate reproducibility standards for frontier AI safety claims, proposing a three-tier disclosure framework to address the current "evidential inversion" where the most consequential safety assertions are the least verifiable due to withheld artifacts.

Original authors: Varad Vishwarupe, Nigel Shadbolt, Marina Jirotka, Ivan Flechais

Published 2026-05-12
📖 5 min read🧠 Deep dive

Original authors: Varad Vishwarupe, Nigel Shadbolt, Marina Jirotka, Ivan Flechais

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a world where a company builds a incredibly powerful new machine. Before they sell it to the public, they publish a report saying, "Don't worry, this machine is safe. It won't hurt anyone."

The problem, according to this paper, is that the company often refuses to show you how they tested the machine. They give you the final score (e.g., "99% safe!") but hide the test questions, the scoring rules, and the machine's settings. They say, "Trust us, we checked it."

The authors of this paper argue that the NeurIPS conference (a major gathering for computer scientists) needs to change the rules. They say: "If you want to publish a safety claim about a super-powerful AI, you must prove it. If you can't show your work, you can't make the big claim."

Here is a breakdown of their proposal using simple analogies:

1. The Problem: The "Evidential Inversion"

Usually, in science, the bigger and more important a claim is, the more proof you need to back it up.

  • Normal Science: If a scientist says, "I found a new planet," they must show the telescope data so others can look too.
  • AI Safety (The Problem): If a company says, "This AI is safe enough to run the power grid," they often hide the data.

The authors call this an "Evidential Inversion." It's like a magician claiming their trick is harmless but refusing to let anyone see the deck of cards. The paper argues this is a failure of science, not just a lack of transparency.

2. The Solution: A Three-Tier "Traffic Light" System

The authors know that sometimes, showing the "cards" (the test data) is dangerous. If you show exactly how to break a security system, bad actors might learn how to do it.

So, they propose a Three-Tier System (like a traffic light) for how much information must be shared:

  • 🟢 Tier 1 (Green Light - Public):

    • When: The test data is safe to show everyone.
    • Rule: You must publish everything: the questions you asked the AI, the code you used, and the results. Anyone can re-run the test.
    • Result: Full trust. The claim is fully supported.
  • 🟡 Tier 2 (Yellow Light - Controlled):

    • When: Showing the data publicly would be dangerous (e.g., it teaches people how to hack), but a trusted expert could look at it safely.
    • Rule: You don't publish the data to the public. Instead, you hand it to a secret panel of trusted, independent experts (like a private security clearance). They check the work behind closed doors and give a "thumbs up" or "thumbs down."
    • Result: The claim is supported, but with a note saying, "We checked this privately, but you can't see the raw data."
  • 🔴 Tier 3 (Red Light - Restricted):

    • When: Even a secret panel can't look at the data because it's too dangerous (e.g., the test involves creating a biological weapon simulation).
    • Rule: You can still write the paper, but you cannot make the big claim that the AI is "safe enough to release." You can only say, "We tried to test this, but the data is too sensitive."
    • Result: The claim is "right-scaled." You can't use the paper to justify releasing the AI to the world.

3. The "Claim Inventory" (The Receipt)

To make sure companies don't cheat, the paper suggests a new form authors must fill out, called a Claim Inventory.

  • Think of this like a receipt at a store.
  • The author must list every safety claim they are making.
  • Next to each claim, they must check a box: "Did we show the data? Did we get a secret panel to check it? Or is it too dangerous to show?"
  • If they check "Too dangerous" but don't have a good reason, the paper gets rejected or the claim is weakened.

4. Why NeurIPS?

The authors chose NeurIPS because it is the "Olympics" of AI research.

  • Currently, companies love to publish their safety reports at NeurIPS because it gives their claims scientific legitimacy. It makes the public and governments trust them.
  • The paper argues: "If you want the gold medal (publication at NeurIPS), you have to follow the rules. You can't just say you're safe; you have to prove it, either publicly or to a trusted referee."

5. How to Stop Cheating

The authors anticipate that some companies might try to game the system (e.g., claiming their data is "too dangerous" just to hide it).

  • The Fix: They propose a Federated Review System. Imagine a group of different "security clearances" (like different national safety institutes or independent labs). If a company says their data is too sensitive for NeurIPS, a neutral expert from this group checks it.
  • If the expert says, "No, this isn't actually that dangerous, you should show it," the company has to show it. If they refuse, the paper is rejected.

Summary

The paper is a call to action for the AI community: Stop accepting "Trust Us" as proof of safety.

If you claim a powerful AI is safe, you must either:

  1. Show your work to everyone.
  2. Let a trusted, independent referee check your work in secret.
  3. If you can't do either, admit that you haven't proven it's safe, and don't use your paper to justify releasing the AI.

The goal isn't to stop AI development; it's to make sure that when we say "This is safe," we actually have the receipts to prove it.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →