← Latest papers
💻 computer science

What Should Frontier AI Developers Disclose About Internal Deployments?

This paper proposes a framework for frontier AI developers to disclose specific information regarding their internal model deployments—categorized by capabilities, usage, safety mitigations, and governance—to ensure safety and oversight while balancing transparency with potential risks.

Original authors: Jacob Charnock, Raja Mehta Moreno, Justin Miller, William L. Anderson

Published 2026-04-28
📖 3 min read☕ Coffee break read

Original authors: Jacob Charnock, Raja Mehta Moreno, Justin Miller, William L. Anderson

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The "Secret Lab" Problem: Why We Need to Peek Inside AI Factories

Imagine you live next door to a massive, high-tech factory. For years, you’ve seen the finished products they ship out—shiny cars, gadgets, and appliances. You know how they work because you use them every day.

But lately, you’ve noticed something strange. Inside the factory, they’ve started using super-intelligent robots to build the products. These robots aren't just moving boxes; they are designing the blueprints, writing the software, and even deciding how the next generation of machines should look.

The problem? The factory doors are bolted shut. You have no idea if these "builder robots" are following the rules, if they are accidentally breaking things, or if they are secretly redesigning the factory to work for themselves instead of the humans.

This paper, written by researchers at MATS, is essentially a "Transparency Manifesto" for these AI factories.


The Core Issue: The "Internal Deployment" Blind Spot

In the AI world, "Internal Deployments" (IDMs) are those super-intelligent robots. They are AI models used inside the companies (like OpenAI or Anthropic) to help researchers build even better AI.

Right now, the world only sees the "finished products" (the AI you chat with on your phone). We are totally blind to the "builder AI" working behind the scenes. If the builder AI makes a mistake in a blueprint or a piece of code, it could create a "glitch" that stays hidden until it’s too late.

The Solution: The Four-Folder Filing System

The authors argue that if these companies want us to trust them, they shouldn't just show us the finished car; they need to show us a "safety report" on their builder robots. They suggest organizing this info into four "folders":

  1. The "Muscle" Folder (Capabilities): How strong are these robots? Can they write code? Can they hack into systems? We need to know if the robots are getting exponentially stronger behind closed doors.
  2. The "Job Description" Folder (Usage): What exactly are they doing? Are they just helping humans write emails, or are they autonomously redesigning the factory's security system?
  3. The "Safety Harness" Folder (Mitigations): What's stopping a robot from going rogue? Do they have "kill switches"? Are there human supervisors watching every move, or are the robots left alone for hours at a time?
  4. The "Rulebook" Folder (Governance): What are the company's laws? If a robot starts acting weird, is there a protocol to report it? Who is allowed to touch the "master controls"?

"But won't that leak our secrets?" (The Counter-Argument)

The companies might say, "If we tell everyone how our robots work, our competitors will steal our recipes!"

The researchers have a clever answer for this. They suggest a "Two-Tiered Disclosure" system:

  • The Public Summary: Like a nutrition label on a cereal box. It tells you the big stuff (calories, sugar) without giving away the secret spice blend.
  • The Private Audit: Like a deep medical exam. The company shares the highly sensitive, "secret recipe" details with trusted doctors (regulators and independent experts) who are legally bound to keep the secrets but can tell us if the patient is healthy.

The Bottom Line

The paper is a call to action for lawmakers and AI developers. It says: "We can't just watch the storefront; we need to know what's happening in the workshop." By creating a standard way to report on these internal "builder AIs," we can make sure the machines building our future are actually following our instructions.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →