Silent Updates: Measuring and Closing the Post-Deployment Disclosure Gap
This paper investigates the "post-deployment disclosure gap" in AI systems, revealing that while providers publish safety documentation, they lack mechanisms to externally verify that deployed models match their reported versions, prompting the authors to propose a "Silent Updates Scorecard" and a "Three-Part Behavioral Trigger System" to enforce transparency and accountability.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are at a massive, magical buffet where the chefs are super-smart robots. You order a specific dish, let's say "The GPT-5 Burger," and the menu tells you exactly what's in it, how safe it is to eat, and what it tastes like based on a famous food critic's review. You feel safe because you trust the menu. But here's the twist: the chefs can secretly swap the burger patty, change the secret sauce, or even replace the whole bun with something completely different while you are still eating, without telling you, changing the name on the plate, or updating the menu. In the world of Artificial Intelligence, this is what happens with "foundation models." These are the giant AI brains that power chatbots and tools. A key concept here is the "chain of custody," which is like a receipt that proves the burger you are eating is the exact same one the critic tasted. Another concept is "silent updates," which are those secret swaps the chefs make. People care about this because if the menu says the burger is safe, but the chef secretly added spicy poison, the review is useless, and you might get sick. We need to know if the thing we are using is actually the thing that was tested.
Now, let's look at what the researchers Sophia Abraham and Ben Bucknall did. They acted like food inspectors, but instead of checking restaurants, they checked nine big AI companies and seven places that host these AI tools. They created a special checklist called the "Silent Updates Scorecard" to see if these companies are honest about their secret ingredient swaps. They looked for proof that the AI you talk to today is the exact same one that was tested in their safety reports.
The results were a bit like finding out the buffet has a "ghost kitchen." The researchers found that while the chefs (the AI companies) are very good at writing detailed menus and safety reports (publishing safety docs and evaluations), they are terrible at proving that the food on your plate matches the menu. In fact, out of the nine main AI providers they checked, zero of them gave a way for an outsider to verify that the specific AI model being served was the exact same one described in their safety reports. It's like the chef saying, "Trust me, this is the burger from the menu," but refusing to let you check the kitchen or see the receipt.
The paper suggests that this happens because the AI systems aren't static; they are constantly changing in the background. The researchers identified four main ways this "silent update" game is played:
- The Shapeshifting Name: Companies use stable names like "GPT-5" or "Claude," but these names are like magic labels that stick to different burgers over time. One day "GPT-5" might be a beef patty, and the next month it's a veggie patty, but the name stays the same.
- The Missing Logbook: While companies are great at announcing when they launch a new burger, they rarely write down when they secretly change the ingredients of an old burger that's already on the menu.
- The Two Faces: Sometimes the burger you get on the website (the chatbot) is different from the one you get if you ask for it through a computer program (the API), and the company doesn't tell you they are different.
- The Unlinked Receipt: The safety reports often talk about a "family" of models (like "the GPT-5 family") rather than a specific, unchangeable version. This means the safety report might be about a burger from last year, but you are eating one from today.
The researchers also found that some companies have rules in their contracts that actually stop people from checking the food. For example, six out of the nine companies they looked at have terms that say you can't run your own tests to see if the AI is behaving differently than the menu says. It's like a restaurant saying, "You can read our menu, but if you try to taste-test the food to see if it matches, we will kick you out."
To fix this, the authors propose a new system called the "Three-Part Behavioral Trigger System." Think of this as a new set of kitchen rules. Instead of waiting for a chef to admit they changed the recipe, the rules would say:
- If the AI starts acting differently (like refusing to answer questions it used to answer), that's a "Drift Trigger," and they must update the menu.
- If they change a specific part of the machine (like the sauce dispenser or the grill), that's a "Component Trigger," and they must log it immediately.
- If the AI suddenly gets much smarter or much more dangerous, that's a "Capability Trigger," and they must re-test the whole thing.
They also suggest a "safe harbor" rule, which would be like a law saying, "If you are a food critic trying to check the safety of the burger for the public good, the restaurant can't sue you or kick you out for doing it."
The paper doesn't claim that these companies are lying or that the AI is dangerous right now. Instead, it suggests that the current system is broken because there is no way to prove the link between the safety report and the actual AI you are using. The authors measured this using public information and found that while the paperwork exists, the "chain of custody" is broken. They suggest that until we can verify that the AI we are using is the same one that was tested, we are eating in the dark.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.