← Latest papers
🤖 machine learning

The Moving Target: A Longitudinal Audit of Trustworthiness Drift Across Twelve Checkpoints of Open-Source Chat LLMs

This paper demonstrates through a longitudinal audit of four open-source LLM release lines that trustworthiness scores significantly drift across successive checkpoints, arguing that such metrics must be treated as dated, checkpoint-specific artifacts rather than static attributes carried forward without remeasurement.

Original authors: Zhichao Fan, Yanhang Li, Zexin Zhuang, Xian Sun, Yingshuo Wang

Published 2026-07-07
📖 5 min read🧠 Deep dive

Original authors: Zhichao Fan, Yanhang Li, Zexin Zhuang, Xian Sun, Yingshuo Wang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you buy a car from a trusted brand, like "Mistral" or "Qwen." The salesperson hands you a brochure (a "Model Card") that says, "This car is 90% safe and 95% honest," based on tests they ran last year.

Now, imagine that six months later, the same brand releases a slightly updated version of that car. They haven't changed the name, but they might have tweaked the engine, updated the software, or even swapped the tires.

The big question this paper asks is: Can you still trust the old brochure? Does the "90% safe" score from the old version automatically apply to the new one?

The authors say: No. Absolutely not.

Here is the breakdown of their findings using simple analogies:

1. The "Moving Target" Problem

The researchers looked at four major open-source AI families (Yi, Qwen, Mistral, and Gemma). For each family, they checked three consecutive versions (Generation 1, 2, and 3).

Think of these AI models like chefs in a kitchen.

  • Generation 1 is the chef's first recipe.
  • Generation 2 is the same chef, but maybe they changed the spice mix, the cooking time, or the type of pan they use.
  • Generation 3 is the chef again, perhaps with a new assistant or a different oven.

Even though the chef's name on the menu hasn't changed, the food they serve does change. The paper found that the "trustworthiness" of the AI shifts significantly between these versions. It's like if a chef's "95% delicious" rating from last month drops to 85% or jumps to 98% just because they changed the recipe slightly.

2. The "Drift" is Real and Big

The team measured how much the AI's performance "drifted" (moved) between versions.

  • They found the average shift was 8.00 percentage points.
  • To put that in perspective, they compared this shift to what you would expect if the AI were just flipping a coin randomly. The actual shift was 3.6 times larger than random noise.

The Analogy: Imagine you are trying to hit a bullseye on a dartboard. If the board itself is moving around wildly every time you throw a dart (even if you throw from the same spot), your score from yesterday tells you nothing about your score today. The paper proves the "dartboard" (the AI's behavior) is moving much more than we thought.

3. The "Certificate" is Expired

Currently, companies often put a trust score on a model card and assume it stays valid for the whole "release line."

  • The Paper's Verdict: A trust score is not a permanent certificate like a driver's license. It is more like a weather report.
  • If you read a weather report saying "It is sunny" from last Tuesday, that report is useless for deciding what to wear today. You need a new report for today's conditions.

The authors argue that every time a new version of an AI is released, the old trust scores should be treated as expired. You cannot carry them forward to the new version without re-testing.

4. The "Longitudinal Model Card" Solution

Since the scores change so much, the authors propose a new way to report them, which they call a Longitudinal Model Card.

Think of this like a fitness tracker log instead of a single photo.

  • Old Way: A photo of you today saying, "I weigh 150 lbs." (This doesn't tell you if you gained or lost weight since last week).
  • New Way (Longitudinal Card): A log that says, "On Jan 1st, I weighed 150 lbs. On Feb 1st, I weighed 152 lbs. The change was +2 lbs."

This new card would include:

  • The specific "snapshot" (Checkpoint): Exactly which version of the AI was tested.
  • The Date: When the test happened.
  • The Drift: How much the score changed from the previous version.

5. What This Does Not Mean

The paper is very careful about what it doesn't claim:

  • It's not about "Better" or "Worse": The AI didn't necessarily get "dumber" or "safer." It just changed. Sometimes the score went up, sometimes it went down. The point is that it is unpredictable.
  • It's not about Closed AI: This study only looked at open-source models that anyone can download. It doesn't say anything about the secret, closed models from big companies (like the ones you chat with on a website) because those are harder to test.
  • It's not about "Magic" fixes: The paper doesn't offer a way to stop the AI from changing. It just says, "Stop pretending it doesn't change."

The Bottom Line

If you are buying, regulating, or using an open-source AI, do not assume the trust score on the brochure applies to the new version. The AI is a "moving target." You must re-test every new version to know what you are actually getting. The old score is just a dated artifact, not a guarantee.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →