← Latest papers
🤖 machine learning

Two Confounds in Cross-Model Value Comparison: Response Determinism and the Access Harness

This paper identifies and corrects for two critical confounds in cross-model value comparisons: response determinism, which conflates genuine value differences with the sharpness of forced-choice commitments, and the access harness, where deployment-specific client layers significantly alter a model's apparent value profile, thereby distorting rankings based on single-draw measurements.

Original authors: Hong-In Won, Jinseok Jang, Hyoseop Kim

Published 2026-07-14
📖 6 min read🧠 Deep dive

Original authors: Hong-In Won, Jinseok Jang, Hyoseop Kim

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to taste-test two different brands of ice cream to see which one is "sweeter." You take one scoop of Brand A and one scoop of Brand B. Brand A is vanilla, and Brand B is vanilla too, but Brand B is served in a tiny, super-concentrated shot glass, while Brand A is in a giant, fluffy bowl. If you only taste one spoonful, Brand B might seem way sweeter just because it's so intense and focused, not because the ice cream itself is fundamentally different.

That is exactly what this paper discovers about Artificial Intelligence (AI) models. For a long time, researchers have been comparing different AI models by asking them a single question: "If you had to choose between Option A and Option B, which one do you pick?" They assumed that if Model X picked A and Model Y picked B, the two models had totally different "personalities" or "values."

The authors, Hong-In Won and Hyoseop Kim, say: "Wait a minute. You're measuring two different things at once, and you're confusing them."

The Two Mix-Ups

1. The "Commitment" Confusion (Determinism)
The first mix-up is about how sharply a model commits to an answer.
Imagine two people answering a quiz.

  • Person A is 100% sure. They scream "A!" every single time you ask.
  • Person B is a bit wishy-washy. They say "A" 60% of the time and "B" 40% of the time.

If you ask them just once, and Person A says "A" and Person B says "B," you might think they are total opposites. But they aren't! They both lean toward "A." Person A is just a loud, decisive "A" fan, while Person B is a quiet, hesitant "A" fan.

The paper shows that when researchers compare AI models using just one single answer (a "single-draw"), they accidentally count this "loudness" or "decisiveness" as a difference in values.

  • The Finding: The authors tested nine different AI models. They found that some models are incredibly decisive (like a model called GPT-5.5, which was 0.95 on a scale of how sure it is), while others are much softer (like the Anthropic "Opus" model, which was only 0.34 when measured through a specific app).
  • The Twist: This "decisiveness" isn't a fixed personality trait you can guess just by looking at the model's name. In one family of models, a smaller, cheaper model was actually more decisive than the big, expensive flagship model. In another family, the bigger model was more decisive. The authors suggest that you can't assume a model's "loudness" just by knowing who made it or how big it is; you have to measure it every time you use it.

2. The "Delivery Truck" Confusion (The Access Harness)
The second mix-up is even sneakier. It's not about the ice cream; it's about the truck that delivers it.
The authors realized that the way you ask the AI a question changes the answer. They call this the "access harness."

  • Imagine you order a pizza. If you call the pizza place directly (the "Raw API"), you get the pizza exactly as the chef made it.
  • But if you order through a specific delivery app (like a "CLI" or subscription tool), that app might add a note to the kitchen: "Make it extra spicy!" or "Don't let the customer say no."

The authors found that for some AI models, the "delivery app" (the software used to talk to the AI) completely changed the model's personality.

  • The Proof: They took the same AI model and asked it the same questions twice: once through the direct connection and once through a popular subscription app.
  • The Result: Through the subscription app, the "Opus" model looked very soft and unsure (0.34). But when they asked it directly through the raw connection, it was actually quite decisive (0.66). The app had "flattened" its personality.
  • The "Refusal" Trick: In one case, the direct connection showed the model refusing to answer 10% of the time because it felt the question was unfair. But the subscription app forced the model to answer every time, making it look compliant. The authors proved this was caused by a hidden "system prompt" (a set of instructions the app sends along with the question) that told the model, "Just pick a letter, don't argue."

What This Means for the "Personality" of AI

So, what did the paper actually prove?

  • It proved that the "distance" we see between AI models is often fake. It's a mix of how loud they are and what software we used to talk to them.
  • It proved that the software we use to talk to AI (the "harness") is not a neutral pipe; it actively shapes the AI's values.
  • It suggests (but doesn't fully settle) that "decisiveness" varies wildly between models and doesn't follow a simple rule like "bigger models are louder."

What is NOT the answer?
The paper explicitly rules out the idea that we can just look at a model's name and know its values. It also rules out the idea that all the differences between models are just "noise."

  • The Good News: Even after fixing these two mix-ups, the authors found that some models do genuinely disagree. For example, a model called "Grok-3" really does lean in a different direction than models from OpenAI or Google on about 73% to 88% of the questions. These are real differences, not just measurement errors.
  • The Bad News: Most of the differences we thought we saw inside a single company's family of models (like between different Anthropic models) were mostly just differences in how loud they were, or how the app they were using was shaping them.

The Takeaway

If you want to know if two AI models have different values, you can't just ask them one question and see what they say. You have to:

  1. Ask them the same question many times to see if they are just "loud" or actually "different."
  2. Make sure you are talking to them through the exact same "door" (the same software connection), because the door itself might be changing their mind.

The authors didn't solve the mystery of AI personalities forever, but they built a better magnifying glass to look at them. They showed us that what we thought was a deep philosophical disagreement might just be one AI shouting and another AI whispering, or one AI talking to a bossy manager and another talking to a quiet friend. The real differences are there, but you have to clean up the noise to see them.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →