Every Act Has Its Price: Compressed Moral Composition in Frontier LLMs
This paper introduces the "Moral Trolley Arena," a two-stage benchmark revealing that frontier LLMs compose multiple moral signals into judgments through a consistently compressed, non-additive mechanism rather than simple summation, suggesting that moral evaluations should focus on these composition rules alongside isolated act rankings.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to understand the "moral compass" of a super-smart AI. For a long time, researchers tested these AIs by asking them simple, isolated questions: "Is it better to save a cat or a dog?" or "Is lying always bad?"
This paper argues that real life isn't that simple. Real choices are like a smoothie, not a single fruit. You don't just taste "strawberry"; you taste strawberry mixed with banana and a little bit of sour lemon. The AI has to figure out how those flavors mix together.
The authors created a new testing ground called Moral Trolley Arena to see how AIs mix these moral "flavors." Here is how they did it and what they found, explained simply:
1. The Setup: Calibrating the Ingredients
First, the researchers had to measure the "strength" of individual moral acts on their own. They took 229 different scenarios (like "saving a stranger from a fire" or "a judge taking a bribe") and ranked them using a game-like system called ELO (the same system used to rank chess players).
- The Result: They found that for all the AI models tested, "Authority" (following rules/laws) was usually the strongest flavor, and "Sanctity" (purity/holy things) was usually the weakest. This is the "single-ingredient" taste.
2. The Experiment: Mixing the Smoothie
Next, they stopped asking about single acts. Instead, they presented the AI with profiles that combined two different acts.
- Profile A: A judge who takes bribes (Bad) but also protects the community from deportation (Good).
- Profile B: A hero who saves a baby from a fire (Very Good) but also reminds a waiter about an unpaid bill (Mildly Good).
The AI had to choose which profile was "better" to save. The researchers then compared the AI's choice against the simple math of adding the two scores together.
3. The Big Discoveries
A. The "Compression" Effect (The Volume Knob)
If AIs were simple calculators, they would just add the scores: Bad (-10) + Good (+10) = 0.
But the paper found something different. The AIs act like they have a volume knob turned down.
- The Metaphor: Imagine you are mixing two loud sounds. A simple mixer would make the result twice as loud. These AIs, however, "compress" the volume. Even if you add a very good act to a very bad act, the final judgment isn't as extreme as the math suggests. The AI dampens the total impact of the combination.
B. The "Intensity Anchor" (The Loud Speaker)
The researchers found that the AI cares more about the loudest voice in the mix than the average of the voices.
- The Metaphor: Think of a band. If you have one guitarist playing a super loud, amazing solo (+2) and one drummer playing a quiet, slightly annoying beat (-1), the audience remembers the solo.
- The Finding: The AI preferred a profile with one "Super Good" act and one "Mildly Bad" act over a profile with two "Medium Good" acts. Even though the total "goodness" was mathematically similar, the AI was "anchored" by the single, intense positive act. It ignored the fact that the other option was more consistently good.
C. The "Secret Sauce" (Residuals)
After accounting for the math of the acts, the researchers looked for leftover biases.
- Loyalty: When "Loyalty" was part of the mix, the AI gave it a little extra credit, as if it had a secret bonus.
- Care: When "Care" (protecting people) was part of the mix, the AI actually gave it slightly less credit than the math predicted.
- The Others: The other moral foundations (Fairness, Authority, Sanctity) behaved exactly as the math predicted, with no secret bonuses or penalties.
D. Everyone Agrees
Finally, the researchers tested 10 different top-tier AI models from different companies (like OpenAI, Google, Anthropic, etc.).
- The Finding: Even though these models are built differently, they all "mix" moral evidence in almost the exact same way. They all compress the volume, all get anchored by the loudest act, and all have the same Loyalty/Care bias. It's not a glitch in one specific robot; it's a shared trait of current advanced AIs.
The Bottom Line
The paper concludes that we can't just look at how an AI ranks single actions to understand its morality. We have to watch how it mixes them.
Current AIs don't just add up good and bad deeds. They compress the total impact, get distracted by the most intense single act, and have hidden biases for loyalty over care. To truly audit an AI's morality, we need to test these "mixing rules," not just the ingredients.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.