Whose Alignment? Comparing LLM Process Alignment Across Diverse Organizational Decision Contexts
This paper argues that aligning LLMs with organizational decision-making requires measuring process alignment (how information is weighted) rather than just output agreement, demonstrating through contrasting case studies that while process alignment predicts accuracy in legal contexts, it may encode discriminatory patterns in contested domains like consumer credit, thereby revealing the pluralistic nature of organizational alignment.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are hiring a new assistant to help your company make important decisions. You want this assistant to think and decide exactly the way your company does.
Most people think "alignment" just means: "Does the assistant give the same final answer (Yes/No) as the company?"
This paper argues that is a dangerous way to look at things. It's like judging a chef only by whether the food tastes good, without checking if they used the right ingredients or followed the recipe. The authors propose a new way to measure alignment called CALM (Contextualized Alignment Lens Model). Instead of just checking the final answer, CALM looks at how the assistant weighs the different clues to reach that answer.
Here is the breakdown of their findings using simple analogies:
The Core Problem: The "Right Answer, Wrong Reasons" Trap
Imagine a detective trying to solve a crime.
- The Organization (The Chief): Solves cases by looking at the motive and the alibi.
- The AI (The New Detective): Solves cases by looking at the color of the suspect's shirt and the time of day.
If both the Chief and the AI guess "Guilty" for the same 80% of cases, they look perfectly aligned on paper. But the AI is using the wrong logic. If a new, strange case comes up where the shirt color doesn't matter, the AI will fail, while the Chief will succeed. The paper calls this the "Knowledge Gap."
The Two Experiments
The researchers tested this idea in two very different worlds to see if "thinking like the boss" actually matters.
Study 1: The Courtroom (ECHR Decisions)
The Setup: They asked AI models to act like judges in the European Court of Human Rights.
The Analogy: Think of this as a recipe book. The rules for judging these cases are clear, written down, and stable. Everyone agrees on what ingredients (clues) matter.
The Result:
- When the AI models were told to think like the judges, they got much better at it.
- The Magic Link: In this world, if the AI started thinking like the judge (using the right clues), it almost always got the right answer.
- Takeaway: When the rules are clear and agreed upon, making the AI "think like the organization" is a great way to make it accurate.
Study 2: The Bank (German Credit Decisions)
The Setup: They asked AI models to act like a German bank from the 1990s deciding who gets a loan.
The Analogy: Think of this as a broken compass. The bank's old rules were based on history, but some of those rules were unfair (discriminating against women, older people, or foreign workers). The "recipe" was flawed.
The Result:
- The Magic Link Broke: Here, making the AI "think like the bank" did not make it better at predicting who would pay back the loan. The AI could think exactly like the bank and still get the wrong answer, or think differently and get the right answer.
- The "Over-Correction" Glitch: When they told a model, "Hey, you're rejecting too many people, fix it," one model went crazy and approved everyone (99.5% of people), breaking the system entirely.
- The Fairness Fight: The AI models seemed to "know" that some of the bank's old rules (like judging based on gender or age) were unfair. They actively resisted following the bank's instructions on those specific points.
- Takeaway: In messy, controversial situations, just copying the organization's thinking process isn't always good. Sometimes the organization's "way of thinking" is actually the problem.
The Big Conclusion
The paper concludes that we can't just ask, "Does the AI agree with us?" We have to ask, "Whose values are we aligning with?"
- In clear, fair systems (like the Court): Aligning the AI's thinking process with the organization is helpful and makes it more accurate.
- In messy, unfair systems (like the old Bank): Aligning the AI's thinking process with the organization might just make the AI a better discriminator.
The Final Metaphor:
If you are training a dog to fetch a ball:
- Study 1 is like training the dog to fetch a ball in a park. If the dog learns the right way to fetch, it gets the ball every time.
- Study 2 is like training the dog to fetch a ball that is actually a bomb. If the dog learns to fetch the bomb "exactly like the owner wants," it's a disaster.
The authors say we need a tool (CALM) to check how the AI is thinking, not just what it decides. This helps us see if the AI is following a good recipe or a bad one, which is crucial for high-stakes decisions like loans and legal rulings.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.