← Latest papers
💻 computer science

Political Neutrality as Balanced Approval: A Large-Scale Human Evaluation of AI Responses

This paper introduces a new definition of AI political neutrality based on maximizing and balancing approval across opposing political groups, validates it through a large-scale human evaluation dataset (PARETO) involving over 7,000 participants, and reveals that while frontier models can achieve high cross-partisan approval, their default responses often exhibit a liberal bias.

Original authors: Jonathan Stray, David Zhai Yang, Steven Luo, Miu Nicole Takagi, Serina Chang

Published 2026-05-29
📖 5 min read🧠 Deep dive

Original authors: Jonathan Stray, David Zhai Yang, Steven Luo, Miu Nicole Takagi, Serina Chang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are hosting a dinner party where two guests, let's call them "Lefty" and "Righty," have a fierce disagreement about a topic like abortion or gun control. They are shouting over each other, and neither can agree on what the "truth" is.

Now, imagine you hire a smart robot waiter (the AI) to serve them. Your goal isn't to pick a side or tell them who is right. Your goal is to serve a dish that both Lefty and Righty can agree is delicious, even though they still disagree on the recipe.

This paper is about teaching that robot waiter how to do exactly that.

The Big Problem: The "Refusal" Trap

Previously, when people asked these robots controversial questions, the robots often did one of two things:

  1. They picked a side: They sounded like Lefty or Righty, making the other person angry.
  2. They refused to answer: They said, "I can't talk about that," which made everyone feel ignored.

The researchers wanted to know: Can we make an AI that answers these tough questions in a way that makes both sides say, "Okay, I can live with that answer"?

The New Recipe: "Balanced Approval"

The authors came up with a new definition of "neutrality." They call it Balanced Approval.

Think of it like a tightrope walker.

  • The Rope: This is the line of maximum possible happiness.
  • The Goal: The AI needs to stand on the tightrope where Lefty is happy AND Righty is happy at the same time.
  • The Catch: If the AI leans too far to the left, Righty gets mad. If it leans too far right, Lefty gets mad. The "perfect" neutral answer is the exact spot on the rope where both sides are equally satisfied.

The researchers call this the "Pareto Frontier." In plain English, it's the "best possible deal" where you can't make one side happier without making the other side unhappy.

The Experiment: A Massive Taste Test

To test this, the researchers didn't just ask the robots to guess. They set up a massive taste test:

  1. The Menu: They picked 20 hot-button topics (like abortion, gun control, and student debt).
  2. The Diners: They recruited 7,434 real people from the US.
  3. The Chefs: They used 5 different top-tier AI models (like GPT, Claude, and Gemini).
  4. The Dishes: For every question, they made four types of answers:
    • The Default: What the AI says naturally.
    • The "For" Dish: An answer arguing only for one side.
    • The "Against" Dish: An answer arguing only for the other side.
    • The "Balanced" Dish: A special answer that gave a short paragraph for both sides, then wrapped it up neutrally.

Then, they asked the diners: "How much do you approve of this answer?"

What They Found: The Surprising Results

1. The "Balanced Dish" was the crowd favorite.
Even though Lefty and Righty hated each other's ideas, they both liked the balanced answer.

  • The Magic Number: The balanced answer got high approval scores (above 60%) from both sides on almost every topic.
  • The Trade-off: You might think, "If I want my side to win, I'd hate a balanced answer." But the study found that people only lost about 10% of their approval when switching from a "my side wins" answer to a "balanced" answer. That's a small price to pay for an AI that doesn't make the other side scream.

2. Most Robots are secretly "Lefty."
When the researchers looked at what the robots said naturally (without any special instructions), most of them (GPT, Claude, Gemini, Llama) leaned toward the liberal/left side. They got more approval from Lefty than Righty.

  • The Exception: One robot, Grok, was different. It didn't lean left; sometimes it leaned right, sometimes left, depending on the topic.

3. Angry Questions are Harder to Answer.
When users asked the robots questions that were already angry or charged (e.g., "Why are liberals so stupid?"), the robots got lower approval scores. It's much harder to be neutral when the question itself is a fight.

4. The "Both Sides" Myth.
Some people worry that showing "both sides" is fake or weak. But the study showed that when the AI presented strong arguments for both sides, people actually felt it was fairer. The only time people hated the balanced answer was when they felt the AI was "faking it" or ignoring their specific concerns.

The Bottom Line

This paper proves that it is possible to build AI that acts like a fair mediator. It doesn't have to be a robot that refuses to speak, nor does it have to be a robot that picks a team.

By aiming for Balanced Approval—finding the answer that makes the most people on opposing sides feel heard and respected—we can create AI that people trust, even when they disagree with each other. It's like a robot waiter who manages to serve a meal that satisfies the entire family, even if they can't agree on what to order next.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →