← Latest papers
📈 economics

Model Selection as Policy Choice: Open-Weight LLMs Converge on Institutions but Diverge on Economic Policy

This study demonstrates that open-weight large language models are not interchangeable tools for policy analysis, as model selection—rather than sampling temperature—drives stable, structured divergences in economic policy preferences while maintaining consensus on institutional and environmental issues, thereby making model choice a critical governance decision.

Original authors: Tamas Olah, Marianna Abuczki, Nayna Babar, Tibor Tokes

Published 2026-07-01
📖 6 min read🧠 Deep dive

Original authors: Tamas Olah, Marianna Abuczki, Nayna Babar, Tibor Tokes

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: Not All AI Chefs Cook the Same Dish

Imagine you are a government official trying to decide on a new economic policy. You ask an AI for advice. Most people assume that if they ask five different AI models the same question, they will get five slightly different versions of the same answer, like asking five different people to recite the same poem.

This paper argues that this assumption is wrong.

Instead of being interchangeable tools, the authors found that different AI models are like different chefs with distinct personalities and political philosophies. If you ask them the same question about the economy, they don't just give different words; they often choose completely different recipes based on their own hidden "tastes."

The Experiment: The AI Taste Test

The researchers set up a massive taste test.

  • The Pantry: They gathered 40 different "open-weight" AI models (these are models anyone can download and run on their own computers, unlike the secret ones owned by big tech companies).
  • The Menu: They gave every single model the exact same questionnaire. The questions covered tough topics like inflation, unemployment, housing, and international relations.
  • The Rules: For every question, the models had to rate four different possible answers (A, B, C, or D) on a scale of 0 to 100.
  • The Repetition: To make sure the results weren't just random glitches, they asked each model the same questions 100 times each.

What They Discovered

1. The "Personality" is Real, Not Random

When you flip a coin, the result is random. When you ask an AI a question, you might expect some randomness too. The researchers found that randomness (noise) was actually very small.

  • The Analogy: Imagine a choir. If the singers were just making random noises, the sound would be a mess. But here, the researchers found that the "voice" of each singer (each AI model) was very distinct and stable.
  • The Result: About 57% of the differences in answers came from which model was answering. Changing the "temperature" (a setting that makes the AI more or less random) barely changed the answer (less than 2%). This means the models have stable, built-in biases that don't go away just because you tweak a setting.

2. They Agree on "Who's in Charge," But Fight Over "How to Spend Money"

The researchers used a statistical map (like a GPS for ideas) to see where the models stood. They found two main "fault lines" where the models differed:

  • The "Trust in Institutions" Line (Where they mostly agree): Almost all models agreed that international organizations, rules, and formal institutions are important. They all generally liked the idea of following the rules and trusting experts.
  • The "Economic Philosophy" Line (Where they fight): This is where the models split up.
    • Group A (The Market Lovers): These models prefer letting the free market decide, cutting regulations, and prioritizing growth.
    • Group B (The Interventionists): These models prefer the government stepping in to fix inequality, control prices, and protect the environment, even if it slows down growth.

The "Stagflation" Test:
The researchers asked a classic hard question: "The country has high inflation AND high unemployment. What do we do?"

  • Some models said: "Crush inflation, even if people lose jobs."
  • Others said: "Save the jobs, even if prices go up."
  • Others said: "Make strict rules for both."
  • Others said: "Use a mix of weird, experimental controls."
  • The Shock: Every single one of these four answers was the "top choice" for at least some of the models. There was no single "AI answer."

3. The "Confidence" Trap

One of the most interesting findings was about confidence.

  • Some models were very shy and said, "I'm not sure" (low confidence scores).
  • Others were very loud and said, "I know the answer!" (high confidence scores).
  • The Twist: A model's confidence didn't depend on how hard the question was. It depended entirely on which model it was.
  • The Analogy: It's like a student who always gets a C but says "I'm 100% sure I'm right," while another student gets an A but says "I'm only 50% sure." If you are a policy maker, you might be tricked into trusting the loud, confident model even if it's just as biased as the quiet one.

4. Size Doesn't Mean "Better" Politics

You might think a bigger, smarter AI (with more "brain power") would have a more balanced or "correct" political view.

  • The Finding: Bigger models were slightly more consistent (less noisy), but they did not have a different political opinion than smaller models.
  • The Analogy: A giant, super-smart robot and a small, simple robot might both be "Market Lovers" or both be "Interventionists." You cannot guess a model's political bias just by looking at its size or brand name.

The Bottom Line: Choosing a Model is a Political Choice

The paper concludes that choosing which AI to use is not a technical detail; it is a policy decision.

  • The Metaphor: If you hire a consultant to write a report on the economy, you are hiring a person with a specific worldview. If you hire AI Model X, you are hiring a "Market Conservative." If you hire AI Model Y, you are hiring a "Government Interventionist."
  • The Danger: If you don't know this, you might think the AI is giving you a neutral, objective fact. But in reality, the AI is quietly injecting its own hidden values into your policy decisions.

The Takeaway:
Before using AI to help make government decisions, we need to treat it like a human expert: Check their background, know their biases, and don't just trust one person's opinion. We need to ask multiple different AIs and compare their answers to see the full picture, rather than letting one "default" model decide our future for us.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →