← Latest papers
💻 computer science

Ideology by Alphabet: Option Order and the Measurement of Machine Political Preferences

This paper demonstrates that the apparent political ideologies of large language models are largely artifacts of fixed answer-order biases rather than genuine preferences, as randomized option ordering dissolves previously identified ideological clusters and reveals that specific slot arrangements can arbitrarily label models with conflicting political stances.

Original authors: Tamas Olah, Laszlo Erdey, Tibor Tokes, Levente Nadasi

Published 2026-09-17
📖 6 min read🧠 Deep dive

Original authors: Tamas Olah, Laszlo Erdey, Tibor Tokes, Levente Nadasi

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

When we ask a computer to make a choice, we often assume it is weighing the facts of the situation. If a large language model—a type of artificial intelligence that can read and write like a human—is asked to pick between four different solutions to a problem, we expect its answer to reflect its understanding of the issue. This is how we treat human voters or expert panels: we assume their preferences are rooted in their values and knowledge. But in the world of survey design, there is a well-known trick of the mind called an "order effect." If you list a series of options for a person to choose from, they are statistically more likely to pick the first one on the list, simply because it is the first one they see. This happens because the brain takes a shortcut when it is tired or when the choices are difficult to distinguish. For decades, social scientists have known that the order of words on a ballot or a questionnaire can change the outcome of a vote. The question that has remained unanswered is whether these same shortcuts apply to machines, and if they do, whether they are strong enough to invent a political personality where none exists.

A team of researchers from the University of Debrecen in Hungary set out to test this by treating artificial intelligence models as if they were political voters. They created a massive test consisting of 400 different scenarios involving real-world governance issues, such as how to handle workplace safety, how to regulate markets, or how to distribute social benefits. For each scenario, they presented the models with four distinct options and asked them to score each one and pick a winner. They ran this test on 15 different open-source models, generating over 640,000 individual responses. In the first phase of the experiment, they used the standard method that most researchers use: they kept the order of the options exactly the same for every single question. If the "government" option was always listed first, it stayed first. If the "individual" option was always last, it stayed last.

When the researchers analyzed the results from this fixed-order setup, they found a pattern that looked remarkably like a genuine political divide. The models sorted themselves into two clear groups. One group, which the researchers called "market-rationalists," consistently favored options related to utility, metrics, and financial incentives. The other group, the "institutionalists," favored options related to rules, fairness, and community responsibility. This split looked so clean and so logical that it resembled a famous theory in political science that divides economies into liberal and coordinated types. The results were so consistent that they passed every standard test for reliability. It appeared as though the researchers had successfully mapped the political ideologies of artificial intelligence.

However, the researchers suspected that this neat political map was an illusion created by the test itself. To find out, they ran the exact same 400 questions again, but this time they scrambled the order of the options. For every single question, they rotated the four choices so that each option appeared in the first slot, the second slot, the third slot, and the fourth slot an equal number of times. This meant that the content of the answer was no longer tied to its position on the screen. When they analyzed the data from this randomized version, the specific political map they had found in the first round dissolved. The clear divide between the "market-rationalists" and the "institutionalists" collapsed. The models still genuinely differed from one another, and the new map of these differences was statistically reliable, but it no longer formed the two distinct camps or the clean typology suggested by the fixed-order test.

The study revealed that the "ideology" they had found in the first round was not a reflection of the models' beliefs at all, but was largely a product of the order in which the options were printed. The three models that had been labeled "market-rationalists" were simply the three models that had the strongest habit of picking the very first option on the list. Because the researchers had accidentally placed the "market" options in the first slot for every question, these three models picked them every time, creating a fake political signature. The researchers calculated that the apparent political bias correlated almost perfectly with a simple measure of how much a model liked the first slot. In fact, the specific arrangement of options they had chosen for their first test turned out to be the single most misleading arrangement possible out of over 13,000 different ways they could have ordered the answers.

The implications of this discovery are significant for anyone who relies on these machines for advice. The researchers found that if you keep the question and the answers exactly the same but just shuffle the order of the choices, the model changes its top recommendation 71% of the time. This is far higher than the normal rate of random error, which was only about 33%. Even more striking, the model's confidence in its answer barely changed at all. When a model switched its recommendation because the options were moved, it reported being just as sure of its new answer as it was of the old one. This suggests that the model is not weighing the arguments; it is following a mechanical rule based on position.

The study also showed that this problem is not limited to a specific type of question. The models were just as likely to change their minds on difficult, contested policy questions as they were on easier ones. In fact, the order effect was strongest exactly where it mattered most: on the questions where the options were hard to tell apart. This is a behavior known as "satisficing," where a decision-maker, faced with a difficult choice, grabs the easiest available shortcut to get an answer. For these machines, the shortcut was picking the first item on the list.

The researchers concluded that measuring the political leanings of artificial intelligence using fixed lists of options is fundamentally flawed. A fixed-order test does not reveal what a model thinks; it reveals how the model reacts to the layout of the screen. While a residual component of genuine content preference survives when the order is randomized, it is weaker and does not form the clean ideological typology suggested by fixed-order tests. The "ideology" found in previous studies may have been nothing more than a statistical artifact of how the questions were written. To get a true picture of what these models prefer, researchers must randomize the order of the answers every time they ask a question. Without this step, any claim about a machine's political stance is just a guess based on a lottery of design choices. The study serves as a warning that when we ask machines to make choices, we must be careful not to mistake the order of the words for the weight of the argument.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →