Effects of Answer Format Variation on Gender Bias in Large Language Models
This paper demonstrates that the format of answer choices (closed-ended, Likert-scaled, or open-ended) significantly alters the measurement of gender bias and distributional alignment in large language models, often reversing model rankings and necessitating multi-format evaluation designs for robust assessment.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the world of artificial intelligence, large language models are the digital engines that power everything from chatbots to search assistants. These systems are trained on vast amounts of text from the internet, learning to predict what words should come next. Because they learn from human writing, they inevitably absorb the patterns, prejudices, and stereotypes present in that data. One of the most persistent and harmful patterns is gender bias, where a model might automatically assume a doctor is a man or a nurse is a woman. To fix this, researchers need to measure how strong these biases are. They do this by asking the models questions and seeing how they answer. For decades, scientists have known that the way a question is asked changes the answer, even for people. If you ask someone to choose between two options, they might pick one; if you ask them to rate their agreement on a scale, they might give a different answer. This is a well-established fact in survey science, but for a long time, researchers testing artificial intelligence assumed the format of the question didn't matter much, or that the model's answer would remain the same regardless of how the question was presented.
A team of researchers at the University of Stuttgart decided to test this assumption directly. They wanted to see if the way they asked a question about gender would change the bias they measured in the machine. They took three different types of large language models and asked them the same set of questions about gender roles, but they changed the format of the answers every time. In one version, the model had to pick a single letter from a list of choices, like a multiple-choice test. In another, the model had to rate its agreement with a statement on a scale from one to ten, similar to how people answer opinion polls. In the third version, the model was allowed to write a free-form answer in its own words. The researchers used two different sets of questions: one set designed specifically to test bias with clear right and wrong answers, and another set taken from real public opinion surveys about how people view gender in society.
The results were striking and showed that the format of the question completely changed the story the researchers were told. When the models were forced to pick a single option from a list, they showed strong, consistent gender biases. They frequently chose the stereotypical answer, such as picking the man for a leadership role. However, when the researchers switched to the open-ended format, where the model could write its own response, the measured bias dropped dramatically. In many cases, the models simply refused to pick a side or stated that the information was insufficient to decide. This wasn't because the models suddenly became less biased; rather, the open format gave them a way to avoid the trap of the multiple-choice question. The most surprising finding was that the ranking of the models changed depending on the format. A model that appeared to be the most biased when forced to pick a letter became the least biased when allowed to write freely. Another model showed almost no bias in the multiple-choice test but revealed significant bias when asked to rate its agreement on a scale.
The researchers also looked at how well these models matched the opinions of real human survey respondents. They found that the format changed the alignment as well. In some cases, a model's answers looked very similar to human opinions when using one format, but looked completely different when using another. For instance, when models were asked to use a scale, their answers often canceled each other out, creating an average that looked neutral even though the models were expressing strong, opposing views at the extremes. This is different from how humans behave; people often use the middle of a scale to show they are undecided, but the models used the middle to balance out extreme stereotypes. The study suggests that the bias we see in these machines is not a fixed, unchangeable trait like a physical feature. Instead, it is a behavior that shifts depending on the constraints we place on it.
This work challenges the idea that a single test can tell us everything about a model's fairness. The researchers found that the choice of answer format is not just a minor detail of the experiment; it is a major factor that shapes the results. If a researcher only looks at multiple-choice tests, they might conclude that a model is deeply biased. If they only look at open-ended responses, they might conclude the model is perfectly fair. Both conclusions are incomplete. The study does not prove that the models are free of bias, nor does it prove that the bias is entirely an illusion. It shows that bias is a complex behavior that manifests differently under different conditions. The researchers argue that to truly understand and evaluate these systems, we must stop relying on a single way of asking questions. We need to look at how the models behave across different formats to get a full picture of their behavior, just as a doctor would not diagnose a patient based on a single symptom but would look at the whole picture of their health.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.