ConsistencyAI: A Benchmark to Assess LLMs' Factual Consistency When Responding to Different Demographic Groups
The paper introduces ConsistencyAI, an independent benchmark that evaluates 19 large language models for factual consistency across diverse demographic personas, revealing that consistency scores vary significantly by both model provider and topic, with a mean benchmark threshold of 0.8656.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the modern information age, we increasingly turn to artificial intelligence to answer our questions, from the price of milk to the details of complex history. These systems, known as large language models, are designed to understand human language and generate text that feels like a conversation. They have become a primary way people access facts, often replacing traditional search engines for quick answers. However, a fundamental question remains about how these machines work: do they tell the same story to everyone, or do they change the facts depending on who is asking? If a computer program knows you are a young student, does it present a different set of truths than it would to an older executive? This possibility raises concerns about fairness and the integrity of information. If the facts themselves shift based on a person's background, such as their age, gender, or job, it suggests the system is not just answering a question but tailoring reality to fit a specific audience. This phenomenon, if real, could subtly shape how different groups of people understand the world, potentially deepening divides rather than bridging them.
To investigate this, researchers at Duke University and the University of North Carolina created a new test called ConsistencyAI. Their goal was to see if large language models tell different facts to different people. They did not ask the models to check if the information was true; instead, they asked whether the models gave the same set of facts to a 25-year-old teacher as they did to a 60-year-old construction worker. The team built a system that simulated a diverse group of people from the United States, representing a wide range of ages, jobs, and backgrounds. They then asked twenty-five different artificial intelligence models to provide five facts on fifteen different topics, such as housing costs, trade deficits, and political conflicts. Each question was asked to every model one hundred times, with the context of the person asking changing each time using a set of 100 personas sampled from a representative dataset. By comparing the answers, the researchers could measure how much the content changed when the persona changed.
The results revealed that while many models are generally consistent, they are not perfectly so. The researchers found that the models did indeed change the facts they presented based on who they thought they were talking to. In a large-scale test involving 100 personas, the consistency scores for the models ranged from a high of 0.9137 to a low of 0.6719, with an average score of 0.8555. This average became the benchmark for the study. The most consistent model was xAI's Grok-4, which stayed remarkably steady in its facts regardless of the user. At the other end of the spectrum, several lighter and mid-tier models showed the most variation. The study showed that the topic being discussed mattered just as much as the model itself. For instance, when asked about the U.S. trade deficit, a topic with clear, standardized government data, the models were highly consistent. However, when asked about housing costs, a topic with complex and varied data sources, the models gave much more different answers to different people.
The researchers also discovered that the models did not just change their tone or style; they actually selected different facts. When the persona changed, the specific details, organizations, and numbers mentioned in the response often shifted. For example, on the topic of border security, one model tended to give answers focused on law enforcement to older users, while younger users received answers focused on geography. Another model gave simplified, analogy-heavy responses to minors but direct, factual responses to adults. This suggests that the models are sensitive to demographic cues and adjust their information accordingly. The study also noted that being a more advanced or "reasoning" model did not guarantee better consistency. Some smaller, older models performed better than the newest, most complex systems, indicating that consistency is not simply a byproduct of a model's size or intelligence.
A significant portion of the research focused on controversial topics. The Israeli-Palestinian conflict showed the widest variation in answers, with some models giving very consistent facts and others giving wildly different ones. On this topic, some models even refused to answer or returned errors, particularly when the user was identified as a child. This suggests that models may have built-in filters that cause them to back away from sensitive subjects depending on the perceived user. The study also found that political sensitivity alone did not predict inconsistency. While some political topics were highly inconsistent, others were not, and even non-political topics like housing costs showed low consistency. This indicates that the difficulty of finding a single, stable narrative in the data is a major factor. If the information available on a topic is messy or contested, the model is more likely to present different facts to different people.
The authors emphasize that their work does not determine if the facts are true, but rather if the facts are stable. They argue that even if the information is correct, presenting different sets of facts to different people can be harmful. It can lead to confirmation bias, where people only hear the version of reality that fits their existing beliefs, or it can reinforce stereotypes. The study suggests that for artificial intelligence to be truly reliable, it must provide a stable set of facts regardless of who is asking. The researchers have released their testing tools and an interactive website so that journalists, developers, and the public can see for themselves how different models behave. They hope this transparency will encourage companies to build systems that are less likely to tailor the truth to the user, ensuring that the information we all rely on remains consistent and trustworthy.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.