Artificial Intelligence in America and China: Investigating prompt response between Human rights and geopolitical questions
This study reveals that while American and Chinese large language models exhibit minimal bias when addressing questions about the United States, they display significant divergent patterns regarding China, with Claude consistently scoring lower, DeepSeek higher, and ChatGPT remaining relatively neutral compared to the group average.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the last few years, a new kind of computer program has become a common part of daily life. These programs, known as large language models, can read and write human language with surprising fluency. They are trained on vast amounts of text from the internet, learning to predict what words should come next in a sentence. Because they are so good at mimicking human conversation, many people now turn to them for answers about history, science, and current events. However, because these programs learn from the data they are fed, they can also absorb the opinions, prejudices, and blind spots hidden within that data. This raises a critical question for anyone who relies on these tools for information: do these digital assistants see the world the same way we do, or do they carry their own hidden biases?
A recent study set out to answer this by treating artificial intelligence like a test subject. The researcher, Zian Ding, wanted to see if different computer models would give different answers when asked about sensitive topics like human rights and international politics. To do this, the study focused on three of the most popular models available at the time: two from American companies and one from a Chinese company. The goal was not to have a long conversation with the machines, but to ask them a specific set of questions and see how they rated the situations on a simple scale. The researcher asked each model to rate sixteen different scenarios, covering ten questions about human rights and six about geopolitical competition between the United States and China. For each question, the models had to give a score from one to ten, where a low number represented a negative view and a high number represented a positive view.
The experiment was designed to be very careful. To ensure the results were fair, the researcher asked every question in a fresh chat session, making sure the computer did not remember previous answers or get confused by past conversations. The researcher also turned off a feature in the American models that allows them to remember details about a user, ensuring that every answer was based only on the specific question asked. By comparing the scores given by the three different models, the study looked for patterns. If all three models gave similar scores, it would suggest they share a common, neutral view. If they gave very different scores, it would suggest that each model has its own unique perspective, likely shaped by the different data it was trained on.
When the researcher asked the models about the United States, the results were largely consistent, though not without minor variations. All three computer programs gave very similar scores for most questions regarding American human rights and geopolitical standing. However, there were exceptions; for instance, the model Claude gave a score for the 6G communication question that was 1.89 points lower than the average, indicating a notable deviation from the other two models. Whether the topic was about racial justice or international competition, the American and Chinese models generally agreed with each other, though specific questions revealed small differences in their assessments. This suggests that when it comes to describing their own country or a close ally, these artificial intelligence systems tend to operate from a largely shared set of facts and values, likely because they were all trained on a large amount of Western data.
The story changed completely when the questions turned to China. Here, the three models diverged sharply, revealing distinct biases. The Chinese model, DeepSeek, consistently gave the highest scores, painting a very positive picture of China's human rights record and its chances in global competition. It often cited many different sources to support its high ratings, though it sometimes hesitated to answer certain sensitive questions at first. The American model from the American company, Claude, on the other hand, gave the lowest scores. It frequently refused to answer questions about minority rights, stating that a simple number could not capture the complexity of the issue. When it did answer, its scores were significantly lower than the average, suggesting a much more critical view of China.
The third model, an American program, sat right in the middle. Its scores were very close to the average of all three, suggesting it remained relatively neutral compared to the other two. This middle ground indicates that this specific model may have been trained on a more balanced mix of data, or perhaps it was updated to avoid taking strong sides on these particular issues. The study also checked if the American model's memory feature changed its answers, but found that turning it on or off made very little difference to the final scores.
The findings suggest that while these artificial intelligence tools are powerful, they are not neutral observers. They reflect the data they are built from and the rules their creators set for them. When asked about the United States, they generally agree, though with some specific exceptions. When asked about China, they tell three different stories. This matters because millions of people, including many teenagers, now use these tools to learn about the world. If a young person asks a computer about human rights in different countries, the answer they get depends entirely on which computer they ask. The study concludes that users need to be aware that these machines have their own perspectives, and that the information they provide is not always an objective truth, but rather a reflection of the biases built into their code.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.