Institutional Prestige as Geographic Bias in Large Language Models: Evidence from Three Factorial Experiments with Bootstrap Confidence Intervals
Through three factorial experiments involving 4,320 API calls across four large language models, this study demonstrates that institutional prestige and publication venue exert significantly stronger discriminatory effects on candidate evaluations than applicant name ethnicity or geographic origin, with a "rescue effect" showing that high-impact publications can disproportionately compensate for low institutional prestige.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the modern world, artificial intelligence has quietly become a gatekeeper. These computer programs, known as large language models, are increasingly asked to make decisions that shape human lives. They help decide who gets a job, who receives a research grant, and who qualifies for a loan. They do this by reading applications and assigning scores, acting as automated judges. For years, scientists have worried that these digital judges might carry the same prejudices as the humans who built them. We know they can be biased against people based on their names, their gender, or their race. But a newer, more subtle question has emerged: do these models also favor candidates simply because of where they went to school? This is not about the quality of the work a person has done, but about the reputation of the institution on their resume. If a computer program gives a higher score to a scientist from a famous university than to an equally qualified scientist from a less famous one, it is not being neutral; it is reinforcing existing hierarchies.
A team of researchers set out to test this specific type of bias using four different artificial intelligence models. They wanted to know if the models were reacting to the prestige of a university or simply to the country where that university was located. To find the answer, they created thousands of fake job and grant applications. They kept the candidate's skills and experience exactly the same in every single scenario. The only things they changed were the name of the university the candidate attended and, in some cases, the name of the journal where the candidate had published their work. They tested candidates from top-tier American universities, from respected institutions in Chile and Colombia, and from a university in Ecuador that does not appear on global ranking lists. They also varied the names on the applications to include Anglo, Latino, and Arabic origins to see if the models still held onto old stereotypes.
The results were clear and statistically robust. The artificial intelligence models consistently gave higher scores to candidates from prestigious universities. This gap was not a small fluctuation; it was a steady, measurable trend. Across the different models and professional fields, a candidate from a top-ranked university received a score that was nearly three-tenths of a point higher than a candidate from an unranked university. In the world of automated scoring, this difference is significant. However, when the researchers looked at the names on the applications, the story changed. The models did not show a statistically significant preference for candidates with Anglo names over those with Latino or Arabic names. The scores for these different name groups were so close that the difference could be attributed to random chance. This suggests that while the models have been successfully trained to ignore ethnic names, they have not been trained to ignore the status of the school on a resume.
To understand if this bias came from the school's reputation or the country it was in, the researchers ran a second set of experiments. They compared a highly ranked university in Mexico with a low-ranked, unranked university in the United States. If the models were simply biased against developing nations, they should have preferred the American school. Instead, the models tended to favor the Mexican university, even though it was less famous than the top American schools but more famous than the small American one. This proved that the models were reacting to the prestige of the institution itself, not just the geography. The bias was about the brand of the university, not the passport of the applicant.
The most surprising discovery came when the researchers introduced a third variable: the journal where the candidate had published their research. They compared candidates who had published in the world's most famous scientific journal against those who had published in a smaller, open-access journal. The difference in scores was massive. The prestige of the journal mattered far more than the prestige of the university. A candidate from a less famous university who had published in the top journal received a much higher score than a candidate from a top university who had published in the smaller journal. In fact, the boost from a famous journal was nearly six times stronger than the boost from a famous university. This created a "rescue effect," where publishing in a top journal completely compensated for the lack of institutional prestige. The models seemed to view a top-tier publication as such a strong signal of quality that it overrode the bias against the school.
The researchers also looked at how consistent the models were in their scoring. They found that when evaluating candidates from less famous institutions, the models were more inconsistent. One time, the model might give a low score to a candidate from a smaller university, and the next time, it might give a slightly higher score for the exact same profile. This inconsistency did not happen as often with candidates from top universities. This suggests that the models are less certain about how to judge people from institutions they know less about, leading to a double disadvantage: lower average scores and more unpredictable evaluations.
Ultimately, this study reveals that while artificial intelligence has made progress in avoiding some forms of discrimination, it has inherited a deep-seated bias toward status. The models act as amplifiers of existing social hierarchies, rewarding the familiar and the prestigious while penalizing the unknown. They do not seem to care about the country of origin as much as they care about the brand name of the school. However, the study also offers a path forward. It shows that high-quality work, when published in the right places, can break through these barriers. For researchers from less famous institutions, the path to fair evaluation by these digital judges lies not in changing their names or their countries, but in the undeniable weight of their published work. The machines are learning to value the work, but only when it is presented on the most prestigious stage.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.