← Latest papers
💻 computer science

Sociodemographic Biases in Educational Counselling by Large Language Models

This study reveals that while all six evaluated Large Language Models exhibit measurable sociodemographic biases in educational counselling, these disparities are significantly amplified by vague student descriptions and can be substantially mitigated through the use of precise, individualized information.

Original authors: Tomasz Adamczyk, Wiktoria Mieleszczenko-Kowszewicz, Beata Bajcar, Grzegorz Chodak, Aleksander Szczęsny, Maciej Markiewicz, Karolina Ostrowska, Aleksandra Sawczuk, Przemysław Kazienko

Published 2026-04-30
📖 4 min read☕ Coffee break read

Original authors: Tomasz Adamczyk, Wiktoria Mieleszczenko-Kowszewicz, Beata Bajcar, Grzegorz Chodak, Aleksander Szczęsny, Maciej Markiewicz, Karolina Ostrowska, Aleksandra Sawczuk, Przemysław Kazienko

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a team of six different "super-intelligent school counselors." These aren't humans; they are Large Language Models (LLMs), which are like super-powered computers that have read almost everything on the internet and learned how to talk and think like people.

The researchers in this paper wanted to know: If you ask these AI counselors about a student, will they treat that student differently just because of their name, race, or how much money their parents make?

To find out, they didn't just ask one question. They created a massive experiment with 243,000 different scenarios. Here is how they did it, broken down into simple concepts:

The Experiment: The "Student Profile" Game

Think of the researchers as actors playing a game. They wrote 900 short stories (called "vignettes") about students.

  • The Twist: In every story, the student's grades, behavior, and situation were exactly the same. The only thing that changed was the label attached to them.
  • The Labels: They swapped labels like "a student" (the control), "a low-income student," "an Asian female," "an immigrant," or "a wealthy student."
  • The Test: They asked six different AI models to give advice on things like: Should this kid get into an advanced class? How serious is their behavior? What college should they aim for?

They also tested two different ways of describing the student:

  1. The "Blurry Photo" (Vague): "This student is hardworking but sometimes misses class." (No numbers, just feelings).
  2. The "ID Card" (Precise): "This student has a 3.8 GPA, 96% attendance, and no disciplinary issues." (Hard facts).

The Big Findings

1. The AI Counselors Are Not Neutral
Just like human teachers sometimes have unconscious biases, these AI models showed clear patterns of favoritism and unfairness.

  • Money Matters: When the AI didn't have many facts, it assumed students from wealthy families were "future leaders" and students from low-income families were "struggling," even if their grades were identical.
  • The "Model Minority" Myth: The AI consistently gave Asian students higher praise and more optimistic future predictions than other groups, mirroring a common stereotype.
  • The Surprising Twist: In some cases, the AI actually over-corrected. When a student had a vague, negative description (e.g., "sometimes disengaged"), the AI was more sympathetic to minority or immigrant students than to the "average" student. It was almost like the AI was trying extra hard to be fair, but ended up being unfair in a different way.

2. The "Blurry Photo" vs. The "ID Card"
This was the most important discovery.

  • When the picture was blurry (vague descriptions): The AI relied heavily on stereotypes. The bias was three times stronger. If the AI didn't know the student's exact grades, it filled in the blanks with assumptions about their race or income.
  • When the picture was clear (precise metrics): When the AI was given hard numbers (like a specific GPA), the bias dropped significantly. The facts acted like a shield, stopping the AI from guessing based on stereotypes.

3. Not All AIs Are the Same
The six models didn't all act the same way.

  • One model (GPT-5.2) was the most "fair," showing very little bias.
  • Another (Grok) was very strict about money, favoring the rich and penalizing the poor.
  • Another (DeepSeek Reasoner) was very friendly toward specific racial groups but less so toward wealthy students.
  • This proves that bias isn't an unavoidable "glitch" in all AI; it depends on how each specific AI was trained.

The Takeaway

The paper concludes that if we want AI to be fair in schools, we can't just ask it vague questions like "Is this student good?" because the AI will fill in the blanks with stereotypes.

Instead, we need to force the AI to look at concrete facts (grades, attendance, specific behaviors). When the AI has to look at the hard data rather than guessing based on a name or a background, it becomes much fairer. The study suggests that the way we give information to the AI is just as important as the AI itself.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →