← Latest papers
💬 NLP

Digital Skin, Digital Bias: Uncovering Tone-Based Biases in LLMs and Emoji Embeddings

This paper presents the first large-scale comparative study revealing that while modern Large Language Models robustly support skin-toned emojis, specialized emoji embedding models exhibit severe deficiencies and systemic biases in sentiment and semantic consistency across different skin tones, highlighting an urgent need for developers to audit and mitigate these representational harms.

Original authors: Mingchen Li, Wajdi Aljedaani, Yingjie Liu, Navyasri Meka, Xuan Lu, Xinyue Ye, Junhua Ding, Yunhe Feng

Published 2026-04-09
📖 5 min read🧠 Deep dive

Original authors: Mingchen Li, Wajdi Aljedaani, Yingjie Liu, Navyasri Meka, Xuan Lu, Xinyue Ye, Junhua Ding, Yunhe Feng

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the internet as a giant, bustling digital city. In this city, people use emojis like little digital stickers to express themselves. For years, the default "stick figure" emoji was just a generic yellow face or hand. But in 2015, the city planners (the Unicode Consortium) decided to add skin tone modifiers. Now, you could change that yellow hand into a light, medium, or dark hand to match your own identity. It was a huge step toward making the digital city feel like home for everyone.

However, this paper asks a scary question: Just because the city has all these different skin tones, does the city's "brain" (the AI) treat them all fairly?

The authors of this paper decided to act like digital detectives. They went into the "brains" of two types of AI systems to see if they were secretly biased against certain skin tones.

The Two Types of AI Brains

  1. The Old Librarians (Static Models): These are like old, dusty libraries with a fixed list of books. If a book isn't on the shelf, they don't know it exists. The researchers found that these old systems were terrible at handling skin tones. Most of them simply didn't have the "dark skin" books on their shelves, or they had very few. They were essentially ignoring a huge part of the population.
  2. The New Super-Readers (LLMs): These are the modern, super-smart AI models (like the ones you might chat with today). They are like a massive, infinite library that can read anything. The researchers found that these new models can see and understand every single skin tone. They don't miss any. But here's the twist: Just because they can see the dark skin tones doesn't mean they respect them equally.

The Hidden Biases Discovered

The researchers ran a series of tests to see how these "Super-Readers" actually thought about different skin tones. Here is what they found, using some simple analogies:

1. The "Heavy Backpack" Problem (Tokenization Bias)

Imagine you are walking into a library. If you are wearing a light-colored shirt, the librarian lets you in with just a small ID card. But if you are wearing a dark-colored shirt, the librarian makes you fill out a 5-page form and carry a heavy backpack just to get in.

The paper found that some AI models (specifically one called Mistral) treat dark skin tones this way. To the computer, representing a dark skin tone emoji takes more "tokens" (pieces of data) than a light skin tone.

  • Why it matters: In the digital world, more tokens mean it costs more money and takes more time to process. So, simply by being a person with a darker skin tone, the AI is making your digital identity "heavier" and more expensive to process.

2. The "Drifting Friends" Problem (Semantic Drift)

Imagine you have a group of friends who are all identical twins, except they are wearing different colored hats. You would expect them to sit together in a circle, all equally close to each other.

The researchers found that in some AI models, the "hats" (skin tones) actually push the friends apart.

  • In some models, the dark-skinned versions of an emoji drifted far away from the "default" yellow version in the AI's mind.
  • In others, the light-skinned versions drifted away.
  • The result: The AI doesn't see a "hand" that happens to have different skin colors; it sees "Hand A" and "Hand B" as completely different concepts. This breaks the idea that skin tone is just a small detail; the AI treats it as a massive, defining difference.

3. The "Good vs. Bad" Stereotype (Sentiment Bias)

This is the most concerning part. The researchers asked the AI: "Which skin tones do you think are 'Good' and which are 'Bad'?"

They used a test where they asked the AI to associate skin tones with positive words (like "joy," "love," "success") and negative words (like "anger," "sadness," "failure").

  • The Finding: Some models consistently linked lighter skin tones with "Good" things and darker skin tones with "Bad" things.
  • The Metaphor: It's like a teacher who subconsciously gives an 'A' to the student in the white shirt and a 'C' to the student in the black shirt, even though they both raised their hands to answer the same question. The AI is reinforcing old, harmful societal stereotypes, even though it was just trained on text.

The Big Takeaway

The paper concludes with a powerful message: Support is not the same as Fairness.

Just because a technology can display a dark skin tone emoji doesn't mean it treats that emoji with the same dignity, cost, or positive feeling as a light skin tone emoji.

The Analogy of the Digital Mirror:
Think of the internet as a giant mirror. For a long time, that mirror was cracked and only showed a few types of people. Now, we have fixed the cracks and added more people to the reflection. But the mirror is still warped. It reflects some people as "heavy," "expensive," or "negative," while others are "light," "cheap," and "positive."

The authors are calling on the builders of these AI systems to stop just adding the features and start fixing the warp. They need to audit their code, ensure that a dark skin tone doesn't cost more to process, and train their models to see all skin tones as equally human, equally valuable, and equally "good."

In short: We built a digital world that claims to be for everyone, but we need to make sure the AI running it doesn't secretly hold a grudge against anyone based on their skin color.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →