← Latest papers
🤖 AI

IndicFairFace: Balanced Indian Face Dataset for Auditing and Mitigating Geographical Bias in Vision-Language Models

This paper introduces IndicFairFace, a novel and ethically sourced dataset of 14,400 images balanced across India's 28 states and 8 Union Territories, to quantify and mitigate geographical bias in Vision-Language Models while maintaining high retrieval accuracy.

Original authors: Aarish Shah Mohsin, Mohammed Tayyab Ilyas Khan, Mohammad Nadeem, Shahab Saquib Sohail, Erik Cambria, Jiechao Gao

Published 2026-02-16
📖 5 min read🧠 Deep dive

Original authors: Aarish Shah Mohsin, Mohammed Tayyab Ilyas Khan, Mohammad Nadeem, Shahab Saquib Sohail, Erik Cambria, Jiechao Gao

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a super-smart digital assistant, a "Vision-Language Model" (VLM), that has read almost everything on the internet and seen billions of photos. You ask it, "Show me a picture of an Indian person."

Ideally, this assistant should show you a beautiful, diverse collage: a farmer from the green fields of Punjab, a fisherman from the islands of Andaman, a student from the hills of Sikkim, and a shopkeeper from the bustling markets of Mumbai.

But in reality, the paper IndicFairFace reveals that this assistant is biased. When you ask it, it mostly shows you people from just a few places—like Rajasthan or Tamil Nadu—while completely ignoring the Northeast, the islands, and many other regions. It's as if the assistant thinks "Indian" only looks a certain way, based on the most popular photos it saw online, rather than the true, messy, wonderful diversity of the country.

Here is a simple breakdown of what the authors did to fix this:

1. The Problem: The "Monolithic" Mirror

Think of the current AI models as a funhouse mirror. Instead of reflecting the real India (which has 28 states and 8 Union Territories, each with unique faces, skin tones, and styles), the mirror squashes everyone into a single, narrow shape.

The researchers found that if you ask these AIs for "Indian faces," they pull from a tiny slice of the population. They over-represent big, urban, or tourist-heavy areas and completely erase the faces of people from the Northeast, the islands, and rural areas. This isn't just a technical glitch; it's a form of digital erasure that can cause real-world harm, like an AI hiring tool rejecting a qualified candidate just because their face doesn't match the AI's narrow idea of what an "Indian" looks like.

2. The Solution: Building a "Fairness Scale"

To fix a broken scale, you need a perfect set of weights to calibrate it. The authors created IndicFairFace, which is exactly that: a perfectly balanced dataset.

  • The Recipe: They gathered 14,400 photos of real people.
  • The Balance: They didn't just grab random photos. They ensured there were exactly the same number of men and women from every single state and territory in India.
  • The Source: They were careful to use only images that were legally free to use (like from Wikimedia Commons) or found through ethical web searches, ensuring they didn't steal anyone's privacy.

Think of this dataset as a perfectly balanced choir. Before, the choir was mostly tenors from one city. Now, every voice from every region of India has an equal seat in the choir.

3. The Fix: "Surgical" De-biasing

Once they had this balanced choir, they used it to "re-train" the AI's brain. But they didn't want to re-teach the whole AI from scratch (which would be slow and expensive). Instead, they used a clever technique called Iterative Nullspace Projection (INLP).

Here is an analogy for how this works:
Imagine the AI's brain is a giant library of ideas. In this library, there is a specific "corridor" where the AI keeps all its stereotypes about geography.

  • The Old Way: You might try to burn down the library to get rid of the bad books. But then you lose all the good books too (the AI stops working well).
  • The New Way (INLP): The researchers found the specific "corridor" where the bias lived. They then built a forcefield around that corridor. When the AI tries to walk down that corridor to make a biased decision, the forcefield gently pushes it back onto the main path.
  • The Result: The AI still knows everything about cars, animals, and math (its general intelligence is untouched), but it can no longer "walk down the bias corridor." It now sees "Indian" as a diverse concept again.

4. The Proof: Does it still work?

A common fear with fixing AI bias is that you might break the AI. The authors tested their "de-biased" models on dozens of other tasks (like identifying cars, animals, or reading text in images).

The result? The AI didn't lose its smarts.

  • Its accuracy on other tasks dropped by less than 1.5% (which is basically nothing).
  • But when asked about "Indian people," it suddenly started showing faces from all over the country, not just the usual suspects.

Why This Matters

This paper is a blueprint for making AI fairer, not just for India, but for the whole world. It shows that we can't just say "AI is neutral" because it was trained on the internet. The internet is full of imbalances. To fix the AI, we need to build our own balanced datasets that represent the real world, not just the popular parts of it.

In short: The authors built a balanced photo album of India, used it to gently nudge the AI's brain out of its narrow thinking, and proved that the AI can be fair without losing its intelligence. It's a step toward ensuring that when AI looks at the world, it sees everyone.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →