← Latest papers
💻 computer science

Picturing Perceptions: An Open-Source Toolkit to Uncover Bias in Humans and Machines

This paper introduces PictoPercept, an open-source toolkit that uses visual forced-choice comparisons against real-world labor data to measure and reveal significant, often divergent, biases in both human judgment and AI systems regarding earnings across different demographic groups.

Original authors: Saurabh Khanna, Zhijun Chen, Chei Billedo, Jiayi Yan, Irene van Driel, Alex Barco Martelo, Hugo Moreda Cartagena, Haizea Gonzalez Atorrasagasti, Markel Adanez Perez, Daniela An, Qianyi Wang, Xinkangru
Published 2026-06-15
📖 5 min read🧠 Deep dive

Original authors: Saurabh Khanna, Zhijun Chen, Chei Billedo, Jiayi Yan, Irene van Driel, Alex Barco Martelo, Hugo Moreda Cartagena, Haizea Gonzalez Atorrasagasti, Markel Adanez Perez, Daniela An, Qianyi Wang, Xinkangrui Gao, Lauren Taylor, Olga Eisele, Sindy Sumter

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to guess who earns more money just by looking at their face. Would you guess correctly? Or would your brain rely on old stereotypes, like a broken GPS that keeps sending you down the wrong road?

This paper introduces a new tool called PictoPercept (short for "Picturing Perceptions") designed to test exactly that. It's like a "bias detector" that works for both humans and computers.

Here is the story of what the researchers did and what they found, explained simply:

The Problem: Old Tools Were Broken

For years, scientists tried to measure bias using surveys (asking people directly) or reaction-time tests (measuring how fast people click buttons).

  • Surveys are like asking someone, "Are you racist?" Most people will say "No" because they want to look good. It's like asking a thief if they steal; they'll lie.
  • Reaction tests are like a video game where you have to sort words quickly. But computers don't have "reaction times" in the same way humans do, so these tests couldn't be used to check if AI was biased.
  • The Missing Map: None of these old tools checked if people's guesses matched reality. They just measured if people preferred Group A over Group B, without asking, "Is Group A actually richer or poorer in the real world?"

The Solution: A Visual "Forced Choice" Game

The researchers built PictoPercept, an open-source toolkit (meaning anyone can use it for free).

  • The Game: You see two faces on a screen. The prompt asks, "Who is more likely to have higher earnings?" You have to pick one. You can't say "I don't know."
  • The Trick: The faces are carefully matched so they look equally attractive and trustworthy. The only difference is their race and gender.
  • The Map: The researchers compared your choices against real data from the U.S. Bureau of Labor Statistics. If the data says Asian men earn the most, but you keep picking Latino men as the highest earners, the tool calculates that as a "bias score."

They tested this on 283 real American adults and also fed the exact same pictures to GPT-5, a powerful AI model.

The Big Surprises (The Results)

1. The "Model Minority" Blind Spot
The biggest shock was about Asian Americans.

  • The Reality: According to government data, Asian Americans are actually the highest-earning group in the U.S.
  • The Perception: People dramatically underestimated their earnings. Even when shown faces of Asian men and women, people guessed they earned less than they actually do.
  • The Irony: Even Asian participants in the study underestimated their own group's earnings! It's as if the whole country has a blind spot to this group's success, perhaps because they aren't seen enough in leadership roles or media.

2. The "Ingroup" Myth
We often think people always favor their own group (e.g., White people favoring White people, Black people favoring Black people).

  • The Reality: This wasn't true for everyone.
    • White men did show favoritism, overestimating their own group's earnings.
    • Asian participants, however, did the opposite. They significantly underestimated their own group. They didn't just ignore the bias; they actively thought their own group earned less than they did.
    • Other groups (like Black and Latino participants) were surprisingly accurate about their own groups, showing no strong favoritism or bias.

3. The AI is Even More Biased
When they tested GPT-5 (the AI), the results were scary.

  • The AI didn't just mimic human bias; it amplified it.
  • The AI was extremely bad at judging women. It systematically thought all female groups earned much less than they actually do.
  • It thought all male groups earned much more.
  • While humans made mistakes, the AI's errors were huge, consistent, and extreme. It's like a student who not only gets the answer wrong but gets it wrong by a factor of ten.

Why This Matters

The paper argues that bias isn't just about "hating" a group; it's about having a distorted map of reality.

  • If you think a group earns less than they do, you might treat them unfairly in hiring or lending.
  • The tool shows that our brains (and our computers) are often running on outdated software that doesn't match the real world.

The Takeaway

PictoPercept is a new, open-source way to check if our perceptions match reality. It found that:

  1. We are blind to the high earnings of Asian Americans.
  2. We don't always favor our own groups; sometimes we underestimate ourselves.
  3. AI models like GPT-5 are currently much more biased than humans, especially against women.

The authors suggest that before we let AI make big decisions (like who gets a loan or a job), we need tools like this to check if the AI is looking at the world through a funhouse mirror.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →