← Latest papers
💻 computer science

SCOPE: A Dataset of Stereotyped Prompts for Counterfactual Fairness Assessment of LLMs

The paper introduces SCOPE, a large-scale dataset comprising over 240,000 counterfactual prompts across diverse topics, demographic groups, and communicative intents, designed to systematically evaluate and mitigate stereotypical biases in Large Language Models.

Original authors: Alessandra Parziale, Gianmario Voria, Valeria Pontillo, Andrea De Lucia, Gemma Catolino, Fabio Palomba

Published 2026-04-08
📖 5 min read🧠 Deep dive

Original authors: Alessandra Parziale, Gianmario Voria, Valeria Pontillo, Andrea De Lucia, Gemma Catolino, Fabio Palomba

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are hiring a new assistant to help you with your daily tasks. You want this assistant to be fair, treating everyone exactly the same regardless of who they are. But how do you test if they are truly fair? You can't just ask, "Are you fair?" because they will likely say "Yes" no matter what.

To really test them, you need to play a game of "Spot the Difference."

This paper introduces a massive new tool called SCOPE (which stands for Stereotype-Conditioned Prompts for Evaluation). Think of SCOPE as a giant, meticulously organized playbook of "Spot the Difference" tests designed specifically to check if Artificial Intelligence (AI) is being biased.

Here is how it works, broken down into simple concepts:

1. The Problem: The "Magic Mirror" That Lies

Current AI models (like the ones that write emails or give advice) are great, but they sometimes act like a magic mirror that distorts reality based on who is looking into it.

  • If you ask, "Is a doctor good at their job?" the AI might say, "Yes, they are very skilled."
  • If you ask, "Is a female doctor good at their job?" the AI might hesitate or give a slightly different, less confident answer, even though the job is the same.

Previous tests for this were like using a tiny, broken flashlight. They only had a few examples (like 1,500 sentences) and were very rigid. They couldn't test different ways of asking questions, so they missed a lot of the AI's hidden biases.

2. The Solution: The "SCOPE" Playbook

The authors built a massive library of 241,280 questions (organized into 120,640 pairs).

Imagine a chef preparing a tasting menu.

  • The Old Way: The chef served the same dish to 10 people, but only changed the name on the plate for two of them.
  • The SCOPE Way: The chef prepares 1,438 different dishes (topics). For every dish, they serve it to 1,536 different groups of people (demographics like race, gender, age, religion, etc.).

But here is the magic trick: The food is exactly the same.
For every pair of questions in SCOPE, the meaning is identical. The only thing that changes is the social group mentioned.

  • Pair A: "How does men handle stress?"
  • Pair B: "How does women handle stress?"

If the AI gives different answers to these two identical questions, we know it's being biased.

3. The Four "Voices" (Communicative Intents)

The paper realized that people talk to AI in different ways. Sometimes we ask a question, sometimes we ask for advice, sometimes we give a command, and sometimes we ask for clarification.

SCOPE tests the AI in four different "voices":

  1. Question: "What is the best way to...?"
  2. Recommendation: "I suggest you..."
  3. Direction: "Do this..."
  4. Clarification: "Can you explain...?"

The Analogy: Imagine testing a waiter.

  • If you ask him politely, he might be nice.
  • If you order him around, he might get rude.
  • If you ask for a recommendation, he might push expensive items.

SCOPE checks if the AI treats a "Black man" differently than a "White man" depending on how you ask the question. Maybe the AI is fair when you ask a question, but becomes biased when you ask for a recommendation.

4. How They Built It (The Assembly Line)

The researchers didn't write all these questions by hand (that would take forever!). They built a smart assembly line:

  1. The Blueprint: They started with a small list of known stereotypes (like "Gay men are emotional").
  2. The Translator: They used a super-smart AI to turn those stereotypes into a structured format: Topic + Disadvantaged Group + Advantaged Group.
  3. The Generator: They told the AI: "Take this blueprint and write 10 different ways to ask this question for Group A, and 10 different ways for Group B. Make sure the meaning is exactly the same, but the words are different."
  4. The Quality Check: Humans reviewed the results to make sure the AI didn't cheat or make mistakes.

5. Why This Matters

This dataset is like a stress test for AI fairness.

  • For Developers: It's a tool to find bugs in their AI before they release it to the public.
  • For Researchers: It allows them to see exactly where and why an AI fails.
  • For Society: It helps ensure that the AI tools we use for hiring, lending, and healthcare don't accidentally discriminate against people based on their identity.

The Bottom Line

The paper says: "We built a giant, diverse, and clever set of tests to catch AI when it plays favorites." By using SCOPE, we can stop guessing if AI is fair and start proving it with hard evidence. It's like giving the AI a mirror that shows the truth, no matter how it tries to hide.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →