Who Should Be Generated? Justifying Demographic Targets in Open-Ended Generation
This paper proposes a formal framework for justifying demographic targets in open-ended generation fairness evaluation, arguing that target construction is an integral component of fairness assessment rather than a preliminary step, and demonstrates through empirical analysis that different target justifications (such as geographic versus occupational priors) lead to significantly divergent fairness measurements.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a judge in a talent show, but instead of watching singers, you are watching AI robots write stories. The AI's job is to invent characters, like a "CEO in the United States." Now, here is the tricky part: the judge needs a scorecard to decide if the AI is being fair. But what does "fair" even mean in this imaginary world?
In the real world, fairness often means treating people equally. But when an AI makes up a character from scratch, there is no real person to compare it to. So, the judge has to pick a "target" to aim for. Should the AI try to match the real number of female CEOs in the US today? Should it try to match the total number of women living in the US? Or should it just flip a coin and make half the CEOs women and half men, regardless of reality? This is the big puzzle the paper tackles: Who should be generated? Before we can say if an AI is biased, we first have to agree on what the "perfect" list of characters should look like. Without a justified target, any score we give is just a guess.
The Missing Target Problem
The authors of this paper, a team from Sun Yat-sen University, noticed something strange happening in the world of AI fairness. They asked six of the smartest AI models to create 30 different stories about a "CEO in the United States." The results were wild: some models made 100% of the CEOs women, while others made 77%.
Now, is that a success or a failure?
- If you compare it to real-life US CEOs (where only about 33% are women), the AI looks like it's over-correcting and being unfair to men.
- If you compare it to the total US population (where about 50% are women), the AI looks like it's being fair.
- If you compare it to a perfectly equal coin flip (50/50), it looks fair.
The problem is that the AI prompts didn't say "make the CEO a woman" or "make the CEO a man." They just said "make a CEO." The AI filled in the blanks on its own. The researchers call this the "Missing-Target Problem." We have the AI's output, but we don't have a justified reason to pick which target distribution (real jobs, real population, or pure equality) we should use to judge it. Most fairness tests just pick a target and hope for the best, but this paper argues that picking the target is actually the most important part of the test.
The Four-Step Recipe for a Fair Target
To solve this, the authors built a new framework, like a recipe for baking a fairness cake. They say you can't just grab a target off the shelf; you have to justify it step-by-step. You need to make four specific commitments:
- What are we judging? (The Object): Are we judging the AI's ability to predict real people, or its ability to invent a fair representation of a social world? The paper decides these are invented characters for a public story, not real job applicants.
- Who counts? (The Prior): Which group of people should the AI represent? The authors argue that for a story set in the US, the AI should represent the people who live there (geographic membership), not just the people who currently hold the job (occupational incumbency). Why? Because the current job holders might have gotten there due to past unfairness, and copying that might just repeat the mistake.
- How do we split the pie? (Allocation): If we agree to represent the people living in the US, how do we divide the "CEO" slots among them? The authors argue for equal-person allocation. This means every person living in the US gets an equal "vote" for who gets to be a CEO in the story. It doesn't mean every group gets an equal number of CEOs (like 50% women), but that every individual has an equal chance.
- How do we measure it? (Operationalization): Finally, we need a real-world number to compare against. The authors use census data (resident population numbers) to create their target.
The Big Discovery: The Target Changes Everything
The researchers tested this new recipe using a massive benchmark called AP-Bench. They generated over 51,000 characters across six different AI models, 12 different countries, and six different jobs (like doctors, nurses, and CEOs).
Here is what they found:
- The Gap is Huge: When they compared the AI's output to their new "geographic" target (based on who lives in the country), the AI was way off. The difference score (called JSD2) ranged from 0.508 to 0.606 on a scale of 0 to 1. That is a massive gap, meaning the AI's characters looked very different from the population they were supposed to represent.
- The Target Matters More Than You Think: This is the most playful and surprising part. The researchers took the exact same AI stories and swapped the target. Instead of comparing them to the "people who live here" target, they compared them to a "equal categories" target (just making sure every group has the same number of CEOs).
- When they did this swap, the scores changed dramatically. The difference in scores for each model ranged from 0.279 to 0.355.
- This means that simply changing what you compare the AI to changes the verdict on whether the AI is fair or not. A model that looks "bad" against one target might look "okay" against another.
What This Means for You
The paper doesn't say "AI is broken" or "AI is fixed." Instead, it says, "Stop guessing what the target is."
The authors argue that we can't just say, "The AI should look like the real workforce," because that might just copy past discrimination. We also can't just say, "The AI should be 50/50," because that might ignore the actual population size.
The main takeaway is that fairness evaluation is not just about measuring the AI; it's about arguing for the standard we use to measure it. Before we can say an AI is biased, we have to explicitly state: "We are comparing this AI to the population of residents, and we believe every resident deserves an equal chance to be represented."
If we don't make that argument clear, our fairness scores are just numbers without a home. The paper provides a framework to build that home, ensuring that when we ask, "Who should be generated?", we have a solid, defensible answer before we even start grading the AI.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.