Who Evaluates AI's Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations
This paper analyzes 186 first-party and 248 third-party evaluation sources to reveal a critical governance gap where developers provide sparse, superficial social impact reporting while independent evaluators offer broader but incomplete assessments, ultimately calling for policies that mandate transparency and strengthen independent evaluation ecosystems.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the world of Artificial Intelligence (AI) as a massive, bustling kitchen where chefs (developers) are creating new, super-powered recipes (AI models) every day. These recipes are so powerful they can write stories, solve math problems, and even generate images. But before these recipes go out to the public, we need to know: Are they safe? Do they taste good to everyone, or do they make some people sick? Do they use too much electricity?
This paper, titled "Who Evaluates AI's Social Impacts?", is like a group of food critics and health inspectors trying to figure out who is actually checking these recipes for safety and fairness. They looked at 186 official reports written by the chefs themselves (First-Party) and 248 reports written by outside critics and independent testers (Third-Party).
Here is what they found, broken down into simple terms:
1. The Chefs Are Bad at Writing Their Own Health Reports
When a chef releases a new dish, they usually hand out a "Model Card" (like a nutrition label). The researchers found that these labels are often vague, incomplete, or missing entirely.
- The Problem: The chefs are great at saying, "This dish cooks fast!" (Technical performance), but they are very quiet about, "This dish might be spicy for some people" (Bias) or "This dish uses a lot of water to make" (Environmental cost).
- The Trend: In the past, chefs were a bit more open about these issues. But recently, the reports on things like environmental costs and bias have actually gotten worse. The chefs seem to be downplaying these topics, perhaps because they don't want to look bad or because they think it's too complicated to explain.
2. The Independent Critics Are Doing the Heavy Lifting
Since the chefs aren't writing good reports, the outside critics (Third-Party evaluators) have stepped in.
- The Good News: These independent testers are doing a much better job. They are finding the "spicy" parts (bias), checking if the food is safe for everyone (harmful content), and measuring how the dish performs for different groups of people.
- The Limit: The critics can only taste the food; they can't see the kitchen. They don't know exactly how much electricity the oven used, how much the ingredients cost, or how the kitchen staff (the humans curating the data) were treated. Only the chefs know those secrets.
3. The "Missing" Ingredients
The researchers looked at seven specific areas of "social impact" (how the AI affects society). Here is the status of each:
- Sensitive Content (Toxicity): This is the most reported area. Everyone is worried about the AI saying mean things, so chefs and critics both talk about it a lot.
- Bias & Fairness: There is some reporting here, but it's getting worse over time. Chefs are starting to say, "It's too hard to measure," or "It's too political right now."
- Environmental Costs: This is a huge gap. Chefs used to report on how much energy their models used, but that reporting has dropped sharply. It's like a restaurant refusing to tell you how much water it wastes.
- Labor & Human Rights: This is the biggest blind spot. Almost no one reports on the people who label the data or moderate the content. It's as if the restaurant never mentions the conditions of the dishwashers.
- Privacy & Cost: These are rarely reported in detail.
4. Why Are Chefs Hiding the Truth?
The researchers interviewed people inside the companies to understand why. They found two main reasons:
- Fear: If a chef admits their dish has a flaw, it might hurt their reputation or get them in trouble with the law. So, they stay silent.
- Incentives: Chefs only do the hard work of testing if it helps them sell the dish or if the government forces them to. If there's no profit in being transparent, they don't do it.
5. The "Famous Chef" Bias
The study also noticed that the outside critics mostly focus on the most famous chefs (like those in the US and China) because their dishes are popular.
- The Result: Smaller chefs or those in less famous regions get almost no attention. Their "recipes" might be just as dangerous, but no one is checking them because they aren't trending.
The Bottom Line
The paper argues that we cannot trust AI to be safe just by reading the labels the companies write themselves.
- The Gap: There is a massive hole in our knowledge about how AI affects the environment, the people who build it, and the privacy of users.
- The Solution: We need new rules (policies) that force chefs to be honest. We also need to build a better system where independent critics can get the data they need to do their jobs, and we need to protect those who speak up so they aren't punished for it.
In short: The chefs are hiding the messy parts of the kitchen, and the independent inspectors are doing their best to peek inside, but they can't see everything without the chefs' help.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.