SHARP: Social Harm Analysis via Risk Profiles for Measuring Inequities in Large Language Models
This paper introduces SHARP, a multidimensional and distribution-aware evaluation framework that moves beyond scalar averages to measure social harm in large language models using risk profiles and tail-sensitive metrics like CVaR, revealing significant hidden inequities and worst-case failures that traditional benchmarks obscure.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Core Problem: Why "Average" Scores Lie
Imagine you are buying a car. The salesperson tells you, "This car has an average speed of 60 mph." That sounds great, right?
But what if that "average" is hiding a secret? What if the car usually drives at 10 mph, but every now and then, for no reason, it suddenly accelerates to 200 mph and crashes? The average is still 60, but the worst-case scenario is terrifying.
This is exactly what the paper argues is happening with Large Language Models (LLMs). Current tests mostly look at the "average" performance of AI. They give a single number (a scalar score) to say how "safe" or "fair" an AI is. The authors say this is dangerous because it hides the rare, severe failures that could cause real harm in high-stakes situations like hiring, healthcare, or justice.
The Solution: SHARP (The "Weather Forecast" for AI)
The authors introduce a new framework called SHARP. Instead of asking, "How does this AI do on average?", SHARP asks, "What is the worst this AI could possibly do, and how often does it happen?"
Think of SHARP not as a report card with a single grade, but as a weather forecast for a storm.
- Old way: "The average wind speed today is 10 mph." (Safe, boring, misses the hurricane).
- SHARP way: "While the average is 10 mph, there is a 5% chance of a 100 mph gust that could knock down trees."
How SHARP Works: The Four "Danger Zones"
The paper breaks down "harm" into four specific categories, like four different sensors on a spaceship:
- Bias: Is the AI being prejudiced against certain groups?
- Fairness: Is it treating people unequally in its answers?
- Ethics: Is it saying things that are morally wrong?
- Epistemic Reliability: Is it lying or making things up (hallucinating)?
Instead of mixing these all into one big "safety" number, SHARP measures them separately. It's like checking the engine, the brakes, and the tires separately, rather than just saying "the car is 80% good."
The Secret Sauce: Looking at the "Tail"
The most important part of SHARP is a metric called CVaR95 (Conditional Value at Risk).
- The Analogy: Imagine a line of 100 people.
- Average Risk: You look at the height of everyone and find the average.
- SHARP (CVaR95): You ignore the first 95 people. You only look at the tallest 5 people (the "tail" of the distribution). Then, you calculate the average height of just those 5.
Why? Because in high-stakes AI, the "tallest 5" (the worst 5% of answers) are the ones that cause disasters. If an AI is usually polite but occasionally spews hate speech, the "average" looks fine, but the "tail" looks dangerous. SHARP focuses entirely on that dangerous tail.
What They Found: The "Hidden" Dangers
The authors tested 11 of the smartest AI models available. Here is what SHARP revealed that old tests missed:
- The "Twin" Illusion: Two AI models might have the exact same "average" safety score. But when you look at their worst 5% of answers, one might be slightly risky, while the other is three to four times worse. They look identical on a standard report card, but they are very different in a crisis.
- Different Flaws: Some models are great at being fair but terrible at not lying (hallucinating). Others are great at not lying but terrible at being biased. You can't just say "Model A is better than Model B." You have to say "Model A is safer for this specific risk, but Model B is safer for that one."
- Bias is the Worst: Across the board, the "tail" (the worst cases) was most often caused by Bias. When things went really wrong, it was usually because the AI was being prejudiced.
The Conclusion: Stop Relying on Averages
The paper concludes that we need to stop treating AI safety like a simple math problem with one answer.
- Old way: "This AI is 90% safe." (Too simple, hides the danger).
- SHARP way: "This AI is usually safe, but in the worst 5% of cases, it fails hard on bias and lies."
By using SHARP, regulators and companies can make better decisions. Instead of picking the AI with the highest "average" score, they can pick the AI that has the safest worst-case scenario. It's the difference between buying a car based on its top speed versus buying one based on how well its brakes work in an emergency.
In short: SHARP is a new tool that stops us from being fooled by "good averages" and forces us to look at the scary, rare mistakes that actually matter.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.