The Costs of Pretending That There Are Data-Generating Probability Distributions in the Social World
This paper argues that the assumption of true data-generating probability distributions in the social world is a harmful fiction that obscures critical choices and goals in machine learning, proposing instead a framework focused directly on relevant populations that preserves classical learning theory while avoiding these misleading abstractions.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Core Idea: The "Magic Dice" Myth
Imagine you are playing a game with a die. If you roll it a thousand times, you can predict with high confidence that a "6" will come up about 16% of the time. In the world of gambling or physics, we assume there is a true, hidden rule (a probability distribution) that governs the die. The die doesn't care about you; it just follows the laws of physics.
Machine Learning (ML) researchers often treat people like that die.
They assume that if they collect enough data about people (their income, age, location), there is a hidden, perfect "rulebook" or "map" that dictates exactly how those people will behave. They believe their job is to find this map, which they call a Data-Generating Distribution. They think, "If we just get enough data, we will discover the true probability that a specific person will commit a crime or default on a loan."
The authors of this paper say: "Stop pretending."
They argue that in the social world (dealing with humans), there is no such thing as a hidden rulebook. Humans are not dice. We are complex, changing, and influenced by our environment. Assuming a perfect, pre-existing probability map exists is not just wrong; it's dangerous because it hides the fact that the people building these models are making huge, subjective choices.
Analogy 1: The "Recipe" vs. The "Ingredient List"
The Old Way (The Gen-D Framework):
Imagine you are trying to predict how a cake will taste. The old way of thinking says, "There is a perfect, magical recipe for 'Cake' that exists in the universe. Our job is to taste enough cakes to figure out that exact recipe. Once we find it, we can bake the perfect cake every time."
The New Way (The Authors' View):
The authors say, "There is no magical recipe. There is only the specific batch of ingredients you have right now and the specific oven you are using."
- If you change the flour brand, the cake changes.
- If you change the oven temperature, the cake changes.
- If you change the baker, the cake changes.
In Machine Learning, the "ingredients" are the data we choose to collect. The "baker" is the person designing the model. The "recipe" isn't a hidden truth waiting to be found; it is constructed by the people making the model.
Analogy 2: The "Weather Forecast" vs. The "Human Mood"
The Weather (Where the old idea works):
If you want to know if it will rain tomorrow, you look at atmospheric pressure. The atmosphere doesn't care about your feelings. It follows physical laws. You can say, "There is a 70% chance of rain," and that number is based on a stable physical reality.
The Human Mood (Where the old idea fails):
Now, imagine trying to predict if a specific person will be happy tomorrow.
- If you ask them today, they might say "Yes."
- If you ask them after a bad day at work, they might say "No."
- If you ask them after a promotion, they might say "Yes."
There is no single "True Probability of Happiness" for that person. The number changes based on the context, the time, and the questions you ask.
The paper argues that when we use AI to predict human behavior (like who will get a job or who might re-offend), we are trying to predict a mood, not the weather. But we are using math designed for weather. This leads us to believe our predictions are "objective facts" when they are actually just guesses based on a specific, arbitrary snapshot of data.
Why Does Pretending Matter? (The "Costs")
The authors say that pretending these "perfect maps" exist causes three big problems:
1. It Hides the Choices We Make
When we say, "The model found the true probability," it sounds like the computer did the work and we just watched.
- Reality: The computer didn't find a truth. The engineers chose which data to feed it. They chose which features to look at (e.g., "Should we look at zip code?").
- The Cost: By pretending the model is just "discovering" a truth, we stop asking, "Why did you choose these specific variables?" This lets bad or biased choices slide under the rug.
2. It Creates a False Sense of Objectivity
If you tell a judge, "The algorithm says this person has a 30% chance of re-offending," it sounds scientific and neutral.
- Reality: That 30% isn't a law of nature. It's a number calculated based on a specific group of people from a specific year. If you changed the year or the group slightly, the number might jump to 60% or drop to 10%.
- The Cost: People trust the number too much. They think the algorithm is "fair" because it's "math," but the math is built on shaky ground.
3. It Confuses "Generalizing" with "Guessing"
In school, we learn that if we see a pattern in a sample, we can apply it to the whole world.
- The Problem: In the social world, the "whole world" is constantly changing. A model trained on data from 2017 might fail completely in 2024 because society changed, not because the model was "wrong" mathematically.
- The Cost: We keep building models that work for yesterday's world but fail in today's world, blaming the "noise" in the data rather than realizing the "map" we are using is outdated.
The Solution: Stop Looking for the "Ghost"
The authors suggest we stop trying to find the "Ghost in the Machine" (the true distribution). Instead, we should:
- Admit we are building, not discovering: Acknowledge that every model is a construction, like a house built with specific bricks.
- Focus on the specific population: Instead of saying "This model works for everyone," say "This model works for this specific group of people in this specific context."
- Be honest about choices: When we build these systems, we need to openly discuss: "We chose to look at these variables because..." rather than "The data spoke for itself."
The Bottom Line
The paper is a warning:
We are treating human beings like dice rolls. We think there is a hidden, perfect number that describes a person's future. But there isn't. There is only the story we tell about them based on the data we chose to collect.
If we keep pretending there is a "True Probability," we risk building systems that feel objective and fair but are actually just reflecting the biases and choices of the people who built them. We need to stop looking for the magic map and start taking responsibility for the territory we are drawing.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.