← Latest papers
🤖 machine learning

Position: Every Ground Truth is a Human Construction, not an Objective Truth

This position paper argues that ground truth datasets in machine learning are not objective facts but human-constructed artifacts, and advocates for acknowledging their contingent nature to improve model reliability, transparency, and accountability through "situated reliability."

Original authors: Charlotte Högberg, Ericka Johnson, Kiri L. Wagstaff

Published 2026-07-14
📖 5 min read🧠 Deep dive

Original authors: Charlotte Högberg, Ericka Johnson, Kiri L. Wagstaff

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot to recognize a "cat." You show it thousands of pictures, and for each one, you say, "Yes, that's a cat," or "No, that's a dog." In the world of machine learning, those labels you give are called Ground Truth.

For a long time, scientists have treated these labels like a magic mirror: they thought the "Ground Truth" was just a perfect, objective reflection of reality that was waiting to be discovered, like finding a hidden treasure map.

But this paper argues that the treasure map isn't hidden; it's drawn by us.

The authors, Charlotte, Ericka, and Kiri, are saying that "Ground Truth" isn't a natural fact that falls from the sky. It is a human construction. It's more like a recipe than a rock. Just as a chef decides how much salt to add, when to stir, and which ingredients to use, humans and machines make a huge series of choices to create the "truth" that trains AI.

The "Recipe" Analogy

Think of building a Ground Truth dataset like making a giant, complex smoothie to test a new blender.

  • The Fruit: You have to decide which fruits to buy. Do you use only red apples, or do you include green ones too?
  • The Cutter: Who slices the fruit? A professional chef? A tired intern? A robot arm?
  • The Blender: How fast do you spin the blades?
  • The Taste Test: Who decides if the smoothie is "good"?

The paper points out that if you change any of these steps, you get a different smoothie. If you change the "truth" (the smoothie recipe), the blender (the AI model) learns something completely different.

Why This Matters (The "Pixel" Problem)

The authors give a real example to show why this matters. There is a famous dataset called MNIST used to teach computers to recognize handwritten numbers. Back in the 1990s, computers were slow and had less memory, so the people who made the dataset decided to shrink the images from 128x128 pixels down to 28x28 pixels.

Because of that one choice made decades ago, thousands of AI models today are trained only to see 28x28 pixel images. If you try to show them a high-definition photo of a number taken today, they can't read it! They are stuck in the past because of a decision made to fit old computers. The "truth" they learned was limited by the tools available at the time, not by the actual numbers themselves.

The "Expert" Myth

Sometimes, we think experts are like Oracle Gods who always know the absolute truth. Or we think they are like Robotic Sensors that just record facts without feeling.
The paper says: Nope. Even experts make mistakes, disagree with each other, and see things differently.

  • In medicine, two doctors might look at the same X-ray and one says "healthy" while the other says "sick."
  • In a crowd-sourced project (like asking thousands of people on the internet to label photos), one person might think a picture is "happy" and another might think it's "sad."

If we pretend the expert is a perfect robot, we miss the fact that their "truth" is actually a human opinion shaped by their training, their mood, and the rules they were given.

The "Situated Reliability" Idea

The authors suggest we stop pretending our AI is universally perfect. Instead, we should talk about "Situated Reliability."

Think of it like a pair of sunglasses.

  • They work great in the bright desert sun (a specific context).
  • They are terrible in a dark cave (a different context).
  • They are useless underwater.

The sunglasses aren't "broken"; they just have a situated use. The paper argues that AI models are the same. A model trained on a specific "Ground Truth" might work perfectly in one hospital or one city, but fail miserably in another. We need to be honest about where and when our models work, rather than claiming they work everywhere.

What Should We Do?

The paper doesn't say "stop making AI." It says we need to be honest builders.

  1. Stop pretending the truth is magic. Admit that we made choices to create the data.
  2. Write down the recipe. When we share a dataset, we should explain exactly how we made it: Who labeled it? What tools did we use? What did we leave out?
  3. Try different recipes. Instead of just one "Ground Truth," maybe we should try making a few different versions to see how the AI changes.
  4. Team up. Machine learning experts need to talk to the people who actually know the subject (like doctors, biologists, or sociologists) to make sure the "truth" makes sense in the real world.

The Bottom Line

The paper suggests that "Ground Truth" is not a fixed, unchangeable fact of the universe. It is a human-made tool, shaped by our choices, our tools, and our limits. By admitting this, we can build better, safer, and more honest AI that knows its own limits, rather than pretending to be a god that sees everything perfectly.

The authors aren't saying we have solved this problem yet. They are saying we need to start talking about it so we don't accidentally build robots that are confident but wrong, or models that work great in a lab but fail in the real world. It's about being a little more humble and a lot more careful about how we teach our machines to see the world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →