← Latest papers
💬 NLP

The Proxy Presumption: From Semantic Embeddings to Valid Social Measures

This paper addresses the "Proxy Presumption" in Computational Social Science by introducing the Construct Validity Protocol (CVP) and Counterfactual Neutralization to rigorously validate semantic embeddings, ensuring they measure social constructs rather than confounding attributes like topic or style.

Original authors: Baishi Li, Ta Yu, Kelvin J. L. Koa, Ke-Wei Huang

Published 2026-05-11
📖 5 min read🧠 Deep dive

Original authors: Baishi Li, Ta Yu, Kelvin J. L. Koa, Ke-Wei Huang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to measure the "creativity" of a new invention. You decide to use a high-tech ruler that measures the distance between the new invention and old ones. If the distance is huge, you declare the invention "highly creative."

This paper argues that in the world of Artificial Intelligence (AI), researchers are doing exactly this, but they are making a dangerous mistake. They are assuming that because two things are "far apart" in the AI's math, they must be "different" in the real world. The authors call this the "Proxy Presumption."

Here is a simple breakdown of their argument, their solution, and their findings.

1. The Problem: The "Confused Ruler"

The authors say that AI models (specifically the ones that turn text into numbers, called "embeddings") are like a kitchen blender.

  • The Target (C): This is what you actually want to measure, like "creativity," "bias," or "novelty."
  • The Noise (Z): This is everything else mixed in, like the topic of the text, the author's writing style, the length of the sentence, or the specific words used.

When you put a text into an AI blender, it spits out a single number (a vector). The problem is that this number is a smoothie of both the Target and the Noise.

The Analogy:
Imagine you want to measure how "spicy" a soup is. You use a thermometer. But the thermometer is also sensitive to how "hot" the soup is (temperature).

  • If the soup is hot and spicy, the thermometer reads high.
  • If the soup is hot but bland, the thermometer also reads high.

If you just look at the high reading and say, "Aha! This soup is very spicy!" you are wrong. The thermometer is actually measuring heat, not spiciness.

In AI, researchers often look at the "distance" between two texts and say, "This text is very creative!" But the paper argues that the AI might just be measuring that the text is about a different topic or written in a different style, not that it is actually more creative. They are mistaking the "heat" (noise) for the "spiciness" (the real concept).

2. The Mathematical Proof: Why You Can't Just "Guess"

The paper uses some heavy math to prove a simple point: You cannot separate the spice from the heat just by looking at the soup.

Unless you have a special tool to filter out the heat, you can never be sure if the high temperature on the thermometer is because of the chili peppers or just the stove. Similarly, without special steps, an AI cannot know if a text is "creative" or just "about a different topic." The math proves that the AI's "ruler" is fundamentally confused.

3. The Solution: The "Validity Protocol" (CVP)

To fix this, the authors propose a new set of rules called the Construct Validity Protocol (CVP). Think of this as a quality control checklist for anyone trying to measure social concepts with AI.

Instead of just saying, "Here is a score for creativity," researchers must now prove their ruler actually measures creativity and not just noise. The protocol has three main steps:

  • Step 1: Define the Recipe. Clearly explain what "creativity" means and what it doesn't mean.
  • Step 2: Clean the Ingredients. Before measuring, use tools to remove the "noise." For example, if you want to measure "political bias," you might use an AI to rewrite the text to remove the specific topic (like "taxes") so you are only measuring the attitude, not the subject.
  • Step 3: The "Validity Card." This is a report card that must be attached to every study. It includes:
    • Stability Test: If you change the font or the order of words slightly, does the score stay the same? (If not, the ruler is broken).
    • The "Different Topic" Test: Does the score change if you talk about a totally different subject? If the score changes just because the topic changed, your ruler is measuring the topic, not the concept.
    • The "Real World" Test: Does this score actually predict real-world outcomes?

4. The "Forensic" Investigation

The authors didn't just talk about theory; they acted like detectives. They looked at 17 recent, famous papers that claimed to measure things like "bias," "ideology," and "creativity" using AI.

What they found:

  • Most papers (10 out of 17) gave a good definition of what they were measuring.
  • Almost all papers showed some examples to prove it looked right.
  • BUT, almost NO papers did the hard work of proving their ruler wasn't just measuring the "noise."
    • Only 1 paper proved their measure was different from other similar measures.
    • Zero papers proved their measure wasn't just tracking the topic or the writing style.
    • Zero papers used the strict math required to separate the "spice" from the "heat."

The Verdict: The authors call this the "Jangle Fallacy." This is when two different studies use the same name (e.g., "Bias") but are actually measuring two completely different things because they didn't check if their rulers were clean. One study might be measuring "rude words," while another is measuring "political topics," but they both call it "Bias."

5. The Conclusion

The paper concludes that we cannot just trust AI to measure complex human ideas like "creativity" or "bias" by simply looking at the distance between words.

If we want AI to be a useful tool for social science, we need to stop treating the AI's math as a magic answer. We need to use the Validity Protocol to:

  1. Clean the data of distractions (like topic and style).
  2. Prove that the score actually changes when the concept changes.
  3. Prove that the score doesn't change when the distractions change.

Until researchers start doing this, the paper argues, we are just measuring the "heat" of the soup and pretending we know how "spicy" it is.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →