← Latest papers
📈 economics

Measuring What Cannot Be Surveyed: LLMs as Instruments for Latent Cognitive Variables in Labor Economics

This paper establishes a theoretical framework for using Large Language Models as valid measurement instruments for latent cognitive variables in labor economics, demonstrating through the construction and validation of the Augmented Human Capital Index that LLMs can reliably quantify occupational task content to reveal distinct dimensions of AI exposure and correct for measurement error in economic estimation.

Original authors: Cristian Espinal Maya

Published 2026-04-06
📖 5 min read🧠 Deep dive

Original authors: Cristian Espinal Maya

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to measure the "secret sauce" of a job.

In economics, we know that some jobs are just about following a recipe (like a factory worker tightening bolts), while others require a chef's intuition (like a nurse diagnosing a complex illness or a lawyer crafting a unique argument). We call this "latent cognitive content." It's real, it matters, but you can't see it, weigh it, or ask a worker to fill out a survey about it easily. It's like trying to measure the flavor of a soup without tasting it.

For decades, economists have tried to measure this "flavor" using two main tools:

  1. Expert Panels: Hiring a bunch of expensive, slow human experts to read job descriptions and guess the "flavor." (Slow, costly, and humans disagree a lot).
  2. Keyword Counting: Looking for specific words like "computer" or "math." (Too simple; it misses the nuance).

The Big Idea: Using AI as a "Super-Taster"
This paper proposes a radical new tool: Large Language Models (LLMs). Think of an LLM not as a chatbot, but as a super-smart, tireless, and incredibly well-read "taster" that has read every job description, textbook, and manual ever written.

The author, Cristian Espinal Maya, asks: Can we trust this AI "taster" to give us accurate measurements for economic research?

To answer this, he treats the AI not just as a tool, but as a scientific instrument (like a thermometer or a scale). Just as a thermometer must be accurate and consistent, an AI scoring a job must meet four strict rules:

The Four Rules of the "AI Thermometer"

  1. Semantic Exogeneity (The "No Cheating" Rule):
    The AI must only look at the description of the job, not the result of the job.

    • Analogy: If you are measuring how "spicy" a dish is, you can't ask the AI, "How much money does the chef make?" The AI must judge the ingredients (the text), not the paycheck. If it looks at the paycheck, it's cheating.
  2. Construct Relevance (The "Does it Make Sense?" Rule):
    The AI's score must actually match what we think the job is.

    • Analogy: If the AI says a "Nurse" is 90% "Spicy" (augmentable by AI) and a "Data Entry Clerk" is 10%, that makes sense. If it got them backwards, the instrument is broken.
  3. Monotonicity (The "More is More" Rule):
    If a job gets a higher score, it must actually have more of that quality.

    • Analogy: If Job A is rated "Very Spicy" and Job B is "Mild," Job A must actually be spicier. The scale can't be random.
  4. Model Invariance (The "Agreement" Rule):
    If you ask two different AI models (like Claude Haiku and Claude Sonnet) to taste the same soup, they should agree on the ranking, even if they disagree on the exact number.

    • Analogy: One taster might say "This is a 7 out of 10" and another says "This is an 8 out of 10." As long as they both agree that this soup is spicier than the one rated a 3, the instrument is valid.

The Experiment: The "Augmented Human Capital Index"

The author put this theory to the test. He fed the AI 18,796 different job tasks (from the O*NET database) and asked it to score them on two things:

  • Augmentation: How much can AI help a human do this better? (The "Chef's Assistant" effect).
  • Substitution: How much can AI replace the human entirely? (The "Robot Replacement" effect).

The Results:

  • It Works: The AI scores matched up very well with other existing (but imperfect) measures.
  • It's Distinct: The AI successfully separated jobs that are augmented (helped) by AI from jobs that are substituted (replaced) by AI. Previous tools often mixed these up.
  • It's Reliable: When two different AI models scored the same jobs, they agreed 76% of the time. That's better than many human experts!
  • It's Cheap: Scoring 18,000 jobs took 6 hours and cost about $5. A human expert panel would take months and cost thousands.

The "Measurement Error" Fix

Here is the cleverest part. The author admits the AI isn't perfect. Sometimes Model A says "7" and Model B says "8." This is "noise" (measurement error).

In statistics, noise usually makes your results look weaker than they really are (a problem called "attenuation bias"). But because the author used two AI models, he could use a statistical trick called ORIV (Obviously Related Instrumental Variables).

  • Analogy: Imagine you are trying to measure the height of a building, but your tape measure is slightly stretched. If you use a second, slightly different tape measure, you can mathematically calculate exactly how much the first one was stretched and "fix" the number.
  • The Result: By using two AI models to cross-check each other, the author was able to "fix" the measurement error and get a truer picture of how AI affects wages.

Why This Matters

This paper is a game-changer for economics. It proves we can use AI to measure things that were previously unmeasurable, cheaply and quickly.

  • For Policymakers: We can finally see which jobs will be helped by AI (and need training) versus which will be replaced (and need safety nets).
  • For Researchers: We can study "cognitive skills" in a way we never could before.
  • For Everyone: It turns the "black box" of AI's impact on work into something we can actually see, measure, and understand.

In short: The author built a new, super-fast, super-cheap "AI ruler" to measure the invisible parts of our jobs, proved it's accurate, and showed us how to use it to fix old mistakes in economic research.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →