← Latest papers
📊 statistics

Causal Inference with Generative Artificial Intelligence: Application to Texts as Treatments

This paper proposes a Generative AI-Powered Inference (GPI) methodology that leverages large language models to generate treatments and utilize their internal representations for more accurate and efficient causal effect estimation from unstructured text, thereby eliminating the need to learn causal representations directly from data and overcoming common challenges like confounding and overlap violations.

Original authors: Kosuke Imai, Kentaro Nakamura

Published 2026-06-12
📖 4 min read☕ Coffee break read

Original authors: Kosuke Imai, Kentaro Nakamura

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to figure out if a specific detail in a story changes how people feel about a character. Let's say you want to know: Does having a military background make voters like a politician more?

The problem is that real-life stories are messy. A politician with a military background might also happen to be older, have a different education level, or write their biography in a more emotional tone. If you just compare two random biographies, you can't tell if the voters liked the candidate because of the military part or because of the education part. In science, we call these messy extra details "confounders."

Traditionally, researchers have tried to fix this by using computers to "read" the text and guess what the confounders are. But this is like trying to clean a muddy window by guessing where the dirt is; it's hard, slow, and often inaccurate.

This paper introduces a new tool called GPI (Generative-AI Powered Inference). Here is how it works, using a simple analogy:

The Magic Copy Machine (The GenAI)

Instead of just reading existing stories, the researchers use a "Magic Copy Machine" (a Large Language Model, or LLM) to write the stories for them.

  1. The Prompt: The researcher tells the machine: "Write a biography of a politician who has a military background." Then, they tell it: "Write a biography of a politician who does not have a military background."
  2. The Secret Blueprint: Here is the superpower. When this AI writes the story, it doesn't just spit out words; it creates a hidden, internal "blueprint" (a mathematical representation) of exactly what it wrote.
  3. The Trick: Because the AI wrote the story, the researchers have access to this perfect, hidden blueprint. They know exactly what the AI put into the text to make it about the military, and they know what it put in for everything else (like education or tone).

The "Deconfounder" (The Filter)

The researchers use this perfect blueprint to build a special filter called a Deconfounder.

  • Old Way: Imagine trying to separate red and blue marbles that are glued together. You have to guess how to pull them apart.
  • GPI Way: Because the AI wrote the story, the researchers have the "instruction manual." They can look at the blueprint and say, "Okay, this part of the blueprint is the 'Military' ingredient, and this other part is the 'Education' ingredient." They can mathematically isolate the military part without messing up the education part.

This allows them to ask: "If we keep the education and tone exactly the same, but only change the military part, how does the voter's score change?"

Why This is Better

The paper claims this method is like upgrading from a hand-cranked calculator to a supercomputer for two main reasons:

  1. Accuracy: Because they use the AI's true internal blueprint instead of guessing the text's meaning, they get a much clearer answer. In their tests, their method had less "noise" (error) and gave more reliable results than the best existing methods.
  2. Speed: The old methods are like trying to solve a giant puzzle by looking at every single piece one by one. The new method is like having the picture on the box; it solves the problem about 100 times faster.

The "Text Reuse" Twist

The researchers also found a cool shortcut. If you take an existing biography and ask the AI to "rewrite this exact same story," the AI creates a new, perfect blueprint for that old text. This means you don't even need to generate new stories from scratch; you can use old data, feed it to the AI, and get the same high-quality results.

The Bottom Line

The paper argues that by using Generative AI not just to generate text, but to understand the hidden structure of that text, we can finally untangle the messy web of cause and effect in social science.

  • The Goal: Measure the true effect of one specific thing (like military service) on an outcome (like voter happiness).
  • The Problem: Other things (confounders) are mixed in.
  • The Solution: Use an AI to generate or rewrite the text, grab its "secret blueprint," and use that to perfectly separate the cause from the noise.

The authors tested this on real voter surveys and found that, yes, military background does seem to make voters feel warmer toward candidates, and they were able to prove this with much more confidence and speed than before. They also note that this same logic could work for images and videos in the future, provided the AI can generate them with similar precision.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →