← Latest papers
📈 economics

Compositional Synthetic Controls

This paper introduces a synthetic control estimator for compositional outcomes based on a random utility model and Aitchison geometry, which is applied to Pennsylvania's electricity mix to reveal a significant, persistent shift toward natural gas and away from renewables following the Alternative Energy Portfolio Standard.

Original authors: Onil Boussim

Published 2026-07-21
📖 4 min read☕ Coffee break read

Original authors: Onil Boussim

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery: "What would have happened if a specific event hadn't occurred?" This is the heart of a field called causal inference, where scientists try to figure out cause and effect. Usually, they look at a single city, state, or country that did something new (like passing a new law) and compare it to a "twin" made up of a mix of other places that didn't do that thing. This method, known as the "Synthetic Control," works like a recipe: you take a little bit of State A, a dash of State B, and a pinch of State C to create a perfect "Synthetic State" that looks exactly like the treated one before the event.

But there's a catch. Many real-world outcomes aren't just single numbers like "temperature" or "GDP." They are recipes themselves—vectors of shares that must always add up to 100%. Think of a pie chart: if the slice for "Apples" gets bigger, the slices for "Oranges" and "Bananas" must get smaller. This is called compositional data. The old way of comparing these pies used a ruler that measured straight-line distance, which is terrible at understanding how the ratios between slices change. It's like trying to judge the flavor of a soup by measuring the weight of the bowl instead of tasting the broth. This paper steps in to fix that ruler, creating a new way to compare these shifting pies that respects the math of how parts relate to the whole.

The paper, titled "Compositional Synthetic Controls" by Onil Boussim, introduces a clever new tool to solve this puzzle. The author argues that when we look at things like employment types, energy sources, or vote shares, we shouldn't just look at the raw numbers. Instead, we should look at the "log-odds," which is a fancy way of comparing how much more likely one choice is than another. By using a special kind of geometry (called Aitchison geometry) that understands these ratios, the new method builds a "Synthetic Twin" that is mathematically perfect. It ensures that if the twin's "Gas" slice goes up, the "Coal" and "Renewables" slices automatically adjust to keep the total at 100%, just like a real pie.

The paper finds that this new method is much better at spotting the truth than the old "straight-line" ruler. The author demonstrates this with a real-world case study: Pennsylvania's electricity grid after a 2004 law called the Alternative Energy Portfolio Standard. The old methods might have missed the full story, but this new tool reveals a massive, persistent shift. It shows that by 2022, Pennsylvania's natural gas usage was nearly 60 percentage points higher than it would have been without the law, while renewable energy, despite growing in absolute terms, actually lost ground relative to gas. The paper suggests that without this new way of looking at the data, we would have missed the fact that the policy's biggest impact was a huge swap from coal to gas, rather than a simple boost for renewables. The author is quite confident in these results, having tested the method against fake scenarios (placebos) to ensure the findings aren't just a fluke, though they note that the small number of states available for comparison limits how statistically "sure" we can be about the exact probability.

In short, this paper gives us a better pair of glasses for watching how the pieces of a whole change over time. It proves that when you are dealing with shares and percentages, you can't just measure the distance between numbers; you have to measure the distance between the relationships those numbers have with each other. By doing so, it uncovers a hidden story in Pennsylvania's energy history that the old methods simply couldn't see.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →