← Latest papers
📊 statistics

Bayesian Uncertainty Quantification for Ranked Choice Voting Polls

This paper proposes a simple Bayesian framework to quantify uncertainty in ranked choice voting polls by estimating candidate win probabilities, addressing the limitations of traditional frequentist methods in capturing the path-dependent nature of RCV outcomes through applied analyses of the 2021 NYC mayoral primary and the 2022 Alaska special election.

Original authors: Evan T. R. Rosenman, Jason Liang

Published 2026-07-01
📖 6 min read🧠 Deep dive

Original authors: Evan T. R. Rosenman, Jason Liang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Problem: Why Standard Polls Get Lost in "Ranked Choice"

Imagine you are trying to predict the winner of a race. In a traditional election (like a simple "First-Past-The-Post" race), it's like a sprint. You just count who has the most runners crossing the finish line first. If Candidate A has 51% of the runners and Candidate B has 49%, the math is easy: A wins. Pollsters can easily say, "We are 95% sure A is ahead," and give you a "margin of error" (like saying the result could be off by 3%).

But Ranked Choice Voting (RCV) is not a sprint; it's a game of musical chairs with a twist.

In RCV, voters don't just pick one person; they rank them (1st, 2nd, 3rd, etc.). The counting process is complex:

  1. Everyone's 1st choice is counted.
  2. If no one has a majority (over 50%), the person with the fewest votes is eliminated.
  3. The people who voted for the eliminated person have their votes transferred to their next choice.
  4. This happens round after round until someone finally gets over 50%.

The Paper's Big Insight:
The authors argue that standard polling tools (like "margin of error") break down here. Why? Because the winner isn't just about who is currently leading; it depends entirely on who gets eliminated first.

Think of it like a domino effect.

  • In a standard race, if you are slightly behind, you are just behind.
  • In RCV, if you are slightly behind, you might get knocked out early. But if you are slightly ahead of the person in last place, you survive, and the person behind you gets eliminated. That changes everything.

The paper shows that in small polls, you might accidentally think the "wrong" person is going to be eliminated first. If you guess the elimination order wrong, your prediction for the final winner is wrong, even if your poll was otherwise accurate. Standard math tools can't easily measure this "path dependency."

The Solution: A "Crystal Ball" Simulation

The authors propose a new way to look at uncertainty using Bayesian statistics. Instead of trying to calculate a single "margin of error" for a specific round, they suggest simulating thousands of possible realities.

Here is their method, broken down into a simple metaphor:

1. The "Ballot Bag" (The Data)
Imagine every possible way a voter could rank candidates is a unique ticket in a giant bag.

  • Ticket A: "I like Adams 1st, Garcia 2nd."
  • Ticket B: "I like Garcia 1st, Adams 2nd."
  • Ticket C: "I only like Wiley."

When a poll is taken, we pull out a handful of tickets (the sample). We want to know what the entire bag looks like based on that handful.

2. The "Magic Paint" (The Prior)
Since we don't know the exact mix of tickets in the bag, we start with a "flat" assumption (a magic paint) that says, "Every type of ticket is equally likely to exist." As we look at the actual tickets pulled from the poll, we update this paint. If we see a lot of "Adams 1st" tickets, the paint turns more "Adams-colored."

3. The "Simulation Factory" (The Posterior)
This is the core of their idea. Instead of just guessing the final winner, they use a computer to:

  • Mix the paint: Create a new, slightly different version of the "whole bag" based on the poll data and the math rules.
  • Run the race: Take that new bag, run the RCV elimination game (knock out the bottom, transfer votes), and see who wins.
  • Repeat: Do this 500 or 1,000 times, creating 1,000 slightly different "what-if" worlds.

4. The Result: "Win Probabilities"
After running the race 1,000 times, they count the results:

  • "In 564 of these worlds, Adams won."
  • "In 428 of these worlds, Garcia won."
  • "In 8 of these worlds, Wiley won."

Instead of saying "Adams is ahead by 3% with a margin of error of 2%," they can now say: "Based on this poll, Adams has a 56% chance of winning, and Garcia has a 43% chance."

Real-World Examples Used in the Paper

The authors tested this on two famous, messy elections to prove it works:

  1. New York City Mayoral Primary (2021):

    • The Situation: Eric Adams won, but only by a tiny 0.8% margin. Many polls thought his rival would be Maya Wiley, but she got eliminated early. The real final opponent was Kathryn Garcia.
    • The Old Way: Polls were confused. They saw a tie between Garcia and Wiley and couldn't tell who would make it to the final round.
    • The New Way: Their simulation showed that while it was a tight race, Adams had a roughly 56% chance of winning and Garcia had 43%. This perfectly captured the "too close to call" nature of the race without getting stuck on the specific elimination order.
  2. Alaska House Special (2022):

    • The Situation: Mary Peltola won, but Nick Begich III (who was eliminated first) actually would have beaten her in a one-on-one race. This is a classic "path dependency" trap.
    • The Old Way: Small polls often predicted Begich would win because they accidentally thought Palin would be eliminated first.
    • The New Way: Their method correctly identified that as the sample size grew, the probability of Peltola winning increased, showing that the "Begich wins" prediction was just a fluke of small sample sizes.

The "Pruning" Trick

One practical problem they solved is that with many candidates, the number of possible rankings is huge (like trying to sort a deck of cards that keeps growing).

They used a trick called "Pruning."
Imagine you are watching a soccer tournament. You don't need to know the exact order in which the teams in the bottom half of the bracket were eliminated to know who is in the Final Four.

  • The authors realized that for predicting the winner, you often only need to know who the top 3 or 4 candidates are.
  • They developed a way to mathematically "cut off" the bottom candidates from the simulation. This makes the computer run faster and the math more accurate, because it stops worrying about the irrelevant details of who came in 10th place.

Summary

The paper argues that for Ranked Choice Voting, we should stop trying to calculate a single "margin of error" for a specific lead. Instead, we should use a simulation approach that asks: "If we ran this election 1,000 times with slightly different poll results, how often would each candidate win?"

This gives voters and journalists a much clearer picture: not just "who is leading," but "how likely is it that this person actually wins?" It acknowledges that in RCV, the path to victory is winding, and a small shift in the polls can change the entire outcome.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →