E-values as statistical evidence: A comparison to Bayes factors, likelihoods, and p-values
This paper advocates for e-values and e-processes as robust measures of statistical evidence, arguing that they synthesize the desirable properties of p-values, likelihood ratios, and Bayes factors while offering unique advantages in handling composite hypotheses, optional stopping, and intuitive wealth-based interpretations.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to solve a mystery. You have a suspect (the Null Hypothesis, or "the innocent person"), and you are gathering clues (data) to see if they are actually guilty.
For decades, statisticians have argued over the best way to measure "how guilty" the suspect looks based on the clues. The paper you provided introduces a new, powerful tool to this debate called the E-value (and its time-traveling cousin, the E-process).
Here is a simple breakdown of the paper's main ideas, using everyday analogies.
1. The Old Tools: Why They Are Frustrating
Before the E-value, detectives had two main tools, and both had annoying flaws:
- The P-value (The "One-Shot" Snapshot):
- How it works: You take all your clues at once and ask, "How unlikely is this if the suspect is innocent?"
- The Flaw: It's very rigid. If you decide to stop looking for clues early because you found something promising, or if you decide to keep looking because you didn't find anything yet, the P-value breaks. It's like a camera that only works if you promise before the photo is taken exactly how many photos you will take. If you change your mind mid-shoot, the photo is ruined.
- The Bayes Factor (The "Opinionated" Judge):
- How it works: You start with a "gut feeling" (a prior belief) about the suspect's guilt and update it as you see clues.
- The Flaw: It depends entirely on your starting gut feeling. Two detectives with different gut feelings can look at the same clues and come to completely different conclusions. It's not "objective" enough for some.
- The Likelihood Ratio (The "Comparison" Tool):
- How it works: It compares how likely the clues are under "Guilt" vs. "Innocence."
- The Flaw: It struggles when you don't have a clear "Guilty" scenario to compare against. What if you just want to know if the suspect is innocent, without having a specific "Guilty" story ready?
2. The New Tool: The E-Value (The "Betting Score")
The authors propose the E-value. To understand it, imagine a casino game.
- The Setup: You are betting against the "House" (the Null Hypothesis/Innocence).
- The Rule: The House claims the game is fair. If the House is telling the truth, you should never expect to make money in the long run. Your average winnings should be zero or less.
- The E-value: This is your current bankroll.
- You start with $1.
- Every time you get a new clue, you place a bet.
- If the House is truly innocent, your bankroll will stay around $1 or go down.
- But, if the House is actually lying (the suspect is guilty), your bankroll will start to explode. It might go from $1 to $10, then $100, then $1,000.
Why is this cool?
- It's Flexible (Optional Stopping): You can stop the game whenever you want! If your bankroll hits $100, you can cash out and say, "I'm done, the House is lying!" It doesn't matter if you planned to play for 10 rounds or 100. The math holds up.
- It's Additive (Combining Evidence): Imagine you play a game with a friend. You both bet against the same House. If you win $10 and your friend wins $10, you can just multiply your scores ($10 100) to get the total evidence. You don't need to worry about whether your bets were independent or dependent; the math just works.
- It's Objective: You don't need a "gut feeling" about the start. You just need a strategy that guarantees you won't get rich if the suspect is innocent.
3. The "E-Process" (The Time-Traveling Detective)
Sometimes, you don't just get one batch of clues; you get them one by one over time.
- An E-value is like a snapshot of your bankroll at a specific moment.
- An E-process is the entire history of your bankroll as you play the game round after round.
- The magic of the E-process is that it allows you to change your mind about when to stop. You can say, "I'll keep betting until my bankroll hits $50," or "I'll stop if I see three red clues in a row." The E-process guarantees that if the suspect is innocent, you will almost never hit that $50 mark.
4. The Big Comparison
The paper compares these tools against a checklist of what a "good" measure of evidence should be. Here is the verdict:
| Feature | P-Value | Bayes Factor | Likelihood Ratio | E-Value / E-Process |
| :--- | :--- | :--- | :--- :--- |
| Can I stop whenever I want? | ❌ No (Breaks the rules) | ✅ Yes (If you planned ahead) | ✅ Yes | ✅ Yes (The best at this!) |
| Can I combine studies easily? | ❌ Hard (Needs independence) | ✅ Yes (If priors match) | ✅ Yes | ✅ Yes (Just multiply them!) |
| Do I need a "Guilty" story? | ✅ No | ❌ Yes (Need two sides) | ❌ Yes | ✅ No (Can test just one side) |
| Is it objective? | ✅ Mostly | ❌ No (Depends on gut feeling) | ✅ Yes | ✅ Yes (Based on math, not feelings) |
| Does it handle complex data? | ✅ Yes | ⚠️ Hard | ⚠️ Hard | ✅ Yes (Very flexible) |
5. The Catch (Limitations)
Is the E-value perfect? Not quite.
- Too Many Choices: Just like there are many ways to bet in a casino, there are many ways to calculate an E-value. Different statisticians might choose different "betting strategies" and get different numbers. However, the paper argues this is no worse than the P-value (where you can choose different tests) or the Bayes Factor (where you choose different priors).
- It's a "Generalized" Likelihood: The authors show that E-values are basically a super-charged version of the Likelihood Ratio. They keep the good parts of the Likelihood Ratio but fix the parts that break when you have messy, complex data.
The Bottom Line
The paper argues that E-values are the "Swiss Army Knife" of statistical evidence.
- They have the flexibility of a betting game (you can stop anytime).
- They have the objectivity of frequentist math (no gut feelings needed).
- They have the power to combine evidence from many different sources easily.
While they don't solve every philosophical problem in statistics, they offer a very intuitive, robust, and practical way to say, "The evidence against this hypothesis is strong," without getting tripped up by the rigid rules of the past.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.