← Latest papers
📊 statistics

Statistical Compatibility, Refutational Information, and Acceptability

This paper proposes a descriptive frequentist framework that reinterprets P-values and S-values not as measures of logical impossibility or direct evidence against a model's truth, but as graded indices of data-model compatibility and refutational information that must be integrated with contextual factors and analyst judgment to determine a model's practical acceptability.

Original authors: Alessandro Rovetta

Published 2026-03-31
📖 6 min read🧠 Deep dive

Original authors: Alessandro Rovetta

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: Don't Confuse "Weird" with "Impossible"

Imagine you are a detective trying to solve a mystery. You have a theory (a Model) about how the crime happened. In statistics, this model is like a "fictional world" where you pretend your theory is 100% true just to see what would happen.

Rovetta's paper argues that when we do statistical tests, we often make a mistake. We see a result that looks "weird" or "rare" in that fictional world, and we immediately scream, "The theory is wrong!"

Rovetta says: Stop. A weird result doesn't mean the theory is logically impossible. It just means the result is unusual for that specific theory. We need to be much more careful about how we decide if a theory is actually "acceptable" in the real world.


1. The "Fictional World" (The Model)

Think of a statistical model like a video game simulation.

  • The Assumptions (A): These are the rules of the game (e.g., gravity works this way, coins are fair, dice are perfect).
  • The Hypothesis (Hi): This is a specific claim you are testing inside the game (e.g., "This coin is fair").

When you run a test, you are stepping inside the video game. You ask: "If the rules are true and the coin is fair, how likely is it that I would see this specific result?"

The Mistake: People often think that if the result is rare in the game, the game rules must be broken. But the game rules are just a tool we invented to measure things. The result is just "rare," not "forbidden."

2. The P-Value: A "Compatibility Score"

Instead of saying a result is "incompatible" (which sounds like a total clash), Rovetta suggests we call it a Compatibility Score.

  • High Score (P-value near 1): The result fits the story perfectly. It's like rolling a 7 with two dice; it's exactly what you expect.
  • Low Score (P-value near 0): The result is a weird outlier. It's like rolling two dice and getting a 12 three times in a row.

The Analogy: Imagine you are wearing a pair of glasses that make everything look blue. If you see a red apple, the apple isn't "incompatible" with the glasses; it's just that the apple looks very strange through them. The P-value tells you how strange the apple looks through your specific pair of glasses. It does not prove the apple isn't red.

3. The S-Value: The "Coin Flip" Meter

Statisticians often use a number called the S-value to make the P-value easier to understand.

  • The Analogy: Imagine you are flipping a coin.
    • If you get 1 Head in 1 toss, that's not a big deal. (Low S-value).
    • If you get 10 Heads in 10 tosses, that's suspicious. (High S-value).
    • If you get 20 Heads in 20 tosses, that's almost impossible for a fair coin. (Very High S-value).

The S-value tells you: "This result is as surprising as getting X heads in a row with a fair coin." It turns a confusing decimal (like 0.00003) into a concrete image (like "20 heads in a row").

4. The Real Question: Is the Model "Acceptable"?

This is the most important part of the paper. Just because a result is "surprising" (like 10 heads in a row) doesn't mean you should immediately throw away your theory. You have to ask: "Is this theory still useful for my decision?"

This depends on Context and Loss (what you lose if you are wrong).

The Betting Analogy:
Imagine you are betting on whether a coin is fair.

  • Scenario A: If you guess wrong, you lose $1. If you guess right, you win $100.
    • Decision: You might take the bet even if the coin looks a little weird. The risk is low, the reward is high.
  • Scenario B: If you guess wrong, you lose your house. If you guess right, you win $100.
    • Decision: You need extreme proof (like 50 heads in a row) before you are willing to bet.

The Medical Analogy:

  • Mild Side Effect (Headache): If a drug causes a weird result suggesting headaches, you might accept the risk because the drug helps with a cold.
  • Severe Side Effect (Heart Attack): If the same drug suggests a risk of heart attacks, you need massive evidence before you accept that the drug is safe.

The Lesson: The "weirdness" of the data (the S-value) is just one piece of information. The final decision depends on how much it hurts to be wrong.

5. "Surprise" vs. "Belief Change"

The paper also clarifies what "surprise" means.

  • Statistical Surprise (Shannon): "Wow, that outcome was rare in the video game." (This is just a calculation).
  • Real Surprise (Bayesian): "Wow, that outcome changed my mind about the world." (This is a belief update).

The Football Analogy:
Imagine a terrible football player kicks a perfect goal.

  • Statistically: It's a "surprise" because it's rare for him to do that.
  • Belief-wise: You might not change your mind. You might think, "He got lucky this one time; he's still a bad player."

Rovetta argues that in frequentist statistics, we are mostly talking about the first kind (Statistical Surprise). We are calculating how rare the event is, not necessarily rewriting our entire worldview yet.

6. The Takeaway for Scientists

If you are doing a single study, don't try to make a final "Yes/No" decision on whether a theory is true. That's too much pressure.

Instead, think of your study as providing a clear description.

  • "Here is the data."
  • "Here is how weird it looks under our assumptions."
  • "Here is the S-value (how many coin flips it would take to match this weirdness)."

Let the real world (doctors, policymakers, other scientists) take that information, mix it with their knowledge, and decide if the theory is "acceptable" for their specific needs.

Summary in One Sentence

Don't treat a rare statistical result as a logical impossibility; treat it as a "warning light" that, combined with real-world consequences and human judgment, helps us decide if a theory is still good enough to use.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →