← Latest papers
🔭 astrophysics

Simple approximations of some statistical functions

This paper proposes simple and sufficiently accurate approximation expressions for the quantiles of the inverse normal distribution, Student's t-distribution, and the outlier rejection criterion to simplify the computation of these statistical functions in hypothesis testing.

Original authors: Zinovy Malkin

Published 2026-04-21
📖 5 min read🧠 Deep dive

Original authors: Zinovy Malkin

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a chef trying to bake a perfect cake for a thousand different parties. In the world of statistics, the "cake" is a complex calculation used to decide if your data is real or just a fluke. Usually, professional statisticians use "super-computers" (or heavy software) to bake these cakes with microscopic precision. They measure every grain of sugar to the billionth of a gram.

But what if you are baking in a small kitchen, or you need to bake a thousand cakes in an hour, and you don't have a super-computer? You don't need a grain-of-sugar measurement; you just need a cake that tastes good enough to serve.

This paper by Zinovy Malkin is essentially a cookbook for "Good Enough" Statistical Recipes.

Here is the breakdown of the three main "dishes" (statistical functions) the author simplifies, using everyday analogies:

1. The "Normal Distribution" (The Bell Curve)

The Problem: Imagine you are measuring the height of people in a city. Most people are average height, with fewer very tall or very short people. This forms a "Bell Curve."
Sometimes, you need to answer the reverse question: "If I want to find the top 1% of tallest people, what height do I need to cut off?"
Mathematically, finding this "cut-off point" (called a quantile) is like trying to solve a complex maze. The standard way to do it involves heavy, slow math.

Malkin's Solution:
Instead of solving the whole maze, Malkin says, "Let's just look at the most common paths people take."
He created a simple shortcut formula (a "back-of-the-napkin" calculation) that works perfectly for the confidence levels we actually use in real life (like 90%, 95%, or 99%).

  • The Analogy: Instead of calculating the exact distance to the moon using complex orbital mechanics, Malkin gives you a ruler that says, "It's about 240,000 miles." It's not perfect, but for planning a picnic, it's accurate enough and takes one second to read.

2. The "Student's t-Distribution" (The Small Sample Detective)

The Problem: The Bell Curve works great when you have millions of data points. But what if you only have 10 measurements? The math gets wobbly and unpredictable. Statisticians use a different shape called the "Student's t-distribution" to handle these small groups.
Calculating the "cut-off" for these small groups is usually a slow, heavy process.

Malkin's Solution:
Malkin realized that for most practical uses, the shape of this curve settles down quickly as you add more data. He created a simple formula that acts like a "smart guess."

  • The Analogy: Imagine you are trying to guess the average weight of a bag of apples. If you have 500 bags, it's easy. If you only have 5 bags, it's harder. Malkin's formula is like a seasoned grocer who looks at the 5 bags and says, "I know the math is tricky, but based on these 5, the answer is roughly X." It's not a supercomputer simulation, but it gets the job done fast.

3. The "Outlier Rejection" (The Bad Apple)

The Problem: Sometimes, in a pile of data, one number is just weird. Maybe you measured a temperature of 200°F in a room that's usually 70°F. That's an "outlier." You need a rule to decide: "Is this a real measurement, or is the thermometer broken?"
The standard rule involves a complex calculation to see if that "bad apple" is far enough away from the rest to be thrown out.

Malkin's Solution:
He created a simple formula to calculate the "throw-away line."

  • The Analogy: Imagine a line of people waiting for a bus. Most are between 5'5" and 6'0". One guy is 7'2". You need a rule to decide if he's a giant or just a guy standing on a box. Malkin's formula is a simple tape measure. If the guy is more than X inches taller than the average, you say, "Okay, that's an outlier, let's check his shoes." The formula tells you exactly what X is without needing a PhD in math.

Why Does This Matter?

The author argues that while "perfect" math exists, it is often overkill.

  • Speed: These simple formulas are like driving a bicycle instead of a rocket ship. You get to the destination (the answer) much faster.
  • Simplicity: They are easy to program into simple devices, like a handheld sensor or an old computer, that can't handle heavy math.
  • Accuracy: The paper proves that for 99% of real-world jobs (astronomy, engineering, quality control), these "bicycle" formulas are accurate enough. You don't lose any real information; you just save a lot of time.

The Bottom Line

Zinovy Malkin is saying: "Stop using a sledgehammer to crack a nut."
He has provided a set of simple, fast, and reliable tools for scientists and engineers to make quick decisions about their data without needing a supercomputer. It's about efficiency: getting a result that is "good enough" to be trusted, in a fraction of the time.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →