← Latest papers
🔢 mathematics

Sharp Sobolev Sandwich and Approximation Rates of Radon-Domain LpL^p Ridge Integral Spaces for ReLUk^k Networks

This paper establishes that the Radon-domain LpL^p space of functions representable by shallow ReLUk\mathrm{ReLU}^k networks forms a sharp Sobolev sandwich around the critical regularity space Hk+(d+1)/2H^{k+(d+1)/2}, with the gap determined by the Seeger–Sogge–Stein loss, and leverages this theory to derive optimal LpL^p approximation rates for discretized neural networks.

Original authors: Juncai He, Zitong Tian

Published 2026-06-24
📖 4 min read🧠 Deep dive

Original authors: Juncai He, Zitong Tian

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to build a complex 3D sculpture (a mathematical function) using only simple, flat sheets of material. In the world of machine learning, these "sheets" are called neurons, and the way they are stacked together is called a neural network.

This paper is like a master blueprint that explains exactly how well you can build any shape using a specific type of sheet called ReLUk. The "k" just means the sheet can be bent or folded kk times (making it smoother or more flexible).

Here is the breakdown of their discovery, using simple analogies:

1. The Problem: How Many Sheets Do You Need?

For a long time, we knew you could build almost any shape with enough sheets (neurons). But we didn't know how many you needed to get a specific level of detail.

  • The Question: If I want my sculpture to be smooth and accurate, do I need 10 sheets, 1,000, or 1,000,000?
  • The Goal: The authors wanted to find the exact "recipe" for the smoothest possible shapes and the most efficient way to build them.

2. The Secret Ingredient: The "Radon Domain"

To solve this, the authors didn't look at the sculpture from the front. Instead, they looked at it through a magical lens called the Radon Transform.

  • The Analogy: Imagine taking a loaf of bread and slicing it into thin pieces from every possible angle. The Radon Transform is the collection of all those 2D slices.
  • The Discovery: The authors realized that if you look at the "slices" (the Radon domain) rather than the whole loaf, the math becomes much clearer. They defined a special "space" (a library of functions) based on how smooth these slices are. They call this the Radon-domain LpL^p space.

3. The "Sandwich" Discovery

This is the paper's biggest "Aha!" moment.

  • The Perfect Case (p=2p=2): When looking at the problem in a specific mathematical way (like measuring average error), the library of shapes you can build with these neurons is exactly the same as a famous class of smooth shapes known as Sobolev spaces. It's a perfect match, like two puzzle pieces fitting together with no gaps.
  • The General Case (1<p<1 < p < \infty): When you change the way you measure the error (looking at different types of "roughness"), the perfect match turns into a Sandwich.
    • The Bread (Top): A slightly smoother class of shapes.
    • The Bread (Bottom): A slightly rougher class of shapes.
    • The Filling: The shapes your neural network can actually build.
    • The Gap: The authors calculated the exact size of the gap between the top and bottom bread. This gap is caused by a known mathematical "friction" (called the Seeger–Sogge–Stein loss) that happens when you slice and reassemble the data. It's the unavoidable cost of turning slices back into a loaf.

4. Why Does This Matter? (The Approximation Rate)

Now that they know exactly what kind of shapes these networks can build, they can predict how fast the network learns.

  • The Recipe: They showed that if you pick your neurons randomly (like grabbing random slices of bread) but in a smart, uniform way, you can build a very accurate sculpture very quickly.
  • The Result: They proved that for the smoothest shapes, the error drops at the fastest possible speed allowed by math.
    • If you double the number of neurons, the error doesn't just go down a little; it drops at a specific, optimal rate.
    • They also showed how to remove a small "logarithmic" penalty that previous methods had, making the process even more efficient.

Summary in Plain English

Think of the authors as architects who finally figured out the exact physics of building with "ReLU" bricks.

  1. They found a special way to look at the bricks (the Radon domain) that reveals their true potential.
  2. They proved that for the most common measurement, these bricks can build exactly the smoothest possible structures.
  3. For other measurements, they proved the structures fit perfectly between two known limits (the "Sandwich"), with the gap size being mathematically unavoidable.
  4. Finally, they showed that by using a simple, random sampling method, you can build these structures with the maximum possible speed and efficiency, proving that these simple networks are incredibly powerful tools for learning smooth patterns.

In short: They didn't just say "neural networks work." They wrote the exact instruction manual for how well they work, why they work, and how fast they can learn, using a clever mathematical lens to see the hidden structure of the data.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →