← Latest papers
📊 statistics

Heteroscedasticity of Denoising Score Matching with Generalised Smooth Noise

This paper reveals that Denoising Score Matching (DSM) suffers from inherent heteroscedasticity due to noise levels and data geometry, and proposes a theoretically derived weighting function to stabilize training variance while providing a justification for existing heuristics in diffusion models.

Original authors: Juyan Zhang, Rhys Newbury, Xinyang Zhang, Tin Tran, Dana Kulic, Michael Burke

Published 2026-08-05
📖 4 min read☕ Coffee break read

Original authors: Juyan Zhang, Rhys Newbury, Xinyang Zhang, Tin Tran, Dana Kulic, Michael Burke

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a world where computers learn to create art, music, or even new molecules by studying a giant, messy library of existing examples. To do this, they use a clever trick called "Score Matching." Think of the data (like a picture of a cat) as a landscape with hills and valleys. The "score" is simply a compass that always points uphill toward the most likely places to find a cat. If the computer can learn to hold this compass perfectly, it can wander through the landscape and eventually find a brand-new, realistic cat to draw.

But here's the catch: the computer can't see the whole map at once. It's like trying to learn the shape of a mountain while standing in a thick fog. So, instead of looking at the real mountain, the computer practices on a version of the mountain that has been covered in static noise, like a TV screen full of snow. It tries to guess how to clean the noise off. This practice method is called "Denoising Score Matching" (DSM). For a long time, scientists assumed this practice was a perfect, free substitute for the real thing. They thought, "If the average direction of the compass is right, we're good to go." But this paper asks a nagging question: Is the practice field actually hiding a secret trap that makes learning unstable?

The authors of this paper, a team from Monash University and Amazon, have discovered that the practice field is indeed a bit of a trickster. They found that Denoising Score Matching is inherently "heteroscedastic." That's a fancy word for saying the amount of "noise" or uncertainty in the computer's learning signal changes wildly depending on where it is in the data landscape.

To use a playful analogy, imagine you are trying to learn to throw darts at a moving target. In a perfect world, every throw would be equally hard or easy. But in this "Denoising" game, some throws are like tossing a dart in a calm room, while others are like trying to throw while standing on a shaking boat in a storm. The paper proves that the "shaking boat" parts happen naturally in specific regions of the data, like the edges between different clusters of information. Because the computer doesn't know which throws are on the boat and which are on solid ground, it treats them all the same. This causes the learning process to wobble and become inefficient, like a student trying to study for a test while someone keeps turning the lights on and off.

The researchers didn't just find the problem; they built a theoretical "stabilizer" to fix it. They derived a special mathematical formula, called a "Godambe weighting," that acts like a smart filter. This filter tells the computer, "Hey, that throw was on a shaking boat; don't trust it as much. But that other throw was on solid ground; pay extra attention to that." By adjusting the importance of each piece of information based on how shaky it is, the computer can learn much more smoothly.

However, there is a twist. The perfect filter requires knowing the exact shape of the shaking boat, which is often impossible to calculate for complex, high-dimensional data (like real-world images). So, the authors also proposed a "good enough" approximation. They showed that a simple, existing trick used by many modern AI models—weighting the learning by the square of the noise level—actually emerges naturally from their math. This explains why that simple trick works so well in practice, even though it's not the perfect solution.

Ultimately, the paper reveals a fundamental trade-off. You can have a mathematically perfect, statistically efficient learning method, but it might be too unstable to train in the real world. Or, you can use a slightly less perfect, "approximate" method that keeps the training stable and gets the job done. The authors prove that the popular methods we use today are essentially making a smart compromise: they sacrifice a tiny bit of statistical perfection to avoid the chaos of the "shaking boat," ensuring the AI can actually learn to create those amazing images and sounds we see today.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →