← Latest papers
🤖 machine learning

Small-Scale Experiments: Are We There Yet?

This paper argues that the perceived failure of small-scale experiments to validate scaling laws stems from hyperparameter sensitivity rather than model size, demonstrating that with proper tuning and a holistic methodology, small models can reliably predict large-scale outcomes such as the superiority of pre-normalization in transformers.

Original authors: Nicholas Lourie, Kyunghyun Cho, Karen Ullrich, Sanae Lotfi

Published 2026-08-13
📖 4 min read☕ Coffee break read

Original authors: Nicholas Lourie, Kyunghyun Cho, Karen Ullrich, Sanae Lotfi

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to bake the perfect loaf of bread, but you only have a tiny, cramped kitchen. You want to know how to bake a giant, city-sized loaf for a festival, but baking that giant loaf would cost a fortune in flour and oven time. So, you decide to test your recipes on tiny, 4-ounce rolls first. This is the dream of "scaling laws" in the world of artificial intelligence: the idea that if you understand how a small AI model learns, you can predict exactly how a massive one will behave without having to spend millions of dollars to build it.

For a long time, scientists hoped this "small-to-large" shortcut would work perfectly. They believed that if they tweaked the recipe for a tiny model, they could just scale up the ingredients and get a giant, perfect model. But lately, things have been messy. When researchers tried to use their small experiments to predict the behavior of huge models, the predictions often failed. It was like the tiny rolls tasted great, but the giant loaves turned out burnt or flat. The big question became: Is the shortcut broken? Do we really have to build the expensive giant loaf every single time to see if it works?

This paper, titled "Small-Scale Experiments: Are We There Yet?", dives into this kitchen mystery. The authors, a team from Meta and New York University, argue that the shortcut isn't broken; we just haven't been looking at the right part of the recipe. They discovered that the problem isn't the size of the model, but how carefully we tune the "knobs and dials" (called hyperparameters) on the small ones.

Think of a small AI model like a very sensitive, high-performance race car engine. If you turn the fuel mixture knob just a tiny bit wrong, the engine sputters and dies. It's incredibly hard to get it running perfectly. A giant AI model, on the other hand, is like a massive, slow-moving steamship. It's much more forgiving; even if you adjust the knobs a bit clumsily, the ship keeps moving forward just fine. The authors found that because small models are so sensitive, researchers often miss the "perfect" setting because they don't try enough different combinations. They might test four or sixteen settings and say, "This is the best we can do," not realizing that the true "perfect" setting is hiding somewhere else in the vast ocean of possibilities.

The paper suggests that if you are willing to do the hard work of testing hundreds of different settings on your tiny models, you can find the perfect recipe. Once you find that perfect setting for the small model, the rules for how it scales up become clear and predictable. They showed this by testing two different ways of building AI brains (called "pre-normalization" and "post-normalization"). In the past, figuring out which one was better took years and massive computers. But by using their new method of intense, small-scale testing, they were able to predict the winner correctly using only tiny models.

The authors also explain why this happens using a cool geometric idea. They say that as a model gets bigger, the "landscape" of possible settings becomes simpler. For a small model, the landscape is a jagged, confusing mountain range with thousands of hidden valleys where the engine might stall. But as the model grows, that landscape smooths out into a gentle, wide valley where it's much easier to find the bottom. This means that while small models are hard to tune, big models are actually easy to tune.

So, the big takeaway is that we don't have to give up on small experiments. We just need to stop skimming the surface with our testing. If we treat small models with the respect they deserve—testing them thoroughly rather than just skimming the surface—we can save a massive amount of money and time. We can finally trust our small-scale experiments to tell us the truth about the giants, proving that the "small-to-large" shortcut is still valid, provided we have the patience to find the right settings first.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →