← Latest papers
🤖 AI

On the Smallness of the Large Language Models Scaling Exponents

This paper argues that current Large Language Model scaling exponents indicate an unsustainable energy regime, a conclusion that persists even after accounting for the "pedestal effect" of non-zero loss limits, while also drawing analogies to fluid turbulence to discuss the influence of data smoothness on these exponents.

Original authors: Sauro Succi, Peter V. Coveney, Alex Hansen

Published 2026-06-24
📖 5 min read🧠 Deep dive

Original authors: Sauro Succi, Peter V. Coveney, Alex Hansen

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Picture: The "Bigger is Better" Problem

Imagine you are trying to teach a robot to speak perfectly. The current rule in the AI world is: "The bigger the robot, the better it speaks." This idea, called the "no-wall" finding, suggests that if you just keep adding more data and more computer power, the robot will keep getting smarter forever.

However, the authors of this paper argue that this strategy is unsustainable. They say we are hitting a wall of energy consumption, even if we haven't hit a wall of intelligence yet.

The Core Problem: Diminishing Returns

The paper uses a simple math concept to explain why this is a problem. Imagine you are trying to clean a very dirty window.

  • The Goal: Make the window perfectly clear (zero errors).
  • The Current Reality: To make the window just half as dirty, you don't need twice the effort. You need 1,000 times the effort.

The authors point out that current Large Language Models (LLMs) have a "scaling exponent" (a number that measures how fast they get better) that is very small (around 0.05 to 0.1).

  • The Analogy: If you want to improve the AI's performance by a tiny bit, you have to feed it a mountain of new data and burn a massive amount of electricity. It's like trying to push a boulder up a hill that gets steeper the higher you go. The paper argues this is a dead end for energy resources.

The "Pedestal" Defense: Why We Can't Just Wait

Some critics say, "Don't worry! We don't actually need the AI to be perfect (zero error). We just need it to be 'good enough' to stop talking nonsense. We stop training before it hits the bottom."

The authors call this the "Pedestal Effect." They imagine a floor (the pedestal) that the AI's errors can never go below, no matter how much data you give it.

  • The Paper's Rebuttal: Even if we accept that we don't need perfection, the math still doesn't work. The authors show that because the "floor" is so high relative to the starting point, the "good enough" zone is reached so quickly that the tiny scaling exponent still applies.
  • The Metaphor: Imagine you are trying to fill a bathtub. You argue, "I don't need it full to the brim; I just need it to have a few inches of water." The authors say, "Even with that low goal, the faucet is dripping so slowly that it will take you a million years to get those few inches, and you'll run out of water (energy) long before you get there."

Why Are the Numbers So Small? (The "Roughness" of Data)

The paper asks: Why is the AI getting better so slowly?

They compare AI training to fluid turbulence (like water swirling in a river or smoke rising).

  1. Smooth vs. Rough: If you are trying to learn a smooth, predictable pattern (like a straight line), you learn fast. But real-world data (language, images, complex systems) is "rough" and chaotic, like a stormy ocean.
  2. The Fractal Connection: The authors use an analogy from physics. In a storm, the energy isn't spread out evenly; it's concentrated in tiny, chaotic swirls. To understand the storm, you have to look at those tiny swirls.
  3. The Result: Because the data is "rough" and chaotic, the AI has to work much harder to find the patterns. The paper suggests that the "roughness" of the data forces the scaling exponent to stay very low.

The "Hidden Map" Theory

The paper references a theory by Sharma and Kaplan (SK) which suggests that the speed of learning depends on the Intrinsic Dimension of the data.

  • The Analogy: Imagine a library.
    • The Embedding Space (the size of the library) is huge—maybe millions of shelves.
    • The Intrinsic Dimension (the actual number of books that contain the real story) is much smaller, but still very large.
  • The Problem: Even though LLMs are amazing at finding the "real story" hidden in the massive library (compressing the data), the "real story" is still too complex (too many dimensions) for the AI to learn efficiently without burning up the world's energy supply.

The Conclusion: Stop Chasing "Bigger"

The authors conclude that the "bigger is better" approach is a recipe for energy burnout.

  • The Verdict: No matter how smart the AI gets, if the data is complex and "rough," the energy required to make it slightly smarter will always be too high.
  • The Suggested Path: Instead of just building bigger, dumber models that eat more electricity, we need to go back to foundational physics-aware models (called "world models"). These are systems that understand how the world works (like gravity or cause-and-effect) rather than just memorizing patterns in a massive database.

In short: We are trying to solve a complex puzzle by throwing more and more people at it, but the puzzle is so hard that we are running out of coffee (energy) before we solve it. We need to change the strategy, not just add more workers.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →