← Latest papers
🔬 condensed matter

Phase transition in large language models and the criticality of natural languages

By treating large language models as controllable effective systems, this study demonstrates that varying a temperature-like parameter induces a phase transition where the critical point exhibits power-law behavior and linguistic complexity indistinguishable from natural languages, thereby providing strong evidence that natural languages operate at a state of criticality.

Original authors: Kai Nakaishi, Yoshihiko Nishikawa, Koji Hukushima

Published 2026-06-09
📖 4 min read☕ Coffee break read

Original authors: Kai Nakaishi, Yoshihiko Nishikawa, Koji Hukushima

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot to speak like a human. You give it a massive library of books and let it learn the patterns of words. This robot is a Large Language Model (LLM). But here's the twist: the researchers in this paper treated this robot not just as a tool, but as a scientific experiment to understand the very nature of human language itself.

They wanted to answer a big question: Is human language "critical"?

In the world of physics, "criticality" is a special state where a system is balanced right on the edge between two very different behaviors. Think of it like a pot of water.

  • Below the boiling point (Low Temperature): The water is calm and ordered.
  • Above the boiling point (High Temperature): The water is chaotic, bubbling, and disordered.
  • At the exact boiling point (Critical Point): The water is in a magical state where it has complex, swirling patterns that connect everything together. It's neither fully calm nor fully chaotic; it's in a "sweet spot" of complexity.

The researchers suspected that human language lives in this "sweet spot." But you can't turn a human language "up" or "down" like a thermostat to test this. So, they used the LLM as a simulated language where they could turn the knob.

The Experiment: Turning the "Temperature" Knob

In these AI models, there is a setting called Temperature.

  • Low Temperature: The AI becomes very predictable and repetitive. It might say, "The number of schools in the United States... The number of schools in the United States..." over and over. It's like a broken record.
  • High Temperature: The AI becomes completely crazy. It spits out nonsense like "detached speeches tailor hello cellular networks early..." It's like a radio tuned to static.

The researchers turned this knob from low to high and watched what happened to the text the AI generated.

The Discovery: The "Phase Transition"

They found that the AI didn't just slowly change from boring to crazy. It went through a sudden, sharp phase transition right around a specific setting (Temperature = 1).

  1. The Low-Temperature Zone (The Repetitive Phase): The text was stuck in loops. It had order, but it was a boring, rigid order.
  2. The High-Temperature Zone (The Nonsense Phase): The text was pure chaos with no connection between words.
  3. The Critical Point (The "Just Right" Zone): Right in the middle, at the transition point, the AI started generating text that looked and felt exactly like human language.

Why is this special?
At this critical point, the text showed a specific mathematical pattern called a power law. This is a pattern where small things happen often, and big things happen rarely, but they are connected in a specific way. This is the same pattern found in real human language (like how common words like "the" appear often, while rare words appear rarely).

The researchers found that:

  • Real human language has this power-law pattern.
  • The AI at the critical point also has this pattern.
  • The AI at any other setting (too cold or too hot) does not have this pattern.

The "Training" Story

To prove this wasn't just a fluke of the AI, they watched the AI learn from scratch.

  • Early in training: The AI was just a random noise machine. Even at the "critical" setting, it couldn't make sense. It didn't know about the "critical point" yet.
  • Later in training: As the AI read more human books, it suddenly "woke up." It developed the ability to sit at that critical point. It learned to balance between being too repetitive and being too chaotic.

This suggests that being "critical" is a fundamental feature of human language. To speak like a human, a system must operate on this edge between order and chaos.

The Final Proof: The "Perplexity" Test

In the world of AI, there is a score called Perplexity that measures how well a model understands a language. Lower is better.

  • The researchers tested the AI at every stage of its training and at every temperature setting.
  • The lowest score (the best match to human language) happened exactly when the AI was fully trained and set to the critical temperature.

The Big Picture

The paper concludes that human language is not just a random collection of words. It is a critical system. It naturally exists on the razor's edge between:

  • Repetition: (Like a broken record)
  • Nonsense: (Like random noise)

And in that tiny, critical space in between, we find the complex, beautiful, and meaningful structure of human speech. The AI didn't just learn to mimic us; by finding this critical point, it accidentally discovered the mathematical "heartbeat" of language itself.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →