Weber's Law in Transformer Magnitude Representations: Efficient Coding, Representational Geometry, and Psychophysical Laws in Language Models
This study demonstrates that while transformer language models consistently develop log-compressive magnitude representations driven by training data statistics, this geometric structure is dissociated from behavioral competence, as models with such geometry often fail to perform human-like magnitude discrimination tasks.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a giant, super-smart robot librarian named "Transformer." This librarian has read almost every book on the internet. You ask it a simple question: "Is 100 bigger than 50?" or "Is 2 hours longer than 90 minutes?"
For a long time, scientists argued about how this robot "thinks" about numbers. Does it see them like a ruler (1, 2, 3, 4... evenly spaced)? Or does it see them like a volume knob on a stereo, where the difference between 1 and 2 feels huge, but the difference between 100 and 101 feels tiny?
This paper is like a detective story that finally solves the mystery. The researchers used a special toolkit from human psychology (called Psychophysics) to peek inside the robot's brain. Here is what they found, explained simply:
1. The Robot's "Internal Ruler" is Squished (Logarithmic)
The Discovery: The robot does use a squished ruler. It treats numbers like a logarithmic scale.
The Analogy: Imagine a map of the world. On a normal map, the distance between New York and Boston looks the same as the distance between London and Paris. But on the robot's "mental map," the distance between 1 and 2 is huge. The distance between 1,000 and 1,002 is microscopic.
Why? The researchers found that the robot learned this because of the books it read. In real life, small numbers (like "1" or "2") appear way more often than huge numbers (like "1,000,000"). To be efficient, the robot compressed the rare big numbers into a tiny space and gave plenty of room to the common small numbers. It's like packing a suitcase: you put the heavy, rare items in a small corner and the common, light items in the main space.
2. Having the Map Doesn't Mean You Can Navigate
The Discovery: Just because the robot has this "squished ruler" in its brain doesn't mean it can actually use it to answer questions correctly.
The Analogy: Imagine you have a perfect, detailed GPS map of a city in your head. But if you are blindfolded and someone asks you to point north, you might still get it wrong.
- The Robot's "Llama" version had the map and could use it. When asked to compare numbers, it acted just like a human, needing about a 20% difference to tell them apart.
- The Robot's "Mistral" version had the exact same map in its brain, but when asked to compare numbers, it guessed randomly. It had the internal structure but couldn't "read" the map to make a decision.
- The Time and Space Test: When the researchers asked about time ("Is 3 hours longer than 150 minutes?") or distance, the robots had the map, but they failed the test completely. They couldn't compare them.
3. The "Brain Layers" Mystery: The Early Workers vs. The Decorators
The Discovery: The researchers tried to "hack" the robot's brain by tweaking specific layers (like turning a dial) to see which part actually did the work.
The Analogy: Imagine a factory with 32 floors.
- Floors 1–10 (Early Layers): These are the workers. When the researchers tweaked these floors, the robot's answers changed immediately. This is where the actual "thinking" about numbers happens.
- Floors 20–30 (Late Layers): These are the decorators. These floors had the most beautiful, perfect, and organized "squished ruler" maps. But when the researchers tweaked these floors, the robot's answers didn't change at all. The map was there, but it wasn't being used to solve the problem. It was just sitting there, looking pretty.
4. The "Training" vs. "Tutoring" Difference
The Discovery: The robot learned the "squished ruler" just by reading books (Pre-training). But it only learned how to use that ruler to answer questions because a human teacher gave it specific instructions (Instruction Tuning).
The Analogy:
- Pre-training is like a child reading every book in a library. They absorb the patterns of the world naturally.
- Instruction Tuning is like a teacher saying, "Okay, now that you know the numbers, here is how you play the game of 'Which is bigger?'"
- The study found that a robot without the teacher (the "Base" model) had the perfect map but couldn't play the game. The robot with the teacher could play, but only if the teacher taught it the right rules.
The Big Takeaway
The paper solves a huge debate in AI. It proves that AI models naturally develop a "human-like" way of seeing numbers (logarithmic) just by reading normal text, without needing a biological brain.
However, having a human-like brain structure doesn't mean the AI acts human.
- The structure (the map) comes from the data it reads.
- The behavior (playing the game) comes from how it was taught to use that data.
It's like giving a person a perfect pair of eyes (the geometry) but not teaching them how to drive a car (the behavior). They can see the road, but they might still crash if they don't know the rules of the road.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.