← Latest papers
🤖 AI

A Rational Account of Categorization Based on Information Theory

This paper proposes a new information-theoretic rational analysis of categorization that demonstrates superior or comparable explanatory power over existing models when evaluated against classic experimental findings.

Original authors: Christophe J. MacLellan, Karthik Singaravadivelan, Xin Lian, Zekun Wang, Pat Langley

Published 2026-04-01
📖 6 min read🧠 Deep dive

Original authors: Christophe J. MacLellan, Karthik Singaravadivelan, Xin Lian, Zekun Wang, Pat Langley

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: How Our Brains Sort the World

Imagine your brain is a massive, chaotic library. Every time you see a new object—a dog, a chair, a weird fruit—you have to decide where to put it on the shelf. Do you put the Golden Retriever with the Poodle? Or do you put it with the "Animals that bark"?

For decades, scientists have debated how we do this. Do we remember every single specific dog we've ever seen (like a photo album)? Or do we just remember a "perfect average dog" (a prototype)?

This paper proposes a new theory: Our brains aren't just memorizing or averaging. Instead, they are acting like super-efficient data compressors. Our goal is to organize the world in a way that gives us the most "information" about what we are seeing.

Think of it like packing for a trip. You don't just throw everything in a bag randomly. You organize your clothes so that if you pull out a "swimsuit," you instantly know you're going to the beach, and if you pull out a "tuxedo," you know it's a fancy party. You want the label (the category) to tell you as much as possible about the contents (the features).

The New Rule: "Maximize the Clues"

The authors, led by Christopher MacLellan, suggest that humans learn categories based on a concept called Mutual Information.

  • The Old Way: Previous theories said we try to minimize errors or just match what we see to a stored list.
  • The New Way: We try to build a mental filing system where the category name gives us the maximum possible clues about the object's features.

The Analogy: Imagine you are a detective.

  • If you see a "Bird," you want to know immediately that it has feathers and a beak.
  • If you see a "Fish," you want to know it has scales and gills.
  • If your category "Bird" also included "Fish," your clues would be weak. You wouldn't know if the animal has feathers or scales.
  • The Theory: Our brains naturally sort things into groups where the group name is the most helpful clue possible.

The Computer Brain: "Cobweb"

To test this, the researchers used a computer program called Cobweb. Think of Cobweb as a robot librarian that is trying to organize a library of books (or in this case, data points) without being told the rules.

  1. The Hierarchy: Cobweb builds a tree structure. At the top, it has broad categories (like "Animal"). As it learns more, it splits them into smaller branches (like "Dog" -> "Poodle").
  2. The Decision: Every time a new item arrives, Cobweb asks: "If I put this item here, will it help me predict the features of future items better?"
  3. The Search: When the robot needs to guess what an item is, it doesn't just look at one file. It looks at the most relevant files in its tree, weighted by how much "information" they provide.

Testing the Theory: Three Classic Experiments

The team tested their robot librarian against three famous human experiments to see if it acted like a real person.

1. The "Central Tendency" Test (The Prototype Effect)

  • The Human Behavior: If you show people a bunch of "Club 1" members who are all tall, blond, and like jazz, and then show them a new person who is exactly that description (the "Prototype"), they recognize them instantly. Even if they've never seen that exact person before, they know they belong to the club.
  • The Result: The robot did this perfectly. It realized that the "average" or "ideal" version of a category is the most informative. Even without being explicitly taught the prototype, the robot's math naturally made the prototype the most recognizable item.

2. The "Tricky Shapes" Test (Linear vs. Non-Linear)

  • The Human Behavior: Humans are great at spotting patterns that aren't just simple "A + B = C." Sometimes, a shape belongs to a group not because of one feature, but because of a specific combination of features.
  • The Result: Older computer models (like the "Independent Cue" model) failed here because they treated features separately. The "Context" model (which remembers specific examples) worked well. Our new robot worked just as well as the best human-like models, proving that its "information maximization" rule can handle complex, tricky patterns without needing to memorize every single example.

3. The "Switching Gears" Test (Prototype to Exemplar)

  • The Human Behavior: This is the most fascinating part. When we first learn a category, we act like Prototype theorists (we look for the "average" member). But as we get more experience, we start acting like Exemplar theorists (we remember specific, weird exceptions).
  • The Result: The robot did exactly this!
    • Early on: It grouped things broadly, ignoring the weird exceptions (like a "Penguin" that doesn't fly).
    • Later: After seeing the weird exceptions enough times, it realized the broad category wasn't giving enough information. So, it created a new, specific branch for the exception.
    • The Metaphor: Imagine you meet a few dogs. You think, "All dogs bark." Then you meet a Chihuahua that doesn't bark. At first, you ignore it. But after meeting 10 Chihuahuas that don't bark, you create a new mental category: "Chihuahuas." Your brain shifted from a general rule to a specific memory because the specific memory became more "informative."

Why This Matters

This paper suggests that our brains aren't broken or lazy; they are highly efficient data scientists.

We don't categorize things just because we are told to. We categorize them because it helps us predict the future. If we group things in a way that maximizes the information we get, we can navigate the world more effectively.

In a nutshell:

  • Old View: We categorize to minimize mistakes.
  • New View: We categorize to maximize clues.
  • The Proof: A computer program built on this "clue-maximizing" rule behaves almost exactly like a human, switching between general rules and specific memories depending on what helps it understand the world best.

This theory bridges the gap between "remembering the average" and "remembering the specific," showing that our brains do both because it's the most rational way to process information.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →