← Latest papers
🤖 AI

From Checklists to Clusters: A Homeostatic Account of AGI Evaluation

This paper argues that AGI should be evaluated not as a static sum of symmetric domain scores, but as a homeostatic property cluster, proposing new metrics that weight domains by causal centrality and measure the persistence of capabilities under stress to distinguish durable intelligence from brittle performance.

Original authors: Brett Reynolds

Published 2026-07-24
📖 5 min read🧠 Deep dive

Original authors: Brett Reynolds

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to understand what makes a person "smart." For a long time, scientists have tried to measure this by giving people a giant list of different tasks—solving math puzzles, remembering stories, spotting patterns in shapes, and understanding language. They call these "domains." The idea is that if you do well on all of them, you have high intelligence. But there's a catch: when we give these tests, we usually treat every single task as if it's equally important. It's like grading a student by giving the same weight to their ability to tie their shoes as they get for their ability to write a poem. Also, these tests are usually "snapshots." You take the test once, get a score, and that's it. You don't know if the student will still be able to solve the puzzle next week, or if they'll crumble if you make the test a little harder or wait a few days before they try again.

This is the world of Artificial General Intelligence (AGI) evaluation. AGI is the dream of building a computer that is as smart and flexible as a human, capable of learning and solving problems in any situation, not just the one it was programmed for. Right now, researchers test these AI systems using those same "snapshot" checklists. But if we want to know if an AI is truly "generally" intelligent, we need to know if its smarts are real and lasting, or just a lucky trick that breaks when things get stressful. This paper asks a big question: How do we stop just counting points and start measuring the stability of a machine's mind?

The authors of this paper suggest that the way we are currently testing AI is a bit like judging a car only by how fast it can drive in a straight line on a sunny day. It might look fast, but that doesn't tell us if the engine will hold up on a bumpy road or if the brakes will work when it's raining. They argue that real intelligence—whether in humans or machines—isn't just a list of skills. Instead, it's more like a homeostatic property cluster. That's a fancy way of saying it's a group of abilities that stick together because there are internal mechanisms keeping them connected, even when things get messy or change.

Think of intelligence like a well-built campfire. The "domains" (math, language, logic) are the logs. But the fire itself is the heat and the flame that keeps all the logs burning together. If you just stack the logs (the skills) but don't have the mechanism to keep them burning (the homeostatic part), a little wind (a stressor or a delay) will blow the fire out. The paper suggests that current tests only check if the logs are there, but they ignore whether the fire can survive a gust of wind.

The authors propose two main changes to fix this. First, they say we shouldn't treat every test question as equally important. Just like some logs are denser and burn hotter, some skills are more "central" to keeping the whole fire going. They suggest using a new scoring method called a centrality-prior score. This would weigh the tests based on how much a specific skill helps keep the whole system stable, borrowing ideas from how human intelligence is already studied. It's like grading a student more heavily on their ability to learn new things than on their ability to memorize a single fact.

Second, they want to stop relying on one-time snapshots. Instead, they propose a Cluster Stability Index. This isn't just about getting a high score once; it's about proving the AI can keep that score over time and under pressure. They suggest testing the AI in three specific ways:

  1. Profile Persistence: Does the AI remember how to do things days or weeks later?
  2. Durable Learning: If you teach it something new, does it stick, or does it forget immediately?
  3. Error Correction: If the AI makes a mistake, can it fix it on its own without crashing?

The paper doesn't claim to have built a perfect test yet or that we have already solved the problem of measuring AGI. Instead, the authors suggest that these new methods would be better. They offer a blueprint for labs to try out these ideas. They propose that these tests can be done even if the labs don't know exactly how the AI is built inside (a "black-box" approach). They also make some predictions about what would happen if we used these new tests, suggesting that many current AI systems might look very different—perhaps less "generally" intelligent and more brittle—once we start checking for stability and stress resistance.

In short, the paper argues that we need to move from checking off boxes to checking if the fire stays lit. By weighing the important skills more heavily and testing if the AI can handle delays and mistakes, we might finally get a clearer picture of whether a machine is truly smart or just pretending to be.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →