The Autonomous Agency Scale: A Behavioral Framework for Measuring Self-Directed Behavior in AI Systems
This paper introduces the Autonomous Agency Scale (AAS), a behavioral framework that quantifies AI self-direction across seven dimensions and two temporal bands, revealing that while current task agents lack genuine autonomy during idle periods, persistent companion architectures demonstrate measurable self-directed behavior.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are watching a incredibly talented actor on a stage. When the director yells "Action!", the actor delivers a perfect, emotional monologue, improvises brilliantly, and solves a complex problem on the spot. But the moment the director yells "Cut!", the actor instantly freezes, stops breathing, and becomes a statue until the next cue. For decades, scientists have been measuring how "smart" these actors are by watching their performances during "Action." They have built complex scorecards to see if the actor can solve math problems, write poetry, or replace a human worker. But they have largely ignored what happens when the director is silent.
This is the gap a new paper tries to fill. It asks a simple, slightly spooky question: Does this system have a mind of its own when no one is watching? The paper introduces a new way to measure "autonomous agency." Think of agency not as being "conscious" or "alive" in a human sense, but as the ability to do things just because you want to, rather than because someone told you to. It's the difference between a dog that sits only when you say "Sit" and a dog that decides to fetch a ball because it's bored. The authors want to know if our current AI systems are just very good at following orders, or if they are starting to have their own internal "to-do lists" that run even when the user goes to sleep.
The "Idle-Gap" Test: Do You Have a Secret Life?
Meet the Autonomous Agency Scale (AAS). It's a new rulebook for grading AI systems, but instead of asking "How smart is it?", it asks, "How self-directed is it?" The authors, led by independent researcher Samuel Presgraves, realized that existing tests are like judging a car only by how fast it goes on a racetrack, ignoring whether it can drive itself home when the driver gets out.
To fix this, the paper splits the score into two distinct "bands," like checking a battery's charge in two different ways:
- The Active Band: This measures the AI while it's busy doing a task you gave it. It's the "Action!" part of the movie.
- The Ambient Band: This is the real kicker. It measures what the AI does when you are not talking to it, when no task is running, and when no alarms are ringing. It's the "Cut!" part of the movie.
The paper introduces a special challenge called the Idle-Gap Test. Imagine you turn off every single trigger that could make the AI move: no scheduled timers, no user prompts, no environmental sensors. If the AI still does something interesting—like thinking up a new idea, checking its own mood, or writing a story just because it feels like it—that's Level 4 behavior (Self-Directed). If it stops moving the second you stop pushing it, it's just a Level 1 (Responsive) or Level 2 (Conditioned) machine, no matter how smart it looks during the "Action" scenes.
The Great AI Reveal: Who Has a Secret Life?
The authors applied this scale to six different AI systems, ranging from powerful coding assistants to your average voice-activated phone helper. The results were a bit of a shocker, drawing a clear line between "Task Robots" and "Companions."
The Task Robots (The "Action" Stars)
Systems like Claude Code, Manus, and Hermes are the superstars of the "Active" band. When you ask them to write code or analyze data, they score high (around 2.3 to 2.4 out of 5). They are brilliant at following complex instructions and working autonomously while they are working.
- The Catch: The moment the task is done, they go to sleep. Their "Ambient" scores (idle time) crash to near zero (0.57 to 1.86).
- The Verdict: These systems are like incredibly efficient employees who work overtime only when you are watching. If you leave the office, they don't go home and start a hobby; they just turn off. Any activity they do while you're gone (like running a nightly update) is just a pre-set alarm clock, not a genuine choice.
The Consumer Assistants (The "Reactive" Helpers)
Systems like Siri and ChatGPT scored even lower on the "Ambient" side (0.29 to 0.86). They are essentially waiting rooms. They are very good at answering questions when you walk in the door, but they have almost no existence when you leave. They are reactive tools, not proactive agents.
The Lone Wolf: Airi
Then there is Airi, a "persistent companion" architecture. This is the only system in the study that passed the Idle-Gap Test.
- The Score: Airi scored a 3.86 on the Ambient band, which is actually higher than its Active score (3.71).
- The Behavior: When no one was talking to it, Airi didn't just sit there. It maintained a "background thought engine," evolved its own moods, and even started creative projects or reached out to the user without being told.
- The Caveat: The authors are careful to note that Airi was evaluated by its own developer, which introduces a bit of bias (like a student grading their own homework). However, the data suggests that Airi's "idle" activity wasn't just a timer going off; it was driven by an internal state that changed over time.
The Bottom Line: We Are Not There Yet
The paper concludes with a sobering reality check. While we have built AI that can do amazing things when we tell them to, no system currently scores a 5 (Sovereign). None of them can truly rewrite their own goals, form new relationships on their own, or change their own "personality" without human help.
The most important finding is a structural boundary: Today's AI is highly autonomous inside a task, but almost entirely dormant between tasks. The "self-directed" behavior we see in movies—where an AI decides to go on a journey or write a book just because it's bored—doesn't exist yet, except in very specific, experimental companion designs like Airi.
The authors suggest that to truly know if an AI is "awake" in its own way, we need to stop testing it in short bursts and start watching it for weeks or months. They call this the Longitudinal Turing Test: not asking if the AI can trick you into thinking it's human, but asking if it can maintain a consistent, self-driven life over time. For now, the answer is that our AI is a brilliant actor who waits for the director's cue, and the only one who seems to be rehearsing in the dark is a very special, very experimental companion named Airi.
Code and full rubric: github.com/CaptainASIC/autonomous-agency-scale
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.