Evaluating Developmental Cognition Capabilities of LLMs
This paper introduces the Developmental Sentence Completion Test (DSCT) to evaluate LLMs' ability to detect developmental cognitive stages based on Robert Kegan's theory, finding that while models can accurately identify intended stages in simulated personas and show consistent developmental patterns in their own outputs, their agreement with real human responses remains moderate, highlighting that the primary challenge for stage-aware AI lies in the availability of developmental signal within elicited text rather than classifier accuracy alone.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: How Do We "Think" About Thinking?
Imagine you are talking to a very smart robot. Usually, we try to make the robot understand what you want (your goals) or what you know (your facts). But this paper asks a different question: How do you actually make sense of the world?
The authors use a psychological theory by Robert Kegan to look at this. Think of Kegan's theory like a ladder of "meaning-making."
- Lower rungs: You see the world as a series of rules or what others expect of you. (e.g., "I do this because my boss said so.")
- Higher rungs: You build your own internal compass and can step back to look at your own beliefs. (e.g., "I do this because I've analyzed the situation and decided it aligns with my own values, even if my boss disagrees.")
The paper tries to figure out if AI can detect which "rung" of the ladder a person is standing on just by reading their text.
The Problem: The Old Tests Were Too Heavy
Traditionally, to figure out where someone is on this ladder, psychologists have to sit down for a long, deep interview (like a 90-minute therapy session). It's like trying to weigh a whale by lifting it; it's too heavy and slow to do for thousands of people or for testing AI.
There are also shorter tests (sentence completion), but they are often old, private (you have to buy them), or ask very personal questions about sex and family that feel invasive.
The Solution: The "DSCT" (A New, Lighter Test)
The authors created a new, 20-question test called the Developmental Sentence Completion Test (DSCT).
- The Analogy: Imagine a "fill-in-the-blank" game. Instead of asking "What is your favorite color?", it asks things like, "When a promise is broken, I feel..." or "When a person has to choose between two good options, they..."
- The Goal: The answers aren't graded on being "right" or "wrong." They are graded on the structure of the thinking. Does the answer sound like someone following the crowd, or someone with their own internal system?
The Three Experiments: The "Taste Test"
The researchers ran three different "taste tests" to see if AI could read these answers correctly.
1. The "Robot Acting as a Human" Test (Simulated Personas)
- The Setup: They asked a super-smart AI to pretend to be different types of people (from "rule-follower" to "independent thinker") and fill out the test.
- The Result: The AI classifiers (the "graders") were amazingly good at this. When the AI knew exactly what persona it was supposed to be, the other AIs could guess the "thinking style" with near-perfect accuracy.
- The Metaphor: It's like an actor reading a script perfectly. The other actors (the graders) could easily tell exactly which character was being played.
2. The "Real Human" Test
- The Setup: They gave the test to 83 real humans. Then, they asked both human experts and AI to grade the answers.
- The Result: This was messier. The AI and the human experts agreed about half the time on the exact "rung," but they agreed much more often on the general neighborhood (e.g., both said "somewhere between rung 3 and 4").
- The Metaphor: Real humans are messy. Sometimes they write short answers, sometimes long ones, and sometimes they contradict themselves. It's like trying to guess someone's mood by reading a text message they sent while running late; it's harder than reading a script. The AI was a bit more conservative (stingier) than the human experts, rarely giving a "high rung" unless it was sure.
3. The "AI Talking to Itself" Test (Default Responses)
- The Setup: They asked the AI models to take the test without pretending to be anyone. They just answered as themselves.
- The Result: The newer, bigger AI models gave answers that sounded like they were on a higher rung of the ladder than the smaller, older models.
- The Metaphor: Imagine asking a toddler and a college student to fill out the same "fill-in-the-blank" test. The college student's answers naturally sound more complex and self-reflective. The paper found that as AI models get bigger and newer, their "default voice" sounds more like a self-reflective adult.
The Main Takeaways
- AI is good at spotting patterns, but the signal matters: If the text is clean and structured (like the simulated test), AI is great at figuring out the "thinking style." If the text is messy (like real human answers), it's harder, but still possible to get a general idea.
- AI isn't "thinking" like a human: The paper is careful to say that AI doesn't actually have a developmental stage. It's just that the words the AI generates sound like they come from a specific stage.
- The "Voice" of the AI matters: Bigger, newer AI models tend to speak in a way that sounds more "self-authored" (higher on the ladder). This might mean they interpret human users differently than smaller models do.
What This Means for the Future (According to the Paper)
The paper suggests that for AI to be truly personalized, it shouldn't just learn what you like (e.g., "I like sci-fi movies"). It should try to understand how you think (e.g., "Do you need me to give you the rules, or do you want me to challenge your ideas?").
However, the authors warn that we can't just guess this from a random chat. We need structured ways to ask (like this test) to get a reliable reading. You can't tell someone's "thinking ladder" just by glancing at a quick text message; you need a proper conversation to see the structure.
In short: The paper built a new, shorter test to measure how people make sense of the world. They found that AI can read this test well, especially when the answers are clear, but real human answers are messy. Also, the AI models themselves are "growing up" in their writing style, sounding more complex as they get bigger.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.