Reinforcing Human Behavior Simulation via Verbal Feedback
This paper introduces DITTO, a model trained via reinforcement learning with verbal feedback to better simulate human behavior, alongside the SOUL benchmark suite, demonstrating significant performance improvements over base models and surpassing GPT-5.4 on multiple human-like simulation tasks.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: Teaching AI to "Act Human" by Talking to It
Imagine you are teaching a child how to behave at a dinner party.
- The Old Way (Scalar Rewards): You give the child a score of "6 out of 10" after they finish eating. You don't tell them why they got a 6. Did they talk too much? Did they use the wrong fork? Did they spill soup? They have no idea how to improve, so they just guess next time.
- The New Way (Verbal Feedback): You sit them down and say, "You did great with the napkin, but you talked over your grandmother, and that was rude. Next time, wait for her to finish speaking."
This paper argues that current AI models (LLMs) are being trained like the child getting a score of "6." They are good at math and coding because those things have clear right/wrong answers (like a test score). But when we ask AI to simulate human behavior—like acting as a patient, a student, or a friend—simple scores don't work. Humans learn social norms through words, explanations, and corrections, not numbers.
The authors built a new system called DITTO that teaches AI to act human by using this "verbal feedback" method.
How DITTO Works: The "Draft and Polish" Game
Think of DITTO as a creative writing workshop where the AI is both the student and the editor. Here is the process:
- The First Draft (The Student): The AI tries to answer a prompt (e.g., "Act like a grumpy teacher"). It writes a response.
- The Critique (The Judge): A powerful AI "judge" reads the draft and gives verbal feedback. It doesn't just say "Good job." It says, "You were too polite. A grumpy teacher would be more impatient. Also, you forgot to mention the homework."
- The Second Draft (The Teacher): The AI takes that specific feedback and writes a new, improved version of the response.
- The Lesson (The Magic): The system trains the AI to look at the original prompt and immediately generate the improved version, without needing the feedback text anymore. It's like the student reading the teacher's notes, fixing the essay, and then memorizing the feeling of how to write a good essay so they don't need the notes next time.
The Result: The AI learns to "internalize" the advice. When you talk to it later, it acts more human because it has learned why certain behaviors are better, not just that they are better.
The Playground: SOUL (Simulation Gym)
To test if this actually works, the researchers built a giant gym called SOUL (Simulation gym Of hUman-Like behavior).
Imagine a gym with 10 different stations, each testing a different type of "human" skill:
- Theory of Mind: Can the AI understand that someone else knows something it doesn't? (Like a game of hide-and-seek).
- Role Play: Can it act like a specific character from a book without breaking character?
- Social Skills: Can it navigate a conversation to get a goal (like asking for a raise) without being rude?
- User Simulation: Can it pretend to be a confused customer talking to a help desk?
They found that existing AI models often failed these tests. They sounded too robotic, too perfect, or too similar to each other.
The Results: DITTO Wins the Gym
The researchers tested their new model (DITTO) against the best existing models (including giant ones from big tech companies).
- The Score: DITTO improved by 36% compared to the basic model.
- The Showdown: DITTO beat the top-tier "GPT-5.4" model on 6 out of the 10 challenges.
- Where it Shined: It was especially good at tasks that require nuance, like social skills and role-playing. It learned to keep secrets better and handle complex conversations without getting "stuck" or sounding fake.
Why This Matters (According to the Paper)
The paper claims that to make AI truly useful for simulating people (like in therapy bots, educational tutors, or customer service training), we have to stop treating them like math students who just need a score. We have to treat them like social beings who need explanations.
By using verbal feedback (words) instead of just scalar rewards (numbers), the AI learns the texture of human behavior. It learns that being "rude" isn't just a low score; it's a specific way of speaking that hurts feelings.
In short: DITTO is an AI that learned to act human by listening to a teacher's detailed critiques, practicing the corrections, and then forgetting the teacher's notes because it finally "got it."
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.