← Latest papers
💬 NLP

Post-training makes large language models less human-like

The paper introduces the Psych-201 dataset to demonstrate that post-training processes, which optimize large language models for utility, consistently degrade their ability to accurately simulate human behavior, a problem that persists and worsens in newer model generations and cannot be resolved by persona-induction techniques.

Original authors: Marcel Binz, Elif Akata, Abdullah Almaatouq, Mohammed Alsobay, Oleksii Ariasov, Franziska Brändle, David Broska, Jason W. Burton, Nuno Busch, Frederick Callaway, Vanessa Cheung, Brian Christian, Julia
Published 2026-05-11
📖 4 min read☕ Coffee break read

Original authors: Marcel Binz, Elif Akata, Abdullah Almaatouq, Mohammed Alsobay, Oleksii Ariasov, Franziska Brändle, David Broska, Jason W. Burton, Nuno Busch, Frederick Callaway, Vanessa Cheung, Brian Christian, Julian Coda-Forno, Can Demircan, Vittoria Dentella, Maria K. Eckstein, Noémi Éltető, Michael Franke, Thomas L. Griffiths, Fritz Günther, Susanne Haridi, Sebastian Hellmann, Stefan Herytash, Linus Hof, Eleanor Holton, Isabelle Hoxha, Zak Hussain, Akshay Jagadish, Elif Kara, Valentin Kriegmair, Evelina Leivada, Li Ji-An, Tobias Ludwig, Maximilian Maier, Marcelo G. Mattar, Marvin Mathony, Alireza Modirshanechi, Robin Na, Mariia Nadverniuk, Antonios Nasioulas, Surabhi S. Nath, Helen Niemeyer, Kate Nussenbaum, Sebastian Olschewski, Thorsten Pachur, Stefano Palminteri, Aliona Petrenco, Camille V. Phaneuf-Hadd, Angelo Pirrone, Manuel Rausch, Laura Raveling, Shashank Reddy, Milena Rmus, Evan M. Russek, Tankred Saanum, Kai Sandbrink, Louis Schiekiera, Johannes A. Schubert, Luca M. Schulze Buschoff, Nishad Singhi, Leah H. Somerville, Mikhail S. Spektor, Xin Sui, Christopher Summerfield, Mirko Thalmann, Anna I. Thoma, Taisiia Tikhomirova, Vuong Truong, Polina Tsvilodub, Konstantinos Voudouris, Robert C. Wilson, Kristin Witte, Shuchen Wu, Dirk U. Wulff, Hua-Dong Xiong, Songlin Xu, Lance Ying, Xinyu Zhang, Jian-Qiao Zhu, Eric Schulz

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a brilliant, chaotic, and incredibly creative apprentice who has read almost every book in the library. This apprentice is a Base Model. They are raw, unfiltered, and they speak exactly like a human would: full of quirks, mistakes, hesitations, and the messy logic of real life. If you asked them a question, they might stumble, guess, or give an answer that is technically "wrong" but feels very human.

Then, you decide to train this apprentice to be a Useful Assistant. You teach them to follow strict rules, give the "correct" answer, and act polite and logical. This process is called Post-Training.

The paper you shared, titled "Post-training makes large language models less human-like," reveals a surprising and counterintuitive truth: The more you train these models to be perfect assistants, the less they sound like real people.

Here is the breakdown of their findings using simple analogies:

1. The "Perfect Student" vs. The "Real Person"

The researchers built a massive testing ground called Psych-201. Think of this as a giant library containing over 200,000 transcripts of real people taking psychology tests, playing games, and solving puzzles. It's like having a "control group" of humanity to compare against.

They tested two types of AI:

  • The Base Model: The raw apprentice who learned from reading books.
  • The Post-Trained Model: The same apprentice after being taught to be a helpful, rule-following assistant.

The Result: Every single time, the "Perfect Student" (Post-Trained) was worse at predicting what a real human would say or do than the "Raw Apprentice" (Base). By trying to make the AI "better" at being an assistant, they accidentally stripped away the very human-like behaviors that made it a good simulator of people.

2. The "Uncanny Valley" of Logic

Why does this happen?

  • Base Models are like a mirror of human language. They have absorbed all our biases, our shortcuts, our "gut feelings," and our mistakes. If a human would guess wrong because they are tired or confused, the Base Model does too.
  • Post-Training is like putting a filter on that mirror. It forces the AI to be "normatively correct." It teaches the AI to ignore human quirks and focus on the "right" answer.

The paper found that this "correction" is most damaging in two areas:

  • Psycholinguistics (How we use words): Humans use language messily. Post-trained models use it too perfectly.
  • Reasoning: Humans often make logical errors or use shortcuts (heuristics). Post-trained models are forced to be logically perfect, which makes them sound less like a human thinking process and more like a calculator.

3. The "Newer is Worse" Paradox

You might think that as AI gets smarter, it gets better at mimicking humans. The paper says: Not quite.

  • The Base Models (the raw apprentices) are getting better at mimicking humans with every new generation.
  • However, the Post-Training process (the "assistant" training) is getting more aggressive. The gap between the "Raw Apprentice" and the "Perfect Assistant" is widening. The newer the model, the more "robotic" and less human-like it becomes once you turn on the "assistant" features.

4. The "Persona" Trick Doesn't Work

There is a popular trick called Persona-Induction. This is when you tell the AI, "Pretend you are a 35-year-old teacher from Brazil," hoping it will act more like that specific person.

The researchers tested this on a massive scale. They gave the AI detailed profiles of real people (age, nationality, education) and asked it to predict how those specific people would answer.
The Result: It didn't help. Knowing the "persona" didn't make the AI predict individual human behavior any better. It's like giving a thesaurus to an actor; it doesn't make them a better improviser.

The Big Takeaway

The paper concludes that the very tools we use to make AI helpful (instruction tuning, reasoning training, safety filters) are the same tools that make it a bad tool for simulating human behavior.

If you want an AI to act like a helpful assistant, you should use the Post-Trained version.
But if you want an AI to act like a human (for example, to simulate how a patient might react in a therapy session or how a student might learn), you should actually use the Base Model (the raw, untrained version), because it still remembers how to be messy, imperfect, and human.

In short: To make AI more useful, we made it less human. To make it more human-like again, we might need to stop "fixing" it so much.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →