← Latest papers
💬 NLP

Can Large Language Models Simulate Human Cognition Beyond Behavioral Imitation?

This paper introduces a novel benchmark based on the longitudinal publication trajectories of 217 AI researchers to evaluate whether large language models genuinely simulate individual human cognitive patterns or merely imitate surface behaviors, utilizing a cross-domain temporal-shift setting and a multidimensional alignment metric to assess current models and enhancement techniques.

Original authors: Yuxuan Gu, Lunjun Liu, Xiaocheng Feng, Kun Zhu, Weihong Zhong, Lei Huang, Bing Qin

Published 2026-03-31
📖 5 min read🧠 Deep dive

Original authors: Yuxuan Gu, Lunjun Liu, Xiaocheng Feng, Kun Zhu, Weihong Zhong, Lei Huang, Bing Qin

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot how to think like a specific human genius, say, a famous scientist named Dr. Smith.

You have two ways to do this:

  1. The "Mimic" Approach: You show the robot Dr. Smith's old papers and say, "Write like him." The robot looks at the words, the sentence structure, and the specific topics Dr. Smith used, and it copies them perfectly. It sounds like Dr. Smith, but it's just a very good actor.
  2. The "Internalize" Approach: You try to get the robot to actually understand how Dr. Smith's brain works. You want it to see a new problem and solve it using Dr. Smith's unique logic, his specific way of asking questions, and his deep-seated beliefs about how the world works—even if the new problem is totally different from anything Dr. Smith ever wrote about before.

This paper is a report card on how well our current "super-smart" AI robots (Large Language Models or LLMs) can do Approach #2.

The Big Question

Can AI truly simulate human cognition (the internal thinking process), or is it just imitating behavior (the external output)?

The authors built a giant test to find out. Instead of asking the AI to write a fake email or a poem, they used something much harder: Scientific Research.

The Experiment: The "Time-Travel" Test

The researchers gathered the entire publication history of 217 real AI scientists. They treated each scientist as a unique "mind."

Here is the tricky part of the test:

  • The Training: They showed the AI a scientist's past papers (e.g., papers about "Image Recognition" from 2015–2020).
  • The Test: They asked the AI to solve a brand new problem in a completely different field (e.g., "How to analyze biological data in 2024") that the scientist had never worked on before.

If the AI is just imitating, it will try to use the same tools and writing style from the old papers, even if they don't fit the new problem. It's like a chef who only knows how to make pizza trying to cook a steak dinner because they are "acting like a chef."

If the AI is simulating cognition, it should look at the new problem and say, "Ah, this scientist usually approaches problems by looking at X, Y, and Z. I will apply that same logic to this new steak dinner."

The Results: The "Uncanny Valley" of Thinking

The results were a bit disappointing for the AI, but very revealing for us.

1. The AI is a Master Actor (Behavioral Imitation)
The AI was fantastic at copying the style. It could mimic the vocabulary, the sentence structure, and even the specific technical methods the scientist used in the past. If you just looked at the words, it seemed like the scientist was writing it.

  • Analogy: It's like a parrot that has memorized a Shakespeare play. It sounds exactly like Shakespeare, but it doesn't understand the tragedy or the human emotion behind the words.

2. The AI is a Struggling Philosopher (Cognitive Internalization)
When the test required the AI to apply the scientist's deep thinking patterns to a new, unfamiliar topic, the AI mostly failed.

  • It couldn't figure out the scientist's unique "problem-solving philosophy."
  • It couldn't predict what the scientist would avoid doing (a key part of how experts think).
  • It mostly just grabbed the closest-looking solution from the past and pasted it onto the new problem, even if it didn't make sense.
  • Analogy: If you asked the parrot to write a new play about a modern issue, it would just mash up old lines from Shakespeare. It wouldn't create a new story with a consistent character; it would just be a chaotic collage of old quotes.

The "Magic" vs. The "Math"

The researchers also tested if they could "teach" the AI better using special tricks (like giving it personality profiles or training it specifically on one person).

  • The Verdict: These tricks helped a little bit, but not enough. The AI still relied on statistical guessing (math) rather than true understanding (magic).
  • The AI is currently very good at predicting the next word based on patterns, but it is not good at building a stable, internal model of how a specific human mind works and then using that model to reason through new challenges.

The Takeaway

Think of current AI like a very talented forger.

  • If you give it a blank canvas and ask it to paint a picture in the style of Van Gogh, it can do a stunning job.
  • But if you ask it to think like Van Gogh when faced with a brand new, weird object (like a glowing alien fruit) and decide how to paint it, it will likely just paint the fruit with the same brushstrokes it used for the sunflowers, because it doesn't actually understand why Van Gogh painted the way he did.

In short: Today's AI can perfectly copy the footprints of human thinking, but it hasn't yet learned how to walk the path itself. It's imitating the behavior, but the "soul" of the cognition is still missing.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →