← Latest papers
💬 NLP

Comparing Semantic Navigation in Humans and Large Language Models using Natural Language Processing

This study reveals that human semantic memory retrieval exhibits a more variable and exploratory search pattern than current large language models, as evidenced by higher entropy, larger semantic steps, and broader dispersion in verbal fluency tasks that no temperature tuning configuration could fully replicate.

Original authors: Gabriel Paris-Colombo, Rodrigo M. Cabral-Carvalho, Felipe D. Toro-Hernández

Published 2026-07-15
📖 5 min read🧠 Deep dive

Original authors: Gabriel Paris-Colombo, Rodrigo M. Cabral-Carvalho, Felipe D. Toro-Hernández

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine your brain as a giant, bustling library where every book is a word or a concept. When you play a game like "Name as many animals as you can in one minute," you aren't just pulling books off a shelf randomly; you are navigating through the library. You might start in the "Pets" aisle, grab a few books (cat, dog, hamster), and then suddenly decide to take a giant leap to the "Ocean" aisle to grab a shark. This is your brain's way of exploring.

A team of researchers decided to see if the super-smart AI chatbots we use today (like GPT-4o, Gemini, and Claude) navigate this same library the same way humans do. They put 82 real people and three different AI models to the test, asking them to list animals. To make the AI act more like a person, they even asked the bots to pretend they were different types of people (like a chef or a pilot) before listing the animals. They also tweaked a "randomness knob" (called temperature) on the AI, turning it from 0.0 (super robotic and predictable) to 1.0 (super chaotic and wild).

Here is what they found, and it's a bit of a plot twist.

The Human "Jungle Gym" vs. The AI "Tightrope"

The researchers measured three things to see how the "walk" through the library looked:

  1. Step Size: How big of a jump did you make between words? (e.g., jumping from "cat" to "shark" is a big jump; "cat" to "dog" is a small one).
  2. Surprise Factor (Entropy): How unpredictable were your steps? Did you switch aisles randomly, or was your path very steady?
  3. Wanderlust (Distance to Centroid): How far did you get from the center of the library? Did you stick to the main path, or did you explore the dusty, forgotten corners?

The Big Discovery:
Humans were the ultimate explorers. Our paths were messy, unpredictable, and full of giant leaps. We took larger steps between words, our path was more unpredictable, and we wandered further away from the center of the category. We were like kids on a jungle gym, swinging from bar to bar, sometimes landing far away from where we started.

The AI models, on the other hand, were like tightrope walkers. They stayed much closer to the center of the aisle. Their steps were smaller, their path was more predictable, and they rarely ventured into the "deep woods" of the library. Even when the researchers turned the "randomness knob" up to the maximum (1.0) to make the AI wilder, the bots still couldn't quite match the human style. They were still too careful and too clustered.

The "Temperature" Trap

You might think, "If I just turn the AI's randomness up high enough, it will act exactly like a human!" The paper suggests this isn't true.

The researchers tested the AI at eight different temperature settings (0.0, 0.15, 0.3, 0.45, 0.6, 0.75, 0.9, and 1.0).

  • At some specific settings, the AI did look human on just one test. For example, at a temperature of 0.75, the Claude model took steps that were about the same size as humans.
  • But here's the catch: No single setting made the AI look human in all ways at once. You could get the step size right, but then the AI would become too predictable. Or you could make it wander far, but then it would stop taking big jumps.

It's like trying to tune a radio to a specific station. You might get the volume right, or the clarity right, but you can't get the perfect human voice to come out of the speaker no matter how you twist the dial. The AI's "brain" is built differently, so it can't replicate the full, messy, beautiful dance of human memory.

What This Means for "Robot People"

For a while, scientists hoped we could just use AI as "silicon participants"—basically, robot humans—to test psychology theories. We could ask the AI a question, see what it says, and assume it's thinking like a person.

This paper argues against that idea. It suggests that while AI is great at sounding fluent and smart, it doesn't actually search for information the way we do. The AI is trained to predict the next word based on statistics, like a super-fast guesser who always picks the most likely option. Humans, however, use a mix of strategy and surprise. We have a "working memory" that helps us remember what we've already said and decide when to jump to a totally new idea. The AI, even when told to pretend to be a person, seems to lack this specific kind of "mental gymnastics."

The Bottom Line

The study didn't prove that AI is "broken" or that it can never be useful. Instead, it suggests that current AI models have a systematic bias toward staying safe and close to home. They are excellent at local exploration (staying in the "Pets" aisle) but struggle with the global exploration (jumping to the "Ocean" aisle) that humans do naturally.

So, while we can tweak the AI to act a little bit like a human in specific moments, we haven't found a way to make it fully mimic the complex, wandering, and surprising way our brains navigate the world of words. The AI is a powerful tool, but it's not a perfect copy of a human mind just yet.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →