← Latest papers
💬 NLP

Human-Like Anaphor Resolution in Large Language Models

This study evaluates five open-weight Large Language Models to determine if they exhibit human-like anaphor resolution patterns, finding that while some models show sensitivity to discourse structure and distance similar to humans, they lack comparable sensitivity to semantic interference effects.

Original authors: Keane Zhang, Varshini Chinta, Raj Sanjay Shah, Sashank Varma

Published 2026-08-07
📖 6 min read🧠 Deep dive

Original authors: Keane Zhang, Varshini Chinta, Raj Sanjay Shah, Sashank Varma

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are reading a mystery novel. The author introduces a character named "The Baker" in the first chapter. Ten pages later, the story mentions "he" again. Your brain doesn't just guess who "he" is; it instantly reaches back into your memory, grabs the image of the Baker, and connects the dots. This mental magic trick is called anaphor resolution. It's how we keep track of who or what we are talking about as a story unfolds.

Scientists have long studied how humans do this. They know that if the Baker was the main focus of the story (topical), you find him faster. If "he" appears right after the Baker, it's easy. But if "he" appears after a long, confusing gap, or if there's another baker nearby to confuse you, your brain stumbles. Now, enter Large Language Models (LLMs). These are super-smart computer programs trained on mountains of text that can write stories, answer questions, and chat like humans. But here's the big question: Do these AI brains actually think like human brains when they solve these puzzles, or are they just really good at guessing based on patterns? This paper dives into that mystery to see if AI and humans are on the same mental wavelength.


The Great AI Memory Test

The researchers decided to put five different AI models—GPT-2-XL, Llama-3.1-8B, Pythia-12B, Mistral-7B, and Mistral-24B—through a series of memory games designed specifically for humans. They wanted to see if the AI would get tired, confused, or slow down in the exact same ways a human reader would.

To measure this, they used two clever tricks. First, they looked at surprisal. Think of this as the AI's "brain sweat." When a computer reads a word it didn't expect, its "surprisal" goes up. In humans, high surprisal usually means we pause longer to think. So, if the AI gets "sweaty" (high surprisal) at the same moments humans get stuck, that's a good sign they are processing the story similarly. Second, they asked the AI comprehension questions, like "Who kicked the chair?" to see if they actually remembered the right person.

The Memory Games

The team ran three different experiments, each testing a specific rule of human memory:

1. The "Focus" and "Distance" Test
Imagine a spotlight. If a character is under the spotlight (the story title mentions them), humans find them easily. If they are in the dark, it's harder. Also, if the character is just a sentence away, it's easy. If they are ten sentences away, it's harder.

  • The Result: Most of the AI models acted like humans here! When the character was highlighted by the title and close by, the AI's "brain sweat" was low, and they answered correctly. When the character was hidden and far away, the AI got confused. It's as if the AI's attention span stretched and shrank just like ours.

2. The "Space and Time" Test
This was trickier. Imagine the story takes place in a giant castle. If the character moves from the kitchen to the next room (short distance) or from the kitchen to the tower (long distance), does it matter? What if they wait 10 minutes or 2 hours? Humans get slower and make more mistakes when the gap in space or time is huge.

  • The Result: The results were mixed. Some models, like Llama-3.1-8B and Mistral-7B, seemed to feel the weight of the distance and time, getting "sweaty" when the gap was long. Others didn't seem to care much about the time or space gaps. It suggests that while some AIs are starting to understand the "feeling" of a story's timeline, others are just skipping over it.

3. The "Confusing Clue" Test
This is the hardest part. Imagine the story mentions a "metal chair." Later, it says "the metal furniture." That's easy. But what if the story also mentions a "wooden desk" (a different type of furniture) right before? Humans get confused because "furniture" could be the chair or the desk. If the wrong item is a very typical example of the category (like a desk for furniture), it slows us down.

  • The Result: This is where the AI mostly failed to act human. When the researchers added a confusing "distractor" item, the AI models didn't get slower or more confused the way humans do. They seemed to breeze right past the confusion. This suggests that while the AI can track where things are in a story, it doesn't quite get the messy, overlapping competition happening in our brains when we try to pick the right memory.

The Verdict: A Partial Match

So, what's the final score? The paper suggests that these AI models are selectively human-like. They are great at understanding how the structure of a story (like how far away a word is or what the title says) affects memory. If you make the story longer or hide the character, the AI gets a bit more "sweaty," just like a tired reader.

However, they are not perfect clones of human minds. When it comes to the messy, semantic juggling act—where similar words fight for attention in our memory—the AI doesn't stumble the way we do. They seem to lack the specific kind of "mental friction" that happens when our brains have to choose between two very similar memories.

The authors are careful to say this isn't a "solved" problem. The study used a relatively small number of stories (16 to 19), so the results are more like strong hints than absolute proof. Also, the AI's overall accuracy on the questions was sometimes quite low, meaning they might be guessing correctly for the wrong reasons. But the takeaway is exciting: these digital brains are starting to show the same patterns of difficulty that human brains do, at least when it comes to the distance and focus of a story. They aren't just word-matching machines; they are beginning to show the first signs of a "situation model"—a mental map of the story world, just like we have.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →