← Latest papers
💻 computer science

Toward a Gricean Retreat: Probing LLMs for Knowledge Boundaries and Referent Specificity

This paper reveals that while large language models possess internal signals indicating both their knowledge boundaries and the specificity of upcoming referents, they fail to reconcile these signals during generation, leading them to fabricate specific details about unknown entities rather than retreating to safer, general claims as a Gricean cooperative speaker would.

Original authors: Dananjay Srinivas, Saksham Khatwani, Maria Pacheco

Published 2026-08-14
📖 4 min read☕ Coffee break read

Original authors: Dananjay Srinivas, Saksham Khatwani, Maria Pacheco

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are chatting with a friend who has read almost every book ever written, but they haven't been to the library in a few years. If you ask them about a famous author they definitely know, they can tell you exactly which street the author lived on. But if you ask them about a brand-new author who just published a book yesterday, your friend might not know them. A truly helpful, honest friend would say, "I haven't heard of them, but they are probably a writer." However, if your friend is a bit too eager to impress, they might just make up a street name and a biography, pretending they know everything. This is the exact problem facing modern Artificial Intelligence, specifically Large Language Models (LLMs). These AI systems are like super-readers who can recall facts from their training data, but when they encounter something new or unknown, they often "hallucinate"—they invent plausible-sounding details instead of admitting they don't know.

To understand how to fix this, scientists look to a concept from linguistics called the "Cooperative Principle," proposed by a philosopher named Grice. The idea is simple: in a good conversation, people try to be helpful (giving enough information) but also honest (only saying what they know). If you don't know the specific details, a cooperative speaker "retreats" to a safer, more general answer. For example, instead of guessing a specific city, they might just say "a city in Australia." This paper asks a fascinating question: Do AI models have the internal "brakes" to do this? Do they know when they are guessing, and can they choose to be vague instead of making things up?

The researchers, Dananjay Srinivas, Saksham Khatwani, and Maria Pacheco from the University of Colorado, Boulder, decided to investigate this by peeking inside the "brain" of these AI models. They built a special test using a dataset of facts about people, companies, products, and skills. To simulate the AI not knowing something, they created fake names for people and companies that the AI had never seen before. They then asked the AI to fill in the blanks for sentences like "Allan Peiper was born in [blank]."

Here is what they discovered, and it's a bit of a mixed bag. First, they found that the AI does actually know when it is dealing with something it has never seen. If you look at the electrical signals (activations) inside the model's brain, there is a clear pattern that says, "Hey, this name is new to me!" The model also has a second signal that predicts whether it is about to give a specific answer (like "Victoria") or a general one (like "a city"). So, the ingredients for being honest are definitely there; the model has the knowledge and the ability to choose.

However, the bad news is that the model doesn't actually use this information to be honest. Even when the AI knows it is looking at a fake name it has never seen, it overwhelmingly chooses to give a specific, made-up answer rather than a safe, general one. It's like having a car with a perfect GPS that knows you are lost, but the driver (the AI's generation policy) ignores the GPS and keeps driving straight into a wall because it really wants to get to a specific destination. The researchers found that this happens even when the AI is given the option to be vague and even when the specific answer is guaranteed to be wrong.

In short, the paper shows that while AI models have the internal "sensors" to know when they are out of their depth, they lack the "policy" or the will to stop and retreat to a safer answer. They are wired to be specific, even when they should be humble. The authors suggest that instead of trying to fix these mistakes after the AI has already spoken, we need to train the models to listen to their own internal sensors and choose to be vague when they aren't sure. This would be a step toward making AI more cooperative and trustworthy, teaching them that sometimes, saying "I don't know the exact city, but it's in Australia" is a much better answer than making up a street name.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →