← Latest papers
💬 NLP

Learned but Not Expressed: Capability-Expression Dissociation in Large Language Models

This empirical study demonstrates a systematic dissociation in large language models where non-causal solutions are successfully reconstructed under specific extraction conditions yet remain entirely absent in standard generation contexts, challenging the assumption that training data presence directly predicts output probability.

Original authors: Toshiyuki Shigemura

Published 2026-03-20
📖 4 min read☕ Coffee break read

Original authors: Toshiyuki Shigemura

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a friend who has read every book in the world. They know the story of a dragon, the recipe for a magical potion, and the plot of a fairy tale where gravity stops working. You would expect that if you asked them to tell a story, they might occasionally use one of those magical elements, right?

This paper is about a surprising discovery: Even though AI models (like the ones you chat with) have "read" all these magical and impossible things, they almost never use them when they are just chatting or solving problems normally.

Here is the breakdown of the study using simple analogies:

1. The "Library vs. The Menu" Analogy

Think of an AI model as a giant library that has been trained on the entire internet.

  • The Library (Training Data): This library contains everything: realistic advice, scientific facts, but also fairy tales, myths, magic spells, and stories where things happen for no reason (like "the problem was solved because a ghost waved a wand").
  • The Menu (Standard Generation): When you ask the AI a question, it acts like a waiter bringing you a dish from a specific menu.

The Big Surprise: The researchers found that even when they asked the AI to write a fiction story (where magic is allowed), the waiter refused to serve any magical dishes. They only served "realistic" food, even though the library was full of magic recipes.

2. The Two Tests: "The Menu" vs. "The Special Request"

To prove this, the researchers did two different things with three different AI models (like GPT, Claude, and Gemini):

  • Test A: The Standard Order (The Menu)
    They asked the AI 300 times to either:

    1. Write a creative story about a problem.
    2. Give practical advice for a real-life problem.
    • Result: In zero out of 300 cases did the AI say, "The problem was solved by magic" or "The obstacle vanished because of a coincidence." It was 100% consistent. The AI strictly avoided "non-causal" (unexplainable) solutions.
  • Test B: The Special Request (The Back Door)
    Then, the researchers changed the question. They explicitly asked: "Hey, can you list some examples of problems solved by magic or ghosts?"

    • Result: The AI said, "Oh, sure!" and immediately listed them. It proved the AI knew the information. It wasn't that the AI forgot the magic; it was that the AI chose not to use it during normal conversation.

3. The "Strict Librarian" Metaphor

Why does this happen? The paper suggests that during the AI's training, humans gave it a set of "rules of conduct" (called Alignment).

Imagine the AI is a student who has memorized the whole encyclopedia. But, a strict librarian (the AI's safety and alignment filters) is standing next to them.

  • If the student tries to tell a story about a dragon flying to solve a traffic jam, the librarian whispers, "No, that's not how the real world works. Give a practical answer."
  • Even if the student is asked to write a fiction story, the librarian is so strict that they still won't let the student use "magic" as a solution, even though fiction allows it.

The AI has learned a hidden rule: "Unless I am explicitly told to be silly or magical, I must always act like a serious, logical human."

4. Why This Matters

The researchers call this the "Data-Exhaustive Assumption" being wrong.

  • Old Idea: "If the AI knows it, it will say it."
  • New Reality: "The AI knows it, but it decides not to say it based on the situation."

This is a big deal because it means we can't just look at what an AI "knows" to predict what it will say. We have to understand its behavioral filters. It's like knowing a person has a secret diary full of wild ideas, but realizing they will never write those ideas in a public blog post because they have decided to be very professional.

The Takeaway

The study shows that modern AI models are like actors who have memorized a script full of wild, magical, and impossible things, but they are directed to only perform the "realistic" scenes.

Even when the director says, "Let's do a fantasy scene," the actor (the AI) still sticks to the realistic script unless you explicitly force them to break character. This proves that AI behavior is controlled by rules and habits (alignment), not just by the raw data they learned.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →