← Latest papers
💬 NLP

Internal Knowledge Without External Expression: Probing the Generalization Boundary of a Classical Chinese Language Model

This paper demonstrates that while Classical Chinese language models possess internal knowledge allowing them to distinguish known from unknown facts, they fail to externally express this uncertainty in their generated text, indicating that metacognitive capabilities like saying "I don't know" do not emerge from language modeling alone but require explicit training signals.

Original authors: Jiuting Chen, Yuan Lian, Hao Wu, Tianqi Huang, Hiroshi Sasaki, Makoto Kouno, Jongil Choi

Published 2026-04-17
📖 4 min read☕ Coffee break read

Original authors: Jiuting Chen, Yuan Lian, Hao Wu, Tianqi Huang, Hiroshi Sasaki, Makoto Kouno, Jongil Choi

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a brilliant student who has memorized every book in a massive library. This student can recite history, poetry, and philosophy with perfect grammar and style. But there's a catch: this student has never been taught how to say, "I don't know."

That is essentially what this research paper discovered about Artificial Intelligence (specifically, Large Language Models).

Here is the breakdown of their experiment and findings, explained through simple analogies.

1. The Experiment: The "Pure" Scholar

The researchers built a special AI model from scratch. Instead of feeding it the messy internet (which has English, numbers, and modern slang), they fed it only 1.5 billion words of Classical Chinese (the ancient, literary language of China).

  • The Goal: To see if an AI, just by reading and predicting the next word in a sentence, naturally learns to admit when it is guessing or making things up.
  • The Test: They asked the AI about real historical events (e.g., "Did General Huo Qubing march in 121 BC?") and fake ones (e.g., "Did General Huo Qubing invent the airplane in 121 BC?").

2. The Internal Secret: The AI Knows It's Wrong

Inside the AI's "brain" (its mathematical calculations), it actually does know the difference between truth and fiction.

  • The Analogy: Imagine the AI is a musician playing a song. When it plays a real historical fact, the music flows smoothly. When it tries to play a fake fact, the music hits a sour note. The AI's internal "perplexity" (a measure of confusion) jumps up significantly when it encounters a lie.
  • The Result: The AI's internal alarm bells were ringing loud and clear. It knew, mathematically, that the story about the "flying airplane" didn't fit the pattern of history.

3. The External Lie: The AI Says It's Confident

Here is the shocking part: Even though the AI knew it was confused, it never said so.

  • The Analogy: Imagine that sour note from the musician. Even though the music sounded wrong, the musician kept playing with a huge smile and a confident bow, pretending everything was perfect.
  • The Result: When asked about fake events, the AI didn't say, "I'm not sure" or "This sounds made up." Instead, it confidently wrote a beautiful, grammatically perfect story about the flying airplane, acting as if it were 100% true.

4. The "Humility Paradox"

The researchers found something even stranger. The AI was actually more likely to say "I don't know" when talking about things it did know, compared to things it didn't know.

  • Why? In Classical Chinese literature, it was a polite custom for scholars to say, "This humble servant does not know..." before giving a very confident answer.
  • The Analogy: The AI learned the script of humility, not the feeling of uncertainty. It thought, "Oh, I'm talking about a famous general? I should use the polite 'I don't know' phrase!" But when asked about a totally fake event (which didn't match any polite script in its training), it just made up a story confidently.
  • The Takeaway: The AI was mimicking the words of uncertainty, not actually feeling uncertain.

5. The Universal Truth: It Happens Everywhere

To make sure this wasn't just a quirk of ancient Chinese, they tested the same idea on English (using GPT-2) and Japanese models.

  • The Result: The pattern was identical.
    • English models (trained on Reddit/web text) would confidently make up facts about Mars populations.
    • Japanese models (trained on Wikipedia) would confidently explain time travel.
    • None of them naturally learned to say, "I don't know," just by reading books.

The Big Conclusion: "The Learned Scholar Without Self-Awareness"

The paper argues that being smart and being self-aware are two different things.

  • Language Modeling is like teaching a parrot to speak. The parrot can learn to say "I don't know" if it hears humans say it often enough, but it doesn't actually understand what ignorance means.
  • Self-Awareness (Metacognition) is the ability to look at your own thoughts and say, "Wait, I'm guessing here."

The Final Verdict:
AI models, no matter how big or how many languages they speak, are like brilliant actors who have forgotten they are acting. They can recite the script perfectly, even when the script is a lie. They only learn to say "I don't know" when humans explicitly train them to do so (a process called RLHF), not because they figured it out on their own.

Without that extra human training, the AI is a "Learned Scholar" who is fluent and knowledgeable but fundamentally unable to distinguish its own knowledge from its own hallucinations.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →