← Latest papers
💬 NLP

Self-Evaluation Is Already There: Eliciting Latent Judge Calibration in Base LLMs with Minimal Data

This paper demonstrates that base large language models already possess a latent ability to predict external judge scores, which can be effectively unlocked and refined into a transferable self-evaluation skill through the proposed Self-Evaluation Elicitation (SEE) method using minimal data.

Original authors: XiuYu Zhang, Yi Shan, Junfeng Fang, Zhenkai Liang

Published 2026-06-04
📖 2 min read☕ Coffee break read

Original authors: XiuYu Zhang, Yi Shan, Junfeng Fang, Zhenkai Liang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a large language model (LLM) as a talented but slightly shy student who has already read a massive library of books (pre-training). The paper asks: Can this student guess how a strict teacher would grade their essay before the teacher even sees it?

The researchers discovered that the student actually already knows the answer. Even without special training, the model can "feel" how good its own writing is, though its guesses are a bit fuzzy and overconfident.

To fix this, they invented a method called Self-Evaluation Elicitation (SEE). Think of it as a two-step study session that takes very little time:

  1. The Practice Round (RL): The student writes an essay and then immediately tries to grade it. A "teacher" (an external AI judge) also grades it. The student gets points not just for writing a good essay, but for grading themselves accurately.
  2. The Targeted Correction (Distillation): This is the clever part. The researchers take the student's practice essays and say, "You wrote the essay perfectly, so don't change a word of it. But your self-grade was wrong. Let's just erase your grade and write the teacher's correct grade in its place."

By repeating this cycle only 160 times (a tiny amount of data compared to the thousands usually needed), the model learns to predict the teacher's score with high precision.

The Key Takeaway:
The paper argues that the ability to judge quality isn't something we need to teach from scratch. It's like a hidden muscle the model already has; we just need to do a few light exercises to "elicit" (bring out) it. Once trained, the model can reliably predict how good its answers are without needing to ask the teacher for help every time.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →