← Latest papers
💬 NLP

Do LLMs Adhere to Label Definitions? Examining Their Receptivity to External Label Definitions

This paper investigates whether Large Language Models truly integrate external label definitions or rely on internal knowledge, finding that while such definitions can improve performance, models often default to their parametric representations, particularly in general tasks, with domain-specific tasks benefiting more from explicit guidance.

Original authors: Seyedali Mohammadi, Bhaskara Hanuma Vedula, Hemank Lamba, Edward Raff, Ponnurangam Kumaraguru, Francis Ferraro, Manas Gaur

Published 2026-02-26
📖 5 min read🧠 Deep dive

Original authors: Seyedali Mohammadi, Bhaskara Hanuma Vedula, Hemank Lamba, Edward Raff, Ponnurangam Kumaraguru, Francis Ferraro, Manas Gaur

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are hiring a very smart, well-read assistant (an AI) to help you sort mail. You have a stack of letters, and you need to decide which ones are "Junk," which are "Important," and which are "Personal."

You give your assistant a rulebook (the Label Definitions) to help them decide. But here's the twist: Does the assistant actually read your rulebook, or do they just ignore it and rely on what they already know from their years of training?

This paper is a scientific experiment to answer exactly that question. The researchers tested four different AI models (like GPT-4, LLaMA-3, and others) to see if they truly listen to new instructions or if they stubbornly stick to their old habits.

Here is the breakdown of their findings, using some everyday analogies:

1. The "Swapped Dictionary" Test (Knowledge Conflict)

The researchers tried a trick: they took the definitions of the labels and swapped them around.

  • Normal: They told the AI, "Entailment means A proves B."
  • The Trick: They told the AI, "Entailment means A contradicts B" (the exact opposite).

The Result:

  • The "Smart" Assistant (GPT-4): When the rules were swapped, GPT-4 got confused. It realized, "Wait, this rulebook makes no sense compared to what I know," and it often refused to answer. It was like a librarian who sees a book with the wrong title and says, "I can't shelve this; the catalog is broken."
  • The "Obedient" Assistant (Smaller models like LLaMA-3 or Mistral): These models often followed the new (wrong) rules blindly. If you told them "Red means Blue," they would call a red apple "Blue." They were so eager to follow your instructions that they ignored their own common sense.

The Takeaway: Smaller models are like eager interns who will follow your instructions even if they are wrong. Bigger, smarter models are like senior experts who might question your instructions if they seem contradictory.

2. The "Specialist vs. Generalist" Test

The researchers tested the AI on two types of tasks:

  • General Tasks: Like sorting everyday sentences (e.g., "Is this sentence true based on that one?").
  • Specialized Tasks: Like sorting mental health posts or hate speech.

The Result:

  • General Tasks: The AI didn't need your rulebook much. It already knew how to sort these things from its training. In fact, giving it a rulebook sometimes confused it, like trying to teach a native speaker how to speak their own language using a textbook.
  • Specialized Tasks: This is where the rulebook became a superpower. For niche topics (like mental health), the AI didn't know much beforehand. When you gave it a clear, expert definition, its performance skyrocketed. It was like giving a general contractor a specific blueprint for building a nuclear reactor; without the blueprint, they'd guess; with it, they build it perfectly.

3. The "Explanation vs. Action" Paradox

This is the most surprising part. The researchers asked the AI to not just pick a label, but to explain why it picked it.

  • The Finding: Sometimes, the AI would write a perfect explanation showing it understood the new rules, but then pick the wrong label anyway.
  • The Analogy: Imagine a student taking a math test. They write a beautiful, step-by-step proof showing they understand the formula (the explanation), but then they circle the wrong answer on the bubble sheet (the prediction).
  • Why it matters: This means an AI can sound very confident and logical in its chat, but still make a critical error in its final decision. The "thinking" part and the "deciding" part of the AI are not always working together.

4. The "Custom Tailoring" Effect

The researchers tried different ways of giving the definitions:

  • Static: "Here is the definition for everyone."
  • Dynamic: "Here is a definition tailored specifically to this sentence."

The Result:
Tailored definitions worked best, especially for smaller models. It's like giving a chef a generic recipe vs. a recipe adjusted for the specific ingredients they have in their kitchen right now. The AI performed much better when the instructions felt relevant to the specific job at hand.

Summary: What Should We Do?

If you are using AI for important jobs (like diagnosing a patient or moderating hate speech), here is the advice from the paper:

  1. Don't assume the AI knows everything: For specialized topics, you must provide clear, precise definitions. Don't just say "Sort this"; say "Sort this based on these specific rules."
  2. Smaller models need more help: If you are using a smaller, cheaper AI, it relies heavily on the instructions you give it. Make those instructions perfect.
  3. Big models might ignore you: If you use a massive model, it might ignore your instructions if they clash with what it already knows. You might need to be very careful about how you phrase things.
  4. Check the work: Just because the AI gives you a great explanation doesn't mean its final decision is correct. Always double-check the result.

In a nutshell: LLMs are like students. Sometimes they are eager to learn your new rules (especially in new subjects), and sometimes they are so smart they think they know better than you. The key is knowing which "student" you are talking to and how to give them the right instructions.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →