On the Limits of LLM Adaptability: Impact of Model-Internalized Priors on Annotation Task Performance
This paper demonstrates that Large Language Models' performance in annotation tasks is primarily constrained by the alignment between their internalized priors and task definitions rather than text-level memorization, revealing that nearly two-thirds of zero-shot errors remain resistant to prompt-based correction.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: The "Stubborn Intern"
Imagine you hire a highly educated intern (the AI) to sort a pile of letters. You give them a specific rule: "Mark any letter that is rude as 'Toxic'."
You might assume that if you explain the rule clearly, the intern will follow it perfectly. However, this paper argues that the intern isn't a blank slate. Before you hired them, they spent years reading millions of books, websites, and social media posts. They have already formed their own internal idea of what "rude" means.
The paper asks: If your rule clashes with the intern's internal idea, who wins?
The answer is surprising: The intern's internal idea usually wins. Even if you give them a perfect written rule, they often stick to their own gut feeling, and they will do so with total confidence.
The Three Main Experiments
The researchers tested this using 9 different AI models and 5 different types of "toxicity" datasets (like gaming chats, news comments, and social media). Here are the three things they discovered:
1. The "Memory" Myth vs. The "Concept" Match
The Question: Does the AI do better because it has memorized the specific letters you are asking it to sort, or because it actually understands the concept of "toxicity" the way you do?
The Analogy: Imagine a student taking a test.
- Memorization: The student has seen these exact questions before and remembers the answers.
- Concept Match: The student actually understands the subject matter and can apply the rules to new questions.
The Finding: The researchers found that memorization doesn't help. In fact, if an AI has memorized the text, it might actually do worse because it's just reciting old data instead of thinking.
Instead, performance depends on Definition-Specific Familiarity (DSF). This is a fancy way of measuring: "Does the AI's internal definition of 'toxic' match your written definition?"
- If the AI's internal "rude detector" matches your rule, it does a great job.
- If their internal detector is different (e.g., they think "rude" means "hate speech against a group," but you meant "any mean comment"), it fails, even if it's a very smart model.
2. The "Sticky" Mistake
The Question: If the AI makes a mistake, can you fix it by giving it more examples or better instructions?
The Analogy: Imagine the intern marks a letter as "Safe" when it's actually "Toxic." You say, "No, look at this example! This is toxic!"
- The Finding: It's like trying to unstick gum off a shoe. Most mistakes (about 65%) cannot be fixed.
- The researchers call this "Decision Stickiness." Once the AI makes a decision, it is very hard to change its mind.
- The Confidence Trap: The worst part? The AI is often most stubborn when it is most confident. If the AI says, "I am 99% sure this is safe," and it's wrong, your instructions almost never change its mind. It's like a confident driver who refuses to look at the GPS even when they are clearly going the wrong way.
3. The "Confidence" Trap
The Question: If you give the AI a wrong rule (a "misaligned" definition), will it realize it's confused and lower its confidence score?
The Analogy: Imagine you tell the intern, "Actually, today we are sorting by color, not by rude-ness."
- The Finding: The AI happily follows your new, wrong rule. But here is the scary part: It doesn't get confused. It still says, "I am 90% sure I am doing this right."
- The AI does not lower its confidence score just because the instructions are weird or wrong. It just applies the new rule with the same high confidence it had before.
- The Result: You cannot trust the AI's "confidence score" to tell you if it's following your instructions correctly. It might be following a completely wrong definition, but it will tell you it's 100% sure.
The Takeaway for Humans
If you want to use AI to label or sort data (like finding toxic comments), the paper suggests three simple rules:
- Don't rely on the AI's "fame." Just because a model is huge or has seen a lot of data doesn't mean it will do your specific job well.
- Check the "Concept Match" first. Before you start, check if the AI's internal idea of the task matches your definition. If they don't match, no amount of prompting will fix it.
- Don't trust the confidence meter. If the AI says it's "very confident," it doesn't mean it's right. It might just be confidently following the wrong instructions.
In short: The AI is not a blank slate waiting for your instructions. It comes with its own "brain" and "habits." If your instructions don't align with those habits, the AI will likely ignore you, make mistakes, and confidently insist it's right.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.