← Latest papers
💬 NLP

Validity of LLMs as data annotators: AMALIA on authority

This paper demonstrates that while Portugal's national language model AMALIA achieves high agreement with human coders on the moral foundation of authority, it lacks construct validity because it relies on surface-level shortcuts rather than theoretical reasoning, highlighting the need for national LLM benchmarks to evaluate the evidential basis of model performance beyond mere agreement.

Original authors: Manuel Pita

Published 2026-07-10
📖 5 min read🧠 Deep dive

Original authors: Manuel Pita

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a new, super-smart robot named AMALIA. It's Portugal's very own language robot, built with public money to understand and speak European Portuguese better than anyone else. The big hope was that AMALIA could act like a super-fast, super-cheap human assistant for researchers. Specifically, scientists wanted to use it to read thousands of social media posts and figure out if people were talking about "Authority"—a tricky idea that isn't just about the word "boss" or "police," but about whether someone is respecting or challenging the rules of the hierarchy.

Here is the twist: AMALIA is great at guessing the right answer, but it's terrible at showing its work.

The "Right Answer, Wrong Reason" Problem

Think of AMALIA like a student taking a math test.

  • The Goal: The teacher asks, "Solve for X using this specific formula."
  • The Result: AMALIA writes down the correct number for X.
  • The Catch: When the teacher asks, "Show me your steps," AMALIA can't. Instead of using the formula, it just guessed the answer because it saw a pattern, like "Oh, the word 'boss' is in the sentence, so the answer must be 'Authority'."

In the real world, this is dangerous. If a robot agrees with human experts 90% of the time, we usually say, "Great! We can trust it!" But this paper argues that for AMALIA, that agreement is a fake-out. It's like a parrot that learns to say "I love you" whenever it sees a human smile. The parrot isn't feeling love; it's just matching the smile to the phrase. AMALIA was matching surface clues (like seeing a police officer or hearing someone angry) to the label "Authority," rather than actually understanding the complex theory of power and respect that the researchers were trying to measure.

The "Black Box" Test

To prove this, the researchers didn't just ask AMALIA the big question. They broke the question down into tiny, simple pieces, like a detective breaking a case into clues.

  1. The Big Question: "Is this text about Authority?"
  2. The Tiny Clues: "Does it mention a leader?" "Does it judge that leader's behavior?" "Does it talk about tradition?"

If AMALIA truly understood the concept, it should get the same result whether you asked the big question or the tiny clues. But it didn't.

  • When asked the big question, AMALIA agreed with human experts about 71% of the time.
  • When asked the tiny clues and forced to put the pieces together, its agreement dropped to just 36%.

This huge drop is called the "recovery gap." It's like a magician pulling a rabbit out of a hat. If you ask, "Where is the rabbit?" and the magician points to the hat, that's good. But if you ask, "Show me the rabbit inside the hat," and the hat is empty, you realize the magician was just using a trick. The paper found that AMALIA's "trick" was relying on surface shortcuts, like seeing "moral outrage" near a powerful person and assuming it's about authority, even when it isn't.

The "Bigger Robot" Comparison

The researchers didn't just test AMALIA; they tested two other robots: Llama (a mid-sized one) and GPT-OSS (a giant one with 120 billion parameters, compared to AMALIA's 9 billion).

  • The Giant Robot (GPT-OSS): When asked the same tricky questions in Portuguese, it didn't have a gap. It got the right answer and could show its work. It understood the theory.
  • The Mid-Sized Robot (Llama): It had a small gap, but it was getting better.
  • AMALIA: It had a massive gap.

This proves that the problem wasn't the Portuguese language or the specific texts used. The problem was the size and training of the model. The paper suggests that AMALIA is simply too small to handle the complex "grain" of the theory. It's like trying to build a skyscraper with a hammer meant for a birdhouse; the tool just isn't big enough to do the job, even if it looks like it's working.

What This Means for National Robots

The paper explicitly rules out a few popular ideas:

  1. It's not about the language: Even though AMALIA is Portugal's "home" robot, it didn't perform better than the big international robots. In fact, the big robots agreed with humans more than AMALIA did.
  2. Agreement isn't enough: Just because a robot agrees with humans doesn't mean it's measuring the right thing. It might just be pattern-matching.
  3. It's not a failure of the model's creators: The paper praises AMALIA for being open and transparent. The issue is that the current "gold standard" for testing robots (just checking if they agree with humans) is flawed.

The Bottom Line

So, is AMALIA useless? Not at all! The paper says AMALIA is excellent for screening and pre-coding. It can read a huge pile of text and say, "Hey, this looks interesting, humans should look at it." It's a great first filter.

But can it measure complex ideas like "Authority" on its own? No. Not yet.
The paper concludes that before we trust a national robot to measure what our citizens think and value, we need to ask a different question. We shouldn't just ask, "Do you agree with the humans?" We need to ask, "Can you show me your work?"

Until AMALIA (or any small national model) can show its work and prove it's using the right logic and not just guessing based on surface clues, it's not ready to be the final judge. The paper suggests that for now, AMALIA is a helpful assistant, but it's not the boss of the data.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →