← Latest papers
💬 NLP

Fine-tuning with Hierarchical Prompting for Robust Propaganda Classification Across Annotation Schemas

This paper introduces a new intent-focused propaganda taxonomy and the Hierarchical Prompting (HiPP) method, demonstrating that fine-tuning language models with this approach significantly improves robustness and performance across varying annotation schemas, particularly for ambiguous, low-agreement datasets.

Original authors: Lukas Stähelin, Veronika Solopova, Max Upravitelev, David Kaplan, Ariana Sahitaj, Premtim Sahitaj, Charlott Jakob, Sebastian Möller, Vera Schmitt

Published 2026-05-14
📖 4 min read☕ Coffee break read

Original authors: Lukas Stähelin, Veronika Solopova, Max Upravitelev, David Kaplan, Ariana Sahitaj, Premtim Sahitaj, Charlott Jakob, Sebastian Möller, Vera Schmitt

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a group of robots to spot propaganda in social media posts. The problem is that propaganda is tricky: it's short, messy, and often depends on why someone wrote it, not just what they wrote.

This paper is like a report card on how well different robot teachers (AI models) learned this task, and it tests three different ways of teaching them.

The Two "Textbooks" (Annotation Schemas)

First, the researchers created two different "textbooks" or rulebooks for the robots to learn from.

  1. The "Technique" Textbook (Sahitaj et al.): This one is like a checklist of specific tricks. It says, "If you see a scary word, that's 'Fear.' If you see a name-calling, that's 'Ad Hominem'." It's very clear-cut, and the human teachers (annotators) agreed on these rules most of the time. It's easy to grade, but it might miss the bigger picture.
  2. The "Intent" Textbook (The New One): This is the researchers' new invention. Instead of just listing tricks, they asked, "What is the goal of this post?" Is the goal to "Shift the Blame"? Is it to "Confuse the Audience"? Is it to "Justify Aggression"?
    • The Catch: This is much harder. It's like asking a student to guess the teacher's motive rather than just spotting a grammar mistake. The human teachers disagreed a lot more on these labels because the motives are fuzzy and subjective.

The Three Teaching Methods

The researchers tested four different robot brains (AI models) using three different teaching strategies:

  1. The "Zero-Shot" Method: You just hand the robot the textbook and say, "Go figure it out." The robots were terrible at this. They guessed randomly, like a student who didn't study.
  2. The "Fine-Tuning" Method: This is the big breakthrough. The researchers showed the robots thousands of examples and let them practice until they got it right.
    • The Result: This was the magic key. It turned weak, confused robots into sharp, competitive ones. The paper claims that without this practice session, the robots simply couldn't learn the task well.
  3. The "Hierarchical Prompting" (HiPP) Method: This is a clever trick. Instead of asking the robot, "What is the main goal?" immediately, they asked it two questions in a row:
    • Step 1: "What specific trick is being used here?"
    • Step 2: "Okay, based on that trick, what is the main goal?"
    • The Analogy: Imagine trying to identify a bird. Instead of guessing the species immediately, you first ask, "Does it have a long beak?" and then, "Based on the long beak, is it a heron?"
    • The Finding: This two-step method worked best after the robots had practiced (fine-tuned), especially on the difficult "Intent" textbook where the rules were messy. It helped the robots stay organized when the answers were confusing.

The Results: Who Won?

  • The Best Robots: The Qwen family of robots (specifically Qwen3-14B) performed the best overall. They were the "honor students."
  • The Runner-Up: The Phi-4 robot was also very strong, beating the GPT-4.1-nano robot consistently.
  • The Lesson on Difficulty: The robots scored higher on the easy "Technique" textbook than on the hard "Intent" textbook. This proves that the new "Intent" rules are indeed more challenging and realistic, even if they are harder to learn.

The "Aha!" Moment (Error Analysis)

The researchers looked at where the robots made mistakes.

  • Before Practice: The robots were "keyword hunters." If they saw words like "War," "Russia," or "Nazi," they immediately screamed "PROPAGANDA!" even if the post was just a news report. They were too sensitive.
  • After Practice: The robots became much smarter. They stopped flagging every news report about war. They only flagged the posts that were actually trying to manipulate people.
  • The Confusion: When the robots did get confused, they didn't pick random answers. They confused similar goals (like mixing up "Shifting Blame" with "Distorting Reality"). This is actually good news; it means they are thinking logically, just getting the fine details wrong, much like a human would.

The Bottom Line

The paper concludes that to build a robot that can really understand propaganda, you can't just give it a rulebook and hope for the best. You have to train it (fine-tune) on real examples. Once trained, using a step-by-step thinking process (HiPP) helps it handle the messy, real-world situations where the "why" behind a message is hard to pin down.

They also released their new "Intent" textbook and the trained robots for others to use, hoping it will help researchers understand not just that propaganda is happening, but what the propagandists are trying to achieve.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →