Lingo_Research_Group at SemEval-2026 Task 9: Evaluating Prompt Variants for Polarization Detection
The Lingo_Research_Group paper presents a systematic evaluation of twelve prompt variants using the Gemma3-27B model for SemEval-2026 Task 9, demonstrating that while prompt-based approaches effectively detect coarse-grained polarization across 22 languages, they face increasing challenges in fine-grained and multi-label sociolinguistic classification.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to solve a mystery that happens in 22 different languages around the world. The mystery isn't about a stolen jewel, but about something called "polarization." In simple terms, polarization is when people on social media start shouting at each other, drawing hard lines between "us" and "them," and refusing to listen to the other side.
The team behind this paper, Lingo Research Group, entered a global contest (SemEval-2026) to see if they could teach a super-smart computer (an AI) to spot this fighting, figure out who they are fighting about, and how they are fighting.
Here is how they did it, explained with some everyday analogies:
The Three Levels of the Mystery
The contest had three levels of difficulty, like a video game:
- Level 1 (The "Is it happening?" Check): The AI just needs to answer "Yes" or "No." Is this post angry and divisive, or is it just a normal post?
- Analogy: Like a smoke detector. It just needs to say, "Is there smoke?"
- Level 2 (The "Who is it about?" Check): If the post is polarized, who are they fighting about? Is it politicians? A specific race? A religion? Gender?
- Analogy: Like a detective asking, "Who is the suspect?" The AI has to pick one or more suspects from a lineup.
- Level 3 (The "How are they fighting?" Check): This is the hardest level. How is the hate being expressed? Are they using stereotypes? Are they calling people names? Are they acting like they don't care about others' feelings?
- Analogy: Like analyzing the style of the fight. Are they using a knife, a blunt object, or just shouting? This requires spotting subtle, tricky behavior.
The Tool: The "Prompt" as a Recipe
Instead of teaching the computer by showing it thousands of examples (which is like forcing a student to memorize a textbook), the team used prompts. Think of a prompt as a recipe or a set of instructions given to a very talented but literal-minded chef (the AI).
The team tried 12 different recipes.
- Some recipes were very short: "Is this angry?"
- Some were detailed: "Look for words that separate people into groups. If you see sarcasm, check if it's mean. Here are three examples of what that looks like..."
They tested these recipes on two different "chefs" (AI models): aya-101 and Gemma3-27B. They found that Gemma3-27B was the better chef, so they used it for the final contest.
The Results: Good at the Basics, Struggling with Nuance
The team's results were a mix of success and struggle, which tells an interesting story:
- Level 1 (Smoke Detector): The AI was very good at this. It correctly identified polarized posts about 76% of the time on average. It's like a smoke detector that rarely misses a fire.
- Level 2 (The Suspects): The score dropped to about 59%. It was harder to pinpoint exactly who was being attacked.
- Level 3 (The Fighting Style): The score dropped even further to about 44%. This was the hardest part.
Why did it get harder?
The authors explain that as the task gets more specific, the AI gets confused.
- The "Conservative Chef" Problem: The AI was trained to be very careful. It only said "Yes, this is polarized" if the evidence was super obvious (like someone screaming a slur).
- The "Sarcasm Trap": In English, people often fight using sarcasm, irony, or clever wordplay (e.g., "Oh, great, another socialist in the name of Jesus"). The AI, being too literal, missed these because there were no obvious "fighting words." It thought, "No direct insults? Then it must be safe."
- The "Direct vs. Indirect" Gap: In many other languages (like Hindi or Chinese), people are often more direct in their arguments. The AI was great at spotting those direct hits but struggled with the subtle, indirect English arguments.
The Language Puzzle
The AI didn't perform the same in every language.
- Nepali and Chinese were easy for the AI because the polarization was often very clear and direct (like "Group A vs. Group B").
- English was surprisingly difficult. Because English speakers use so much sarcasm and cultural shorthand, the AI kept missing the fights. It's like trying to understand a joke in a language you don't fully know; you hear the words, but you miss the punchline.
The Big Takeaway
The paper concludes that AI is great at spotting the "loud" fights (Level 1) but struggles with the "quiet" or "clever" fights (Levels 2 and 3).
The team learned that you can't just give an AI a simple instruction and expect it to understand complex human emotions across 22 different cultures. The more you ask it to analyze the nuance (the "how" and "who"), the more it tends to miss the subtle signs, especially when people are being sarcastic or indirect.
In short: The AI is a good detective for obvious crimes, but it needs more training to understand the subtle, tricky, and sarcastic ways people argue online.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.