SwanNLP at SemEval-2026 Task 5: An LLM-based Framework for Plausibility Scoring in Narrative Word Sense Disambiguation
The SwanNLP team proposes an LLM-based framework for SemEval-2026 Task 5 that utilizes structured reasoning, dynamic few-shot prompting, and model ensembling to effectively predict human-perceived plausibility of word senses in narrative texts, demonstrating that commercial large-parameter models closely replicate human judgments.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are reading a short story to a friend. Suddenly, you come across a word that has two very different meanings, like the word "bat." It could be a flying animal, or it could be a piece of sports equipment.
In a simple sentence like "The bat flew out of the cave," it's obvious. But in a complex story, the clues might be subtle. Maybe the story is about a baseball game, but the character is wearing a cape. Is the "bat" the animal or the sports gear? Humans are pretty good at guessing which meaning makes the most sense based on the whole story.
This paper is about teaching computers to do exactly that: guess which meaning of a word makes the most sense in a story.
Here is a breakdown of what the researchers at Swansea University did, using some everyday analogies.
1. The Problem: The "Ambiguous Story" Puzzle
The researchers participated in a competition called SemEval-2026 Task 5. Think of this as a giant puzzle contest.
- The Challenge: They were given short stories with tricky words.
- The Goal: Instead of just picking one answer (like a multiple-choice test), they had to give the story a "Plausibility Score" from 1 to 5.
- 5: "Oh, that makes perfect sense! The word definitely means X here."
- 1: "No way. That meaning fits the story about as well as a square peg in a round hole."
- The Twist: They had to mimic human judgment. Sometimes, humans disagree. One person might think the "bat" is an animal (score 4), while another thinks it's sports gear (score 2). The computer needed to understand why humans might feel differently.
2. The Solution: Three Different "Brain" Strategies
The team tried three different ways to teach their computer models how to think like humans.
Strategy A: The "Apprentice" (Fine-Tuning Small Models)
Imagine you have a smart but young student (a smaller AI model). You want them to learn how to grade stories.
- The Old Way: You just show them the story and ask for a grade.
- The New Way: You act like a strict teacher. You say, "First, look at the clues. Is the story about baseball or a cave? Is it easy to figure out, or is it tricky? Then give me a grade."
- The Result: By forcing the computer to "think out loud" (a process called Chain-of-Thought) and check if the story is "easy" or "hard" first, the student got much better at guessing the right score. It was like teaching the student to use a checklist before answering.
Strategy B: The "Consultant" (Dynamic Few-Shot Learning)
For the bigger, smarter computers (Commercial LLMs like GPT-4), they didn't need to be retrained from scratch. Instead, they used a Reference Library.
- The Analogy: Imagine you are a judge in a contest. Before you make a decision, you look at a few past cases that are very similar to the one you are judging right now.
- How it worked: The computer was given the story, and then it was shown 1 or 3 examples of other stories and how humans scored them.
- The Result: This was like giving the computer a cheat sheet of "similar situations." It helped the computer realize, "Ah, this story feels like that other one where the score was high," leading to much more accurate human-like scores.
Strategy C: The "Council of Judges" (Ensembling)
The researchers noticed that sometimes one computer makes a mistake, just like one human might have a bad day.
- The Analogy: Instead of asking one person to grade the story, they asked five different computers to grade it independently. Then, they took the average of all five scores.
- The Result: This is like a panel of judges in a talent show. If one judge is too harsh and another is too lenient, the average score is usually fairer and closer to what the audience (humans) actually thought. This method worked the best overall.
3. The Big Takeaways
- Thinking Matters: Computers that were forced to "reason" through the story (checking clues, checking difficulty) did much better than those that just guessed.
- Context is King: The bigger, more expensive computers that could look at "similar past stories" (the Consultant strategy) were the best at mimicking human nuance.
- Teamwork Wins: When they combined the opinions of multiple models (the Council), the results were even closer to human agreement than any single model could achieve alone.
4. The Verdict
The team finished in 10th place in the competition. While they didn't win first place, they proved that computers are getting very good at understanding the feeling and logic of a story, not just the grammar.
In short: They taught computers to stop just reading words and start reading between the lines, using checklists, reference books, and group discussions to figure out what a story is really trying to say.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.