Multiperspectivity as a Resource for Narrative Similarity Prediction
This paper proposes leveraging an ensemble of 31 diverse LLM personas to embrace narrative multiperspectivity rather than treating it as a challenge, demonstrating that such an approach improves similarity prediction accuracy on the SemEval-2026 Task 4 dataset while revealing limitations in current single-ground-truth evaluation frameworks.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to decide which of two stories is more similar to a third "original" story. In the world of computers, this is usually treated like a math problem: there is one right answer, and the computer just needs to find it.
But this paper argues that storytelling isn't math; it's art. And in art, there is no single "right" answer. A literary critic might see a story about a woman as a tale of female empowerment, while a historian might see it as a critique of 19th-century society. Both are valid, but they lead to different conclusions about what the story "means."
The authors of this paper asked: What if we stopped trying to force computers to find the "one true answer" and instead let them argue like a panel of diverse experts?
Here is the breakdown of their experiment, explained with some everyday analogies.
1. The Setup: The "Super-Panel" of 31 Characters
Instead of asking one AI model to make the decision, the researchers created a team of 31 different AI "personas."
Think of this like hiring a jury for a trial, but instead of regular people, you hire:
- The Experts (Practitioners): A Feminist Literary Critic, a Postcolonial Theorist, a Data Scientist, a Psychologist. These are the "serious" thinkers with specific rulebooks for analyzing text.
- The Regulars (Lay People): A High School Student, a Football Player, a Pirate, a Taxi Driver. These are the "everyday" folks who just use their gut feeling.
They asked each of these 31 characters to vote on which story was more similar. Then, they took a majority vote (like a democratic election) to get the final answer.
2. The Surprise: The "Clueless" Crowd vs. The "Over-Thinkers"
The researchers expected the "Experts" to win. After all, they have fancy degrees and specific frameworks for analyzing stories, right?
Wrong.
- The Individual Result: When you looked at each character alone, the "Regulars" (like the Football Player or the Tour Guide) were actually better at guessing the right answer than the "Experts." The experts often got bogged down in their own complex theories and missed the obvious similarities.
- The Team Result: However, when you combined them all into a group vote, the Experts saved the day. Even though they were individually worse, their errors were different from each other. The Football Player might get confused by one thing, while the Feminist Critic gets confused by something else. When you mix them all together, their mistakes cancel each other out, and the group becomes incredibly smart.
The Analogy: Imagine trying to guess the weight of a giant pumpkin.
- One expert might overthink the soil density and get it wrong.
- Another might focus on the stem and get it wrong.
- But if you take 30 people with different, slightly wrong guesses and average them, you often get a result that is shockingly close to the real weight. This is called the "Wisdom of Crowds."
3. The Twist: The "Gender" Trap
Here is where the paper gets really interesting. They noticed a strange pattern: whenever the AI personas started using words related to gender, feminism, or patriarchy (like "female protagonist," "gender roles," or "objectification"), their accuracy dropped.
Why?
- Theory A: The benchmark (the test they were taking) was designed by people who didn't care about gender. So, when the AI started focusing on gender, it was looking at the "wrong" clues for this specific test.
- Theory B: The AI was actually finding valid, deep meanings in the text that the test creators completely missed. The test said "Wrong," but the AI said, "Actually, this is a very important part of the story."
This suggests that our current tests for AI might be unfair. They might punish an AI for being too smart or too culturally aware, forcing it to ignore valid interpretations just to get a high score.
4. The Conclusion: Diversity is the Secret Sauce
The paper concludes that to make AI understand stories better, we shouldn't just try to make it "smarter." We should make it more diverse.
- Don't just have one brain: Have a room full of different brains.
- Don't fear the "wrong" perspective: Even if a specific perspective (like a feminist critique) gets the "test score" wrong, it adds valuable diversity to the group, which helps the whole team make better decisions in the long run.
In a nutshell:
If you want a computer to understand a story, don't ask it to be a robot. Ask it to be a room full of people: a poet, a plumber, a historian, and a teenager. Let them argue, let them vote, and let their differences create a smarter, more human-like understanding of the story.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.