Poller: Are LLMs Suitable for Evaluating the Poetry Understanding Task?
This paper introduces Poller, a novel LLM-based evaluation method that adopts the poet's perspective to significantly reduce errors in assessing Chinese poetry understanding across specialized dimensions, effectively bridging the gap between automated efficiency and human expertise.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a box of beautiful, complex, modern Chinese poems. You want to know if a computer (specifically, a Large Language Model or LLM) can understand them well enough to grade someone else's explanation of those poems.
The Problem: The "Polite Robot" vs. The "Real Poet"
The researchers found that standard AI models are terrible at this job. It's like asking a polite, well-read robot to judge a painting. The robot sees the colors and the shapes, but it misses the soul of the piece.
In the paper's experiments, when the AI tried to grade a human's understanding of a poem, it was overly generous. It gave scores of 90 or 95 out of 100, even when the human's explanation was shallow or missed the point. The AI was essentially saying, "Oh, you used big words? That's great! Here's an A+!" It lacked the deep, cultural, and emotional context that a real poet would have.
The Solution: "Poller" (The Role-Playing Trick)
To fix this, the researchers invented a method called Poller (Poetry LLM Evaluator).
Think of it like this: Instead of asking the AI to be a generic "teacher," they ask the AI to pretend to be the actual author of the poem.
Here is how they do it:
- The Costume Change: They feed the AI a "character sheet" for the real poet. This includes the poet's biography, their personal struggles, their specific views on art, and even what famous critics have said about their work.
- The Perspective Shift: The AI is told, "You are now this poet. You wrote this poem. Now, look at this student's explanation of your work. Does it capture what you were trying to say?"
- The Verdict: By stepping into the poet's shoes, the AI stops being a polite robot and starts thinking like a creator. It begins to notice the subtle metaphors, the specific rhythm, and the hidden emotions that a generic AI would miss.
The Results: From "Too Nice" to "Spot On"
The results were dramatic.
- Before Poller: The AI was like a friend who never wants to hurt your feelings, giving high scores to almost everything. The gap between the AI's grade and a real human expert's grade was huge.
- After Poller: The AI became a strict, insightful critic. When it played the role of the poet, its grading became much closer to what a human expert would give.
For example, when judging how well someone understood the rhetorical techniques (the fancy tricks poets use) or defamiliarization (making the familiar look strange to make you think), the AI's errors dropped by nearly 95%. It went from being wildly inaccurate to being almost as good as a human expert.
The Takeaway
The paper concludes that AI can be a good judge of poetry, but only if you trick it into wearing the poet's hat. By forcing the AI to adopt the specific perspective, background, and soul of the author, we bridge the gap between "fast computer processing" and "deep human expertise."
What the Paper Does NOT Say
- It does not claim this method works for all types of writing (like news articles or legal contracts); it is specifically tested on modern Chinese poetry.
- It does not say this will replace human poets or critics entirely; it just offers a better way to automatically grade poetry understanding tasks.
- It does not claim the AI actually feels emotions; it just simulates the poet's perspective so well that the grading becomes accurate.
In short: To get a computer to understand poetry, you have to make it pretend to be the poet. Once it does, it finally gets the joke.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.