← Latest papers
💬 NLP

Do Large Language Models Perform Well on Comprehending Poetic Logic in Modern Chinese Poetry?

This paper introduces Peony, the first benchmark designed to evaluate the "poetic logic" of modern Chinese poetry across stanza, line, and imagery levels, revealing significant limitations in current large language models' ability to comprehend such texts through holistic reasoning.

Original authors: Tian Lan, Shanshan Wang, Zehua Duo, Jiang Li, Guanglai Gao, Derek F. Wong, Xiangdong Su

Published 2026-08-25
📖 5 min read🧠 Deep dive

Original authors: Tian Lan, Shanshan Wang, Zehua Duo, Jiang Li, Guanglai Gao, Derek F. Wong, Xiangdong Su

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Language is more than a tool for exchanging facts; it is a vessel for emotion, imagery, and the subtle, often illogical connections that define human experience. For decades, artificial intelligence has learned to read and write by studying vast libraries of news articles, scientific papers, and everyday conversations. In these texts, meaning usually follows a straight line: a cause leads to an effect, a statement is followed by evidence, and the goal is clarity. But there is a different kind of writing where the rules are inverted. Poetry, particularly modern Chinese poetry, often relies on silence, sudden jumps in time, and images that do not fit together in the real world but make perfect sense in the mind of a reader. This style demands a unique form of reasoning, one that pieces together a whole story from fragments that seem unrelated on the surface.

Until recently, it was unclear whether the most advanced computer programs could understand this kind of writing. These programs, known as large language models, have become incredibly skilled at summarizing news or writing code, but they have rarely been tested on the complex, emotional logic of poetry. The question was not just whether a machine could recite a poem, but whether it could grasp the hidden thread that ties a line about a withered flower to a feeling of deep sadness, or understand why a poet might describe the night rising from the earth. To answer this, a team of researchers from universities in China and Japan set out to build a new way of testing these machines, creating a specialized challenge designed to measure their ability to navigate the unique logic of modern Chinese verse.

The researchers created a new testing ground called Peony, named after a flower often found in poetry, to see if artificial intelligence could truly comprehend the hidden structures of modern Chinese poems. They gathered 800 high-quality poems written by contemporary poets, ensuring the material was fresh enough that the computers had not simply memorized the answers from their training data. Instead of asking the machines to write poems or translate them, the researchers designed four specific puzzles that forced the models to think like a reader. The first two puzzles asked the computers to put shuffled pieces of a poem back in the correct order, either line by line or stanza by stanza. The other two puzzles asked the models to fill in missing words or sentences, choosing the one option that fit the emotional and logical flow of the piece, while rejecting other options that might sound similar but broke the poem's internal logic.

When the researchers ran these tests on six of the most powerful language models available, the results revealed a significant gap between human understanding and machine processing. While some models performed reasonably well on tasks that required matching simple images, they struggled profoundly when asked to follow the deeper narrative or emotional journey of a poem. The most capable models, which use a "thinking" process to reason through a problem step-by-step, did better than those that answered instantly, but even the best performers failed to consistently get the logic right. In the most difficult tasks, where the model had to infer the connection between two distant lines of a poem, the success rate was surprisingly low. The data showed that while these machines could often identify a single correct word or line in isolation, they frequently lost the thread when trying to hold the entire poem in their mind at once. They could see the individual bricks but often failed to understand the shape of the building.

The study also highlighted a specific weakness in how these models handle the "poetic logic" that defines the genre. Modern Chinese poetry often relies on juxtaposing images that seem contradictory, such as a night rising from the ground or a face that is like a lotus blooming and falling. The researchers found that the models tended to rely on surface-level meanings and common associations rather than the deeper, often abstract emotional connections that a human reader would make. For instance, when a poem referenced a famous historical figure or a cultural concept, the models often missed the subtle link between that reference and the poem's theme, treating the words as mere data points rather than carriers of meaning. This suggests that while artificial intelligence has mastered the mechanics of language, it still lacks the intuitive ability to navigate the abstract, emotional landscapes that poets inhabit.

The researchers concluded that current artificial intelligence systems are not yet ready to fully comprehend the poetic logic of modern Chinese poetry. The models can mimic the form and even get the right answer by chance or by spotting familiar patterns, but they struggle to reconstruct the holistic meaning that emerges from the interplay of images, emotions, and silence. The "Peony" benchmark serves as a clear map of where these systems stand today, showing that the leap from processing words to understanding the soul of a poem is still a vast distance. As these technologies continue to evolve, this new test offers a way to measure not just how well a machine can speak, but how deeply it can feel the unspoken connections that make poetry a uniquely human art form.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →