← Latest papers
🤖 AI

Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark

This paper introduces ProverbIT, an Italian proverb benchmark that reveals large language models struggle to distinguish correct proverb endings from literal synonyms in multiple-choice tasks, suggesting their performance relies more on memorized patterns than deep semantic understanding of culturally embedded figurative language.

Original authors: Enrico Mensa, Lorenzo Zane, Calogero Jerik Scozzaro, Matteo Delsanto, Tommaso Milani, Daniele Paolo Radicioni

Published 2026-08-06
📖 4 min read☕ Coffee break read

Original authors: Enrico Mensa, Lorenzo Zane, Calogero Jerik Scozzaro, Matteo Delsanto, Tommaso Milani, Daniele Paolo Radicioni

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are walking through a giant library where the shelves are stacked not just with books, but with the entire history of human conversation. This is the world of Large Language Models (LLMs). Think of these AI systems as super-fast, super-obsessive readers who have devoured almost everything written on the internet. They are incredibly good at spotting patterns, like knowing that if someone says "It's raining cats and," the next word is almost certainly "dogs." They use this pattern-matching to write emails, solve math problems, and even write code.

But here's the tricky part: just because an AI has read a million stories doesn't mean it truly understands the jokes, the sarcasm, or the old sayings hidden inside them. This is where proverbs come in. A proverb is a short, wise saying passed down through generations, like "Don't count your chickens before they hatch." These aren't just random words; they are cultural shortcuts that rely on shared human experience. For a computer, figuring out a proverb is like trying to solve a riddle where the answer depends on knowing the feeling of the culture, not just the dictionary definition of the words. Scientists are curious: Do these AI models actually get the joke, or are they just guessing based on what words usually hang out together?

This is exactly what the researchers behind the ProverbIT paper wanted to find out. They created a special test, a "quiz" made entirely of Italian proverbs, to see how well different AI models could handle these cultural riddles. They didn't just ask the AI to finish the sentence; they set a trap. They gave the models a list of endings, but they made sure the correct ending wasn't on the list. Instead, the list was filled with clever fakes: some that sounded similar, some that meant the same thing but sounded silly, and some that were just plain wrong. The only correct choice was "None of the above."

The results were a bit of a shocker. When the researchers simply asked the AI to finish the proverb on its own, almost every model got it right. It was like asking a student to recite a poem they had memorized; they nailed it. But the moment the researchers switched to the multiple-choice quiz with the "trick" question, the scores plummeted. Even the most advanced, "reasoning" models—those designed to think step-by-step like a human—stumbled badly. Instead of realizing the correct answer was missing and picking "None of the above," they stubbornly picked one of the fake options.

Why did this happen? The researchers looked inside the AI's "brain" (its thought process) and found a fascinating flaw. The models did know the correct ending. In their internal thinking, they would mention the right phrase over and over again. But when it came time to make a final choice, they seemed to ignore their own knowledge. They were so focused on finding a word that sounded right or looked like a synonym that they couldn't say, "Wait, the real answer isn't here." It's like a student who knows the answer is "42," sees a test with options "40," "44," and "46," and then confidently circles "44" because it's the closest number, completely forgetting that the real answer isn't even on the page.

The study suggests that while these AI models are amazing at memorizing patterns and recalling facts, they still struggle with a specific kind of thinking: negative reasoning. This is the ability to look at a set of choices and realize, "None of these are right." They rely heavily on what they've seen before rather than truly understanding the meaning behind the words. So, while these digital brains are getting smarter every day, they still have a lot to learn about the subtle, tricky, and sometimes missing pieces of human wisdom.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →