NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap
The paper introduces NOLLI, a procedurally generated and difficulty-calibrated benchmark of 7,500 English-Korean puzzle tasks that reveals Korean language models perform comparably to English on direct translations but face significant gaps in writing-system-intensive tasks due to challenges in multi-step sub-syllabic execution, while also demonstrating that structural task size is an unreliable proxy for empirical difficulty.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a world where computers are learning to read, write, and think just like humans. For a long time, scientists have been testing these "Large Language Models" (LLMs) with tricky puzzles to see if they truly understand logic or if they are just memorizing patterns. But there's a catch: most of these tests are written in English. When we try to test the same models in other languages, like Korean, the scores often drop. This raises a big question: Is the computer just bad at the new language, or is it struggling with something deeper, like how the letters are built or the cultural rules behind the words? To answer this, we need a fair playing field where the only thing changing is the language, not the difficulty of the game itself.
Enter NOLLI, a new set of logic puzzles designed specifically to figure out exactly where and why AI struggles with Korean. Think of NOLLI as a giant, procedurally generated obstacle course. Instead of using the same old riddles, the researchers built a machine that can create thousands of unique puzzles on the fly, from Sudoku grids to secret codes. The magic trick here is "calibration." Usually, making a puzzle "harder" means making the grid bigger or adding more words. But the NOLLI team realized that bigger doesn't always mean harder for a computer. So, they tuned the puzzle generators like a radio dial until a specific "reference" computer got the right amount of questions right (75% for easy, 50% for medium, 25% for hard). This ensures that when they compare English and Korean, they are comparing apples to apples, not apples to oranges.
The team tested 15 different AI models, ranging from the super-smart "frontier" models that everyone talks about to smaller, open-source ones. They organized the puzzles into three levels. The first level was Direct Translations: the exact same puzzle written in English and Korean. The second level was Script Adaptations, where the puzzle uses the building blocks of the Korean alphabet (called jamo) instead of English letters. The third level was Korean-Only, featuring puzzles that rely on Korean culture, like figuring out complex family relationships or traditional calendar math.
Here is what they found, and it's a bit of a plot twist. First, the language itself isn't the villain. When the puzzles were just direct translations, the AI performed almost exactly the same in English and Korean. The gap was tiny, suggesting that simply reading Korean isn't what's tripping the models up. However, the trouble started when the puzzles required manipulating the Korean writing system. In a game called "Korean Cipher," where the AI had to decode messages using the sub-syllabic parts of Korean letters, some models crashed completely, dropping up to 68.7 percentage points compared to their English performance. Interestingly, another puzzle using the same letters but treating them as simple symbols didn't cause a drop. This suggests the problem isn't the letters themselves, but the complex, multi-step mental gymnastics required to break them apart and put them back together.
Finally, the team looked at the Korean-only cultural puzzles. They found a strange pattern: while some tasks got easier or harder depending on the model, one specific task—figuring out Korean family titles (like "maternal grandfather's younger brother's wife")—was hard for every single model, no matter how smart it was. Even the best AI models couldn't close this gap. The study concludes that while AI has gotten great at translating, it still stumbles when asked to perform complex, step-by-step operations on the unique building blocks of the Korean language or to navigate its specific cultural rules. It's not that the computer can't read Korean; it's that it hasn't quite learned how to do the mental dance required to manipulate it.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.