← Latest papers
💻 computer science

Evaluating Large language models on Understanding Korean indirect Speech acts

This study evaluates the ability of various large language models to understand Korean indirect speech acts by constructing a specialized dataset and employing both automated and human assessments, revealing that while proprietary models like Claude3-Opus outperform open-source alternatives, none yet match human-level proficiency in interpreting context-dependent intentions.

Original authors: Youngeun Koo, Jiwoo Lee, Dojun Park, Seohyun Park, Sungeun Lee

Published 2026-08-10
📖 4 min read☕ Coffee break read

Original authors: Youngeun Koo, Jiwoo Lee, Dojun Park, Seohyun Park, Sungeun Lee

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are chatting with a friend, and they say, "It's freezing in here." If you are standing next to a broken air conditioner, they are likely just stating a fact. But if you are sitting in a warm living room with the windows closed, they are probably hinting, "Please close the window" or "Turn up the heat." This gap between what words literally say and what they actually mean is the secret sauce of human conversation. In the world of science, this is called pragmatics: the study of how context changes meaning. While computers have gotten incredibly good at translating languages or solving math problems, they often struggle with this "reading between the lines" skill. They might take your hint literally and just say, "Yes, the temperature is low," missing the point entirely. This matters because as we start relying on AI assistants for everything from scheduling meetings to giving advice, we need them to understand not just our words, but our intentions.

This paper is like a detective story where researchers put a bunch of Artificial Intelligence (AI) models to the test to see if they can crack the code of indirect hints. The scientists focused on Korean, a language where context is king, and they built a special quiz based on a classic theory of how humans use language. They created 240 scenarios where the same sentence could mean two very different things depending on the situation. For example, saying "The ice is thin" could be a simple fact, or it could be a warning to kids not to step on a frozen pond. The researchers wanted to see if AI could tell the difference between a direct statement and an indirect request, promise, or expression of emotion.

The results were a mix of "not bad" and "still a long way to go." The researchers tested 12 different AI models, ranging from massive, expensive systems to smaller, open-source ones. The clear winner was a model called Claude3-Opus. It scored about 71.94% on a multiple-choice test and 65% on a more open-ended test where it had to explain its reasoning. While this is impressive, it's important to note that even the best AI model didn't beat the humans. The human participants in the study scored around 77.64%, proving that people are still the masters of understanding subtle social cues.

The study found that most AI models are great at taking things literally but terrible at guessing the hidden meaning. If a sentence was a direct request, the AI got it right most of the time. But when the request was indirect—like saying "It's nice outside" to suggest going for a walk instead of just talking about the weather—most models stumbled. They tended to get stuck on the literal meaning of the words. Interestingly, the AI models were surprisingly good at understanding indirect compliments or idioms (like saying someone is a "chef" because they cooked a great meal), but they struggled the most with indirect requests.

The researchers also noticed some funny quirks in how the AI behaved. Some smaller models, when asked to answer in Korean, accidentally replied in English, as if they were confused about who they were talking to. Others tried to force multiple-choice answers even when the question asked for a free-form explanation, as if they were trying to game a test they had seen before. The paper suggests that while these AI models are powerful, they still lack the deep, intuitive grasp of human context that even a teenager has. They are like students who have memorized the dictionary but haven't yet learned how to have a real conversation. The study concludes that to make AI truly helpful, we need to teach them to look past the surface of words and understand the messy, context-filled world of human intention.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →