← Latest papers
💬 NLP

KorNAT: LLM Alignment Benchmark for Korean Social Values and Common Knowledge

This paper introduces KorNAT, the first benchmark for evaluating Large Language Models' alignment with South Korea by assessing their understanding of nation-specific social values through a large-scale survey and common knowledge via textbook-based questions, revealing that current models often fall short of the required alignment standards.

Original authors: Jiyoung Lee, Minwoo Kim, Seungho Kim, Junghwan Kim, Seunghyun Won, Hwaran Lee, Edward Choi

Published 2026-08-21
📖 5 min read🧠 Deep dive

Original authors: Jiyoung Lee, Minwoo Kim, Seungho Kim, Junghwan Kim, Seunghyun Won, Hwaran Lee, Edward Choi

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the rapidly evolving world of artificial intelligence, large language models have become powerful tools capable of writing stories, solving problems, and answering questions with a fluency that once seemed impossible. These systems are trained on vast amounts of text from the internet, learning patterns of human language and knowledge. However, a significant gap has emerged in how these models interact with different cultures. While they excel at understanding general concepts, they often struggle to grasp the specific social values, historical context, and everyday common knowledge that define life in a particular country. A model might know the facts of physics, which are universal, but it may fail to understand the nuanced opinions of a society regarding local laws, or the specific way a nation remembers its history. For an artificial intelligence to be truly useful and safe in a specific country, it must align with the collective mindset and basic knowledge of that nation's people.

Researchers in South Korea have taken a major step toward solving this problem by creating a new way to measure how well these digital minds understand the Korean way of life. They introduced a benchmark called KorNAT, which acts as a comprehensive test for artificial intelligence, specifically designed to see if a model can think and feel like a Korean citizen. The team built this test around two main pillars: social values and common knowledge. Social values refer to the shared opinions people hold about important societal issues, such as whether certain laws should be changed or how to handle social conflicts. Common knowledge covers the basic facts and cultural touchstones that most people learn in school, from history and literature to science and geography. By testing models on these two fronts, the researchers could determine if an artificial intelligence is merely reciting facts or if it truly understands the cultural fabric of the nation.

To build this test, the researchers did not simply ask a few experts for their opinions. Instead, they conducted a massive survey involving 6,174 unique Korean participants. They asked these real people a wide range of questions about current social issues, gathering a diverse set of responses to establish what the "correct" answer looks like for the country. For the social values portion of the test, there is no single right answer; rather, the goal is for the artificial intelligence to match the distribution of opinions found in the general population. If most people agree with a statement, the model should agree too. For the common knowledge portion, the questions were drawn from standard Korean textbooks and educational materials, covering subjects like Korean history, social studies, and science. These questions have definite right answers, similar to a school exam, to test if the model knows the basic facts that every educated Korean person should know.

The team then put seven of the world's most advanced artificial intelligence models to the test. The results revealed a clear divide between models that are generally capable and those that are truly aligned with Korean culture. While some of the top global models performed reasonably well, they often fell short of the standard set by the survey data. One model, which had been specifically trained on a large amount of Korean text, stood out as the most aligned. It demonstrated a much deeper understanding of both the social values and the common knowledge of the country, outperforming the others significantly. This suggests that while general intelligence is powerful, it is not enough on its own; a model needs specific cultural training to truly resonate with a local population.

The study also highlighted that even the best-performing models have room for improvement. In the social values section, only a few models managed to reach the level of alignment the researchers considered a reference score. In the common knowledge section, the results were similar, with most models struggling to reach the threshold of basic knowledge expected of an adult in the country. The researchers noted that some models were hesitant to answer questions about social values, perhaps because they were programmed to avoid taking sides on controversial topics. This hesitation, while well-intentioned, meant they failed to reflect the actual opinions of the people they were meant to serve. The findings suggest that for artificial intelligence to be effectively deployed in a specific country, it must be carefully tuned to understand not just the language, but the values and the shared knowledge of that society.

This work represents the first time a benchmark has been created to measure national alignment in this specific way, moving beyond simple language translation to test cultural understanding. The dataset, which contains thousands of carefully curated questions, has been reviewed and approved by a government-affiliated organization dedicated to ensuring data quality. By making this test available, the researchers hope to encourage the development of artificial intelligence that is more inclusive and better suited to the diverse needs of different nations. The ultimate goal is to ensure that as these powerful tools become part of daily life, they can assist people in ways that feel natural, respectful, and grounded in the reality of their own culture.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →