TurkCuisineBench: A source-grounded short-answer benchmark for large language models’ factual knowledge of Turkish cuisine and culinary heritage
This paper introduces TurkCuisineBench, a 72-item, source-grounded short-answer benchmark that evaluates large language models' factual knowledge of Turkish cuisine and heritage, demonstrating that semantic evaluation significantly improves accuracy assessment over exact-string matching while highlighting substantial performance variations across different model endpoints.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the rapidly evolving world of artificial intelligence, computers have learned to speak human languages with increasing fluency, often answering questions about history, science, and culture with surprising ease. However, a significant gap remains in how we test these machines. Most standard tests rely on multiple-choice questions or simple word-for-word matching, which can miss the nuance of a correct answer that is phrased differently. This is particularly true for languages like Turkish, which uses a flexible structure where words change form to fit their role in a sentence, and for specialized fields like culinary heritage, where a dish might be known by several names or described with varying details depending on the region. Researchers are increasingly asking whether these digital minds truly understand the specific cultural knowledge of a community or if they are merely guessing based on patterns they have seen before. To answer this, we need tests that are not just broad, but deeply rooted in the specific facts and traditions of a place, verified by real human experts rather than just automated scoring systems.
A new study introduces a specialized test designed to measure exactly this kind of knowledge: the ability of large language models to answer factual questions about Turkish cuisine and its history. The researchers, led by Muhammed Buğra Yılmaz, created a benchmark called TurkCuisineBench, which consists of seventy-two short-answer questions. These questions were not made up; every single one was derived from official records, such as government databases for geographical indications and UNESCO lists of intangible cultural heritage. The goal was to see if artificial intelligence could provide answers that matched these trusted sources, even if the wording was not identical. The study involved eight different AI models, ranging from well-known commercial systems to open-source versions, all asked to answer the questions in Turkish without looking up information or using outside tools. The process was strictly controlled: the questions were locked in place before the models began, and the answers were evaluated by human experts who checked for meaning rather than just spelling.
The results revealed a wide gap in performance between the different models. When the answers were judged strictly by whether they matched the source text exactly, the scores were modest, with the best-performing model getting about sixty-two percent correct. However, when human experts reviewed the answers to see if they were semantically correct—meaning they carried the same factual meaning even if the words were different—the scores improved significantly. The top model achieved a semantic accuracy of nearly seventy-one percent, while the lowest-performing models struggled to reach fourteen percent. This difference highlights a critical flaw in how many AI systems are currently evaluated: a machine might give a perfectly correct answer that is simply phrased differently from the one in the database, and a rigid computer program would mark it as wrong. By having humans review the responses, the study recovered forty-three correct answers that would have been missed by strict automated scoring, raising the overall accuracy of the best models by more than seven percentage points.
The study also looked closely at how the models failed. When they got an answer wrong, the most common mistake was substitution, where the model would provide the right type of information but the wrong specific detail, such as naming the wrong ingredient or the wrong region for a dish. Another interesting finding was how often the models refused to answer. One specific model chose to say it did not know the answer in nearly half of the cases, while others almost always attempted to provide an answer, even if it was incorrect. This suggests that different AI systems have very different strategies for handling uncertainty. The researchers also tested whether the presence of certain clues in the question would make it easier for the models to answer correctly, but the data showed no clear evidence that this was the case; the difficulty of the question seemed to depend more on the specific topic, such as whether it was about a specific dish or a broad historical tradition, rather than the wording of the prompt.
Perhaps the most important takeaway from this work is not just which AI model is the best, but how we should measure them. The study demonstrates that for culturally specific knowledge, relying solely on automated, exact-match scoring gives an incomplete and often unfair picture of a model's capabilities. By combining strict source verification with careful human review, researchers can create a more accurate and fair assessment of what these systems actually know. The benchmark itself is deliberately narrow, focusing only on facts that can be traced back to official documents, which means it does not claim to represent the entire, living culture of Turkish cooking. Instead, it offers a transparent, auditable way to test if machines can handle the specific, documented facts of a heritage. This approach provides a blueprint for future evaluations, showing that to truly understand how artificial intelligence handles human culture, we must look beyond simple word matching and engage with the meaning behind the answers.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.