← Latest papers
💬 NLP

Progressing beyond Art Masterpieces or Touristic Clichés: how to assess your LLMs for cultural alignment?

This paper addresses the limitations of existing datasets for assessing Large Language Models' cultural alignment by proposing new design guidelines, constructing a corresponding dataset, and demonstrating through contrastive experiments that this approach yields more discriminative results in distinguishing culturally specialized models.

Original authors: António Branco, João Silva, Nuno Marques, Luis Gomes, Ricardo Campos, Raquel Sequeira, Sara Nerea, Rodrigo Silva, Miguel Marques, Rodrigo Duarte, Artur Putyato, Diogo Folques, Tiago Valente

Published 2026-04-29
📖 5 min read🧠 Deep dive

Original authors: António Branco, João Silva, Nuno Marques, Luis Gomes, Ricardo Campos, Raquel Sequeira, Sara Nerea, Rodrigo Silva, Miguel Marques, Rodrigo Duarte, Artur Putyato, Diogo Folques, Tiago Valente

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to hire a new employee to work in a specific neighborhood. You want to know if they truly understand the local vibe, the inside jokes, and the unwritten rules of that community, or if they are just a tourist who has memorized a few guidebook facts.

This paper is about how we test Large Language Models (LLMs)—the AI chatbots we use every day—to see if they are truly "culturally aligned" with a specific group of people (in this case, Portuguese people), or if they are just faking it with surface-level knowledge.

Here is the breakdown of the paper's journey, using some simple analogies:

1. The Problem: The "Tourist Brochure" Test

The authors argue that most current tests for AI culture are like tourist brochures. They ask questions like:

  • "What is the most famous painting in Portugal?"
  • "What do Italians usually eat for dinner?"
  • "What is the capital of France?"

Why this fails:

  • The "Google" Effect: Any AI trained on the internet already knows these facts. It's like testing a local resident by asking them the name of the Eiffel Tower; even a tourist knows that. It doesn't prove they live there.
  • The Stereotype Trap: These questions often rely on clichés (e.g., "Italians love pasta"). This is like judging a whole neighborhood by a single cartoon character. It misses the real, messy, everyday reality of the culture.
  • The "Crutch" Problem: Many tests explicitly tell the AI, "Answer this question as if you are in Portugal." The authors say this is cheating. It's like giving a student a cheat sheet that says "You are in the exam room" before they even read the question. It inflates their score without proving they actually know the material.

2. The Solution: The "Local Pub" Test

The authors propose a new way to build these tests, which they call Tuguesice-PT. Instead of asking about famous landmarks or history books, they designed the test to feel like a casual conversation in a local pub.

They created a set of strict rules (guidelines) for the people writing the questions:

  • No "Tourist" Clues: Don't say "In Portugal..." in the question. Just ask the question naturally. If you ask, "Who was the dictator for most of the 20th century?" without mentioning the country, a true local should know the answer, but a generic AI might guess wrong.
  • Everyday Knowledge: Ask things that a local kid learns growing up, not things found in an encyclopedia. For example, asking about the specific bus route between two local cities, rather than the population of the country.
  • One Clear Answer: Avoid vague questions like "What do people usually drink?" (which could be wine, beer, or water depending on who you ask). Ask specific, factual things that have one right answer.
  • The "AI Fail" Rule: Ideally, the questions should be tricky enough that big, famous AI models (like the ones from Google or OpenAI) get them wrong if they haven't been specifically trained on that culture.

3. The Experiment: The Showdown

The team built their new "Local Pub" test (Tuguesice-PT) and compared it against an old "Tourist Brochure" test (a translated version of a dataset called BLEnD).

They ran a series of models against both tests:

  • The "Locals": AI models that had been specifically fine-tuned (trained) on Portuguese data.
  • The "Tourists": Generic AI models that had not been trained on Portuguese data.
  • The "Big Names": Massive, powerful models like Gemini and Llama.

The Results:

  • On the Old Test (Tourist Brochure): Everyone scored high. The "Locals" and the "Tourists" got almost the same score. The test couldn't tell the difference between a model that truly understood the culture and one that just memorized Wikipedia. It was like a test where everyone got an 'A' because the questions were too easy.
  • On the New Test (Local Pub): The difference was huge.
    • The models specifically trained for Portugal ("Locals") did much better.
    • The generic models ("Tourists") struggled significantly.
    • The "Oracle" Test: When the researchers gave the generic models a "hint" (telling them, "Assume you are Portuguese"), their scores skyrocketed on the old test but stayed low on the new one. This proved that the old test was just measuring if the AI could follow instructions, not if it actually knew the culture.

4. The Conclusion

The paper concludes that to truly know if an AI understands a culture, we need to stop asking it about art masterpieces and tourist clichés.

Instead, we need to ask it about the unspoken, everyday details that only someone who has lived there would know naturally. By removing the "crutches" (like telling the AI which country it's in) and focusing on specific, local, factual knowledge, we can finally see which models are truly aligned with a culture and which ones are just pretending.

In short: The old tests were like asking a fish, "Do you know what water is?" (Everyone says yes). The new test asks, "What is the best spot to catch a sardine off the coast of Lisbon on a Tuesday?" (Only the local fish knows).

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →