Computational Hermeneutics: Evaluating generative AI as a cultural technology
This paper proposes "computational hermeneutics" as a new framework for evaluating generative AI by shifting from standardized accuracy metrics to an interpretive approach that treats culture as fundamental, addressing challenges of situatedness, plurality, and ambiguity through iterative, human-centered, and context-aware benchmarks.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: AI Isn't a Calculator; It's a Storyteller
Imagine you are trying to judge a new chef.
- The Old Way (Current AI Testing): You give the chef a recipe for a perfect chocolate cake. You check if the cake looks exactly like the picture. If it does, they get an A. If it's slightly burnt, they get an F. This works great for math or coding, where there is one "right" answer.
- The Problem: Generative AI (like ChatGPT or image generators) isn't just baking cakes. It's writing poetry, giving advice, painting portraits, and having conversations. In these areas, there is no single "right" answer. A poem can be sad or happy, and both are valid. A painting can be abstract or realistic.
The authors argue that we are currently judging these "storyteller" AI systems using "calculator" rules. We are treating culture (the messy, human stuff of meaning, history, and emotion) as just a small ingredient to be measured, rather than the main course.
The New Approach: "Computational Hermeneutics"
The paper suggests a new way to look at AI called Computational Hermeneutics.
- Hermeneutics is a fancy word from the humanities (like literature and history) that means "the art of interpretation." It's how we figure out what a book, a law, or a painting means in a specific situation.
- The Shift: Instead of asking, "Is this AI answer factually correct?" we should ask, "How does this AI answer make sense in this specific context?"
The authors say AI systems are actually "Context Machines." They don't just spit out facts; they try to guess what you mean based on the situation, the history, and the vibe of the conversation.
The Three Big Challenges (The "Three S's")
The paper says that because AI is dealing with human culture, it faces three specific hurdles that math problems don't have:
1. Situatedness (The "Where and When" Problem)
- The Analogy: Imagine someone tells a joke. If they tell it at a funeral, it's a disaster. If they tell it at a comedy club, it's a hit. The joke didn't change, but the context did.
- The AI Issue: Current AI often acts like it has a "God's eye view," pretending it knows the one true meaning of everything. But meaning only exists in a specific time and place. The paper argues we need to stop expecting AI to be neutral and start acknowledging that every answer comes from a specific perspective.
2. Plurality (The "Many Truths" Problem)
- The Analogy: Think of a song. To a music critic, it might be "innovative." To a parent, it might be "too loud." To a teenager, it's "an anthem." All three people are right.
- The AI Issue: AI is trained on the whole internet, which is full of people who disagree. Current tests try to force the AI to pick one "correct" answer. The paper says we need to accept that there can be multiple valid interpretations of the same thing, and AI should be able to handle that diversity without breaking.
3. Ambiguity (The "Fuzzy Edge" Problem)
- The Analogy: Great art often leaves things open to interpretation. If a movie explains every single detail and leaves no mystery, it feels boring. The "mystery" is what makes us think.
- The AI Issue: Right now, AI developers try to "fix" ambiguity. They want the AI to be 100% clear. But for creative tasks, ambiguity is a feature, not a bug. The paper argues we should evaluate AI on how well it navigates the gray areas, not how well it erases them.
Three Rules for Better Testing
If we want to test these "Context Machines" properly, the authors propose three new rules for how we build our tests (benchmarks):
1. Make Tests Iterative (The "Conversation" Rule)
- Old Way: Ask one question, get one answer, give a score.
- New Way: Have a conversation. Meaning changes as you talk. A good AI test should be a back-and-forth chat, not a multiple-choice quiz. You need to see how the AI adapts as the context evolves.
2. Include People (The "Human-in-the-Loop" Rule)
- Old Way: Let a computer grade the AI's output against a database.
- New Way: Real humans need to be part of the test. Since culture is about human feelings and values, only humans can truly judge if an AI response feels "right" or "appropriate" in a specific situation. We need to test the relationship between the human and the AI, not just the AI alone.
3. Measure Context, Not Just Output (The "Setting the Scene" Rule)
- Old Way: Did the AI write a sentence with perfect grammar?
- New Way: Did the AI understand who it was talking to, where they were, and why they were asking? We need to test the AI in the messy, real-world situations where it will actually be used, not in a sterile lab.
The Bottom Line
The paper is a call to action for the AI world. It says: "Stop treating culture like a bug to be fixed. Start treating it like the main feature."
By using Computational Hermeneutics, we can move away from asking "Is this AI smart?" to asking "Is this AI wise?" We can build systems that understand that meaning is messy, that there are many ways to see the world, and that the best answers often depend on the story we are telling together.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.