Pun Intended: Multi-Agent Translation of Wordplay with Contrastive Learning and Phonetic-Semantic Embeddings for CLEF JOKER 2025 Task 2
This paper presents a novel three-stage multi-agent framework combining large language models, phonetic-semantic embeddings, and contrastive learning to successfully translate English puns into French, achieving top rankings in the CLEF JOKER 2025 competition by prioritizing linguistic creativity over literal translation.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to translate a joke from one language to another, but the joke relies on a word that sounds like two different things at once. In English, you might say, "I'm reading a book on anti-gravity. It's impossible to put down!" The humor comes from the phrase "put down," which means both "to stop holding" and "to place on a surface." If you translate this literally into French, you lose the double meaning, and the joke dies. This is the tricky world of "puns," a type of wordplay where a single word carries two meanings or sounds like another word. For decades, computers have been terrible at this because they are trained to be literal, like a robot that follows a recipe exactly but has no sense of humor. They try to swap words one-for-one, which often destroys the punchline. But what if we could teach computers to be less like robots and more like comedians, understanding that sometimes you have to break the rules of translation to keep the joke alive?
This paper, titled "Pun Intended," is a report from a team of researchers at Georgia Tech who entered a competition called CLEF JOKER 2025. Their goal was to translate English puns into French using the latest, most powerful computer brains (called Large Language Models). Instead of just asking the computer to "translate," they built a three-step system designed to capture the spirit of the joke rather than just the dictionary definitions.
First, they tried a "baseline" approach, simply asking different AI models to generate French puns and using a second AI to act as a judge, checking if the result was actually a joke or just a boring sentence. This was like asking a group of friends to tell a joke and having one friend nod if it was funny. While this worked okay, the jokes often drifted too far from the original story.
To get better, they built a "guided chain-of-thought" pipeline. Think of this as giving the computer a map and a compass. They taught the AI to first identify the tricky word, then look up its two different meanings, and finally use a special mathematical tool called "phonetic-semantic embeddings." This tool is like a super-sensitive ear and a dictionary combined; it helps the computer find French words that not only mean the right things but also sound similar to the original English words. It's like trying to find a French word that sounds like "cat" but means "dog," just to see if you can make a joke out of it.
Finally, they added a "multi-agent" system. Imagine a panel of four different judges, each with a specific job: one checks if the meaning is right, one checks if the translation sounds natural, one checks the emotion, and one checks if it feels authentic. These judges didn't just give a score; they gave feedback. If the joke wasn't quite funny or the translation was clunky, the system would rewrite it and try again, up to five times, until the panel was happy.
The results were fascinating. When the researchers used standard computer tests (which measure how closely the new sentence matches the old one word-for-word), their system actually scored quite low. This is because their system intentionally avoided literal translations to keep the humor alive. However, when real human experts who speak French fluently read the jokes, the team's system scored in first and second place out of 51 teams. The humans loved the jokes because they were actually funny and made sense in French, even if they didn't look exactly like the English originals.
The paper suggests that to translate wordplay successfully, we must stop trying to be perfect translators and start trying to be creative partners. The researchers found that while the latest AI models are great at spotting where a pun is, they need help to find the right French words that sound and mean the right things. By combining a "guide" that looks at sounds and meanings, and a "panel" of judges that refine the joke, they showed that computers can learn to tell jokes that work across languages. They didn't solve the problem of translating every single pun perfectly, but they proved that shifting the goal from "literal accuracy" to "funny equivalence" is the right way to go.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.