Target-Side Paraphrase Augmentation for Sign Language Translation with Large Language Models
This paper proposes a target-side paraphrase augmentation method using GPT-4o to generate controlled variants of reference sentences for training Sign Language Translation models, demonstrating improved BLEU scores and semantic fidelity on the PHOENIX14T dataset while highlighting the approach's limitations on highly repetitive or extremely sparse data.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot how to translate a person's hand gestures (Sign Language) into spoken words. The biggest problem is that the robot doesn't have enough practice books. It sees the same few sentences over and over, or it sees words it has never encountered before.
This paper is like a clever teacher who realizes: "If the robot only learns one way to say 'The weather is rainy,' it might get confused if the sign means 'Rain is falling' or 'It's wet outside.'"
Here is how the researchers fixed this, using simple analogies:
1. The Problem: The Robot's "One-Track Mind"
Sign Language is complex. To teach a computer to understand it, you need thousands of examples of videos paired with text. But these examples are rare.
- The "Long Tail" Problem: Imagine a library where 50% of the books only have one copy. If the robot needs to translate a word that appears only once in its training data, it's like asking a student to write a book report on a book they've never read.
- The "Repetitive" Problem: On the other hand, some datasets are like a broken record. They repeat the exact same sentences so perfectly that the robot just memorizes the answers without really understanding the meaning.
2. The Solution: The "Rewriting Machine"
Instead of trying to film more videos (which is hard and expensive), the researchers used a super-smart AI (GPT-4o) to act as a creative writing coach.
- The Trick: They kept the video of the sign language exactly the same. But for the text part, they asked the AI to rewrite the sentence in three different ways.
- Original: "The low-pressure areas determine our weather."
- Rewrite 1: "Our weather is determined by low-pressure areas."
- Rewrite 2: "Low-pressure areas are what control our weather."
- The Goal: This teaches the robot that the same hand movements can result in different word choices. It stops the robot from memorizing a single sentence and forces it to learn the actual meaning behind the signs.
3. The Training Method: "Broadening then Focusing"
They didn't just throw all these new sentences at the robot at once. They used a two-step training schedule, like a musician learning a new song:
- Phase 1 (The Jam Session): The robot practices with the original sentence plus all the AI-generated variations. This helps it understand that there are many ways to express the same idea.
- Phase 2 (The Final Rehearsal): The robot goes back to practicing only with the original, perfect sentences. This ensures that when it takes the final test, it still knows the "standard" answer, but now it understands the concept better.
4. The Results: It Depends on the "Recipe"
The researchers tested this method on three different "kitchens" (datasets), and the results were very different:
- Kitchen A (German Sign Language - PHOENIX14T): This was a normal kitchen with a good mix of ingredients. The new method worked great! The robot's translation scores went up. It learned to be more flexible and accurate.
- Kitchen B (Greek Sign Language - GSL): This kitchen was already perfect. The robot was already getting 94% correct because the sentences were so simple and repetitive. Adding more variations actually confused it slightly because the test only looked for one specific answer. It was like trying to teach a master chef a new way to boil water when they already know it perfectly.
- Kitchen C (Argentinian Sign Language - LSA-T): This kitchen was missing half its ingredients. The robot was so confused by the lack of data that adding more sentence variations didn't help. You can't fix a broken recipe if you don't have the main ingredients (the video data) to begin with.
5. The "Hidden" Victory: The Judge's Score
The researchers noticed something interesting. The standard scoring system (BLEU) is like a teacher who only gives points if you use the exact same words as the answer key.
- If the robot said, "It is raining," and the key said, "Rain is falling," the standard teacher gave it zero points.
- But the researchers used a Super-Judge AI to read the answers. The Super-Judge said, "Hey, those mean the same thing! Give it points!"
- The Result: Even when the standard score didn't go up much, the Super-Judge confirmed the robot was actually understanding the meaning much better.
Summary
The paper shows that using an AI to rewrite text (without changing the video) helps Sign Language translation, but only if the data is in the "sweet spot."
- If the data is too simple, it doesn't help.
- If the data is too messy (missing too much), it doesn't help.
- But if the data is just right, it teaches the robot to understand the meaning of the signs, not just memorize the words.
The researchers also noted that this is the first time this specific "AI rewriting" trick has been tried for Sign Language, and they made all their code and data available for others to try.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.