SMART: MLLM-guided Temporal Alignment for Unifying Sign Language Recognition and Spotting
The paper proposes SMART, a unified framework that leverages MLLM-generated motion descriptions and a Multi-Scale Temporal Adapter to enhance continuous sign language recognition and spotting through efficient small-batch video-text alignment and joint temporal localization.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
For millions of people who are deaf or hard of hearing, sign language is not merely a collection of hand gestures but a complete, flowing language with its own grammar, rhythm, and nuance. Just as spoken words blend together in a sentence without pauses, signs in a continuous stream often merge into one another, making it difficult for computers to tell where one sign ends and the next begins. For decades, researchers have tried to teach machines to understand these videos, but they have faced a stubborn hurdle: the data available to train these systems usually only provides the final sentence, like a transcript, without telling the computer exactly which moment in the video corresponds to which specific word. This lack of precise timing forces computers to guess, often resulting in a jagged, uneven understanding where the machine might recognize a word but fail to know how long it lasted or where it started. Without this fine-grained knowledge, building a truly fluent translator for sign language remains out of reach.
A team of researchers at Dankook University in Korea has now proposed a new approach to bridge this gap, introducing a system they call SMART. Their work tackles the twin challenges of recognizing what signs are being made and pinpointing exactly when they happen in a video. Instead of relying on the old method of guessing word boundaries based on sentence length alone, the researchers built a framework that uses a powerful type of artificial intelligence known as a multimodal large language model. Before the system even begins to learn, this AI acts as a careful observer, watching the sign language videos and writing down detailed descriptions of the hand movements and facial expressions happening at every single moment. These descriptions serve as a rich, semantic guide, teaching the computer to associate the visual motion with specific language concepts, much like a teacher explaining the meaning of a gesture while pointing to it.
The researchers then designed a specialized engine to process these videos, one that pays close attention to how movements change over time. While standard video analysis tools often look at frames in isolation, this new system is built to understand the flow of motion across multiple scales, capturing both quick, sharp movements and slower, more deliberate gestures. By combining this temporal awareness with the detailed text descriptions generated earlier, the system learns to align the visual video with the language meaning in a stable and efficient way, even when the computer is processing only a small number of videos at a time. This allows the model to build a much clearer picture of the sign language dynamics than was previously possible.
Perhaps the most significant innovation in this work is how the system handles the two tasks of recognition and timing simultaneously. The researchers created a unified structure where the part of the system that identifies the words feeds its knowledge directly into the part that locates the boundaries. In previous attempts, these two tasks were often treated separately, or the timing information was too sparse to be useful. Here, the system uses the confidence it has in recognizing a specific sign to sharpen the edges of that sign in time, effectively telling the computer, "I am sure this is the word 'stop,' so let's mark exactly when it starts and stops." This creates a feedback loop where the ability to recognize words improves the ability to find their boundaries, and the precise timing of those boundaries, in turn, helps the system recognize the words more accurately.
The team tested this new framework on four different sign language datasets, including videos in German, Chinese, and two distinct Korean sign language collections. The results showed that their method outperformed all existing state-of-the-art systems in both recognizing the sequence of signs and locating them precisely within the video. On one large dataset, the system reduced the error rate in recognizing signs to less than one percent, a significant improvement over previous methods. Furthermore, when asked to pinpoint the exact start and end times of signs, the system achieved a level of accuracy that far surpassed traditional methods, proving that the two tasks are indeed complementary. The researchers found that by using the language descriptions as a guide and letting the recognition and spotting tasks support each other, they could overcome the limitations of weak supervision that had long plagued the field.
This work suggests that the path to fluent sign language understanding lies not just in better cameras or faster processors, but in teaching machines to understand the semantic richness of human movement. By generating detailed descriptions of motion and using them to guide the learning process, the SMART framework demonstrates that computers can learn to see sign language with a level of detail that approaches human perception. The findings indicate that combining the broad understanding of language models with the precise timing of video analysis creates a powerful synergy, offering a promising new direction for building tools that can finally break down communication barriers for the deaf and hard-of-hearing community.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.