← Latest papers
💬 NLP

Toward Signing Activity Projection in Sign Language Interaction

This paper investigates the transfer of Voice Activity Projection (VAP) to sign language interaction using the Public DGS Corpus, finding that while hand-cue-based models show promise for predicting signing continuations, accurate turn-taking prediction remains challenging and requires sign-language-specific event definitions beyond speech-derived categories.

Original authors: Takao Obi, Wang Yusong, Koji Inoue, Kotaro Funakoshi

Published 2026-06-09
📖 4 min read☕ Coffee break read

Original authors: Takao Obi, Wang Yusong, Koji Inoue, Kotaro Funakoshi

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a social robot as a new dance partner. In a normal conversation, people take turns speaking. If you stop talking, your partner knows when to jump in or when to wait. Computers are getting pretty good at predicting this "dance" when people speak, using a system called Voice Activity Projection (VAP). It's like a radar that listens to the silence and guesses, "Is the person about to speak again, or is it my turn?"

But what happens when the "dance" isn't spoken words, but sign language?

This paper is a first attempt to teach a robot to predict turn-taking in sign language using the same "radar" technology, but instead of listening to sound, it's watching video.

The Big Challenge: Listening with Eyes

In spoken language, the robot listens for the sound of a voice. In sign language, there is no voice to hear. The "voice" is the movement of hands, the look in the eyes, and the shape of the mouth.

The researchers wanted to see if they could simply swap the robot's "ears" for "eyes." They took a system designed for speech and tried to make it work with sign language data from a public database (the Public DGS Corpus).

How They Did the Experiment

Since the robot couldn't actually "hear" signing, the researchers had to create a simplified version of the game:

  1. The "On/Off" Switch: Instead of trying to understand complex sentences, they turned the video into a simple "On/Off" light. If a person was making a sign, the light was ON. If they were still, the light was OFF.
  2. The Two Questions: They asked the robot two specific questions about the future:
    • Question A (The "Who's Next?" Test): After a moment where both people are silent, who starts again? Is it the same person (HOLD), or does the turn pass to the other person (SHIFT)?
    • Question B (The "Are They Done?" Test): If someone is currently signing, are they about to stop and let the other person speak very soon?

The Tools: What Did the Robot Watch?

The robot didn't just watch the whole video; it focused on specific body parts, like a detective looking for clues:

  • Hands: The main actors in sign language.
  • Eyes: Where the person is looking.
  • Mouth: The shape of the mouth (which often helps in signing).

They trained the robot using these visual clues to predict the future "On/Off" lights.

The Results: A Mixed Bag

The findings were a bit like a student who is great at one subject but struggling with another:

  • The Good News (The "Who's Next?" Test): The robot was actually quite good at guessing who would speak next after a silence, especially when it watched the hands. It turns out that hand movements are a very strong clue for knowing who holds the "floor" in a sign language conversation.
  • The Bad News (The "Are They Done?" Test): The robot struggled to guess if a signer was about to stop and hand over the turn. Even when the robot looked at the eyes or mouth, it couldn't reliably predict the exact moment a turn would end.

Why Was It Hard?

The authors explain that the "rules" they taught the robot were borrowed from spoken language, and they might not fit sign language perfectly.

  • The Wrong Map: They used a map designed for "voice" to navigate "signs." In sign language, the end of a turn isn't just about stopping a sound; it involves complex visual cues like holding a pose, looking at the other person, or specific hand shapes that the simple "On/Off" light didn't capture well.
  • Missing Details: The robot was watching 2D stick-figure outlines (keypoints) of the body. It missed the depth of the hand movements and the subtle nuances of facial expressions that are crucial for knowing when a turn is truly over.

The Bottom Line

This paper is a "proof of concept." It shows that:

  1. We can try to use the same predictive technology for sign language as we do for speech.
  2. Hand movements are a powerful clue for knowing who is talking.
  3. However, simply copying the speech rules doesn't work perfectly. To make a robot truly good at signing conversations, we need to invent new rules and definitions that are specific to sign language, rather than just forcing speech rules onto it.

In short, the robot is learning to dance, but it's still tripping over the specific steps of the sign language dance, because it's trying to learn them using a map meant for a different kind of dance.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →