Research on table tennis motion recognition under limb occlusion
This paper presents a reliable table tennis motion dataset and proposes a spatio-temporal graph convolutional neural network model that effectively predicts occluded joint positions to improve motion recognition accuracy under limb occlusion.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the high-stakes world of competitive table tennis, a split-second decision often determines the outcome of a match. Athletes must constantly scan their opponent's body, reading subtle cues in the swing of an arm or the shift of a weight to predict where the ball will go next. However, the game is rarely played with a clear, uninterrupted view. The net, the table, and the opponent's own body frequently block the line of sight, hiding the very movements that hold the secret to the next shot. This phenomenon, known as occlusion, creates a blind spot where critical information disappears. For decades, researchers have tried to teach computers to see through these gaps, using artificial intelligence to fill in the missing pieces of a moving picture. The challenge lies in teaching a machine not just to recognize a pattern when it is complete, but to understand the whole story even when parts of it are missing, relying on the logic of how human bodies move together over time.
A team of researchers from Shandong Sport University and the Shanghai University of Sport set out to solve this specific puzzle within the context of table tennis. They began by capturing the precise movements of thirty professional players using a specialized system of infrared cameras that tracked the position of reflective markers placed on the athletes' bodies and rackets. This process generated a massive library of data, recording the exact location of fourteen key joints on the body for every single frame of a player's motion. With this digital archive in hand, the team created a controlled experiment to simulate the chaos of a real match. They used computer software to build stick-figure animations of the players, then systematically erased different parts of these figures to mimic the way a net or a body might hide a limb during a game. They created scenarios where the left arm was hidden, where the legs were invisible, where the entire torso was obscured, and even where only the path of the racket remained visible.
Thirty other professional players were then asked to watch these incomplete animations and identify the specific type of stroke being performed, such as a forehand attack or a backhand drive. The results offered a clear map of what information is truly essential. When the full body was visible, the players identified the moves with nearly perfect accuracy. When the left arm was hidden, their performance barely dipped, suggesting that the non-dominant arm is less critical for reading the opponent's intent. However, when the racket's trajectory was the only thing visible, accuracy dropped significantly, though it remained high enough to prove that the path of the ball and racket carries immense weight. The most telling finding came when the researchers removed the lower body or the entire torso; even with the lower limbs obscured, recognition accuracy was 91.4%, and with the torso completely obscured, it remained at 96.2%, indicating that while the lower body provides context, the upper body and racket trajectory alone are sufficient for highly accurate motion identification.
Having established what humans need to see, the researchers turned to artificial intelligence to see if a computer could do the same job, but with a twist. They built a specialized neural network, a type of computer program designed to learn patterns, to act as a digital restorer. The goal was to feed the system an animation with missing joints and have it calculate exactly where those hidden parts should be. The system learned the relationship between the visible bones and the invisible ones, understanding that if the right shoulder moves a certain way, the left arm must move in a specific, coordinated fashion. When tested, the computer successfully reconstructed the missing limbs, filling in the gaps with a level of precision that matched the judgments of human coaches. The reconstructed movements were so accurate that when experts scored the original incomplete animations and the computer-filled versions, the scores were nearly identical, with the computer's version restoring the motion to a state that felt complete and natural.
The study concludes that while the racket and the dominant arm are the primary sources of information for predicting a table tennis stroke, the rest of the body provides the necessary context to make that prediction reliable. The research demonstrates that by understanding the spatio-temporal connections between joints—how they move in relation to each other across time—it is possible to mathematically infer the position of a hidden limb with high confidence. This work does not just offer a way to improve motion tracking technology; it provides a deeper understanding of how athletes themselves process visual information under pressure. It suggests that even when the view is blocked, the brain and the machine can both rely on the remaining visible clues to reconstruct the full reality of the movement, turning a fragmented view into a clear prediction.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.