← Latest papers
💻 computer science

Learning Athletic Humanoid Tennis Skills from Imperfect Human Motion Data

This paper introduces LATENT, a system that enables humanoid robots to learn robust and naturalistic tennis skills by leveraging imperfect, fragmentary human motion data as priors for primitive skills, which are then refined and composed to achieve stable multi-shot rallies on the Unitree G1 robot.

Original authors: Zhikai Zhang, Haofei Lu, Yunrui Lian, Ziqing Chen, Yun Liu, Chenghuai Lin, Han Xue, Zicheng Zeng, Zekun Qi, Shaolin Zheng, Qing Luan, Jingbo Wang, Junliang Xing, He Wang, Li Yi

Published 2026-03-17
📖 4 min read☕ Coffee break read

Original authors: Zhikai Zhang, Haofei Lu, Yunrui Lian, Ziqing Chen, Yun Liu, Chenghuai Lin, Han Xue, Zicheng Zeng, Zekun Qi, Shaolin Zheng, Qing Luan, Jingbo Wang, Junliang Xing, He Wang, Li Yi

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine trying to teach a robot how to play professional tennis. The biggest problem? We don't have perfect video footage of how humans actually do it.

Real tennis matches are chaotic. Players run huge distances, swing with lightning speed, and their wrists twist in ways that are incredibly hard to film accurately with cameras. Trying to record a perfect, full-length tennis match for a robot to copy is like trying to film a hummingbird's wingbeat with a shaky, low-resolution camera from 50 feet away. You get the general idea, but the details are blurry.

The paper introduces a system called LATENT (which stands for Learns Athletic humanoid TEnnis skills from imperfect human motion daTa). Think of LATENT as a genius cooking instructor who teaches a robot chef not by showing a perfect recipe, but by showing them a few rough sketches of how to chop, stir, and flip, and then letting the robot figure out the rest.

Here is how it works, broken down into simple steps:

1. The "Rough Sketch" Data (Imperfect Motion)

Instead of trying to film a whole tennis match, the researchers asked five amateur players to just practice the basic moves in a small room:

  • How to swing a forehand.
  • How to do a backhand.
  • How to shuffle sideways.
  • How to cross your feet to run.

They filmed these short clips. The data is "imperfect" because:

  • It's blurry: The cameras couldn't capture the tiny, fast wrist movements perfectly.
  • It's incomplete: The clips don't show how to use these moves to hit a ball; they just show the moves themselves.

2. The "Secret Language" (Latent Action Space)

The robot needs to learn these moves, but if it tries to copy the blurry wrist movements exactly, it will fail. So, the researchers created a "Secret Language" (a latent space).

Imagine the robot has a library of "move cards."

  • Instead of memorizing exact muscle coordinates, the robot learns to pick a card that says "Do a Forehand Swing."
  • The robot knows that this card usually results in a good swing, but it also knows that the "wrist part" of the card is a bit fuzzy.

3. The "Coach's Correction" (High-Level Policy)

This is the magic part. The robot has a High-Level Coach (the AI brain) that watches the game.

  • The Coach picks the right "move card" from the library (e.g., "Run to the ball and swing forehand").
  • Crucially, the Coach is allowed to fix the fuzzy parts. If the "move card" says "swing wrist like this," but the ball is coming fast, the Coach says, "Actually, twist the wrist more to hit it harder."

The researchers built a special rule called a "Latent Action Barrier." Think of this as a safety net. It tells the Coach: "You can fix the wrist, but don't go crazy. Stay close to the human's natural style." This prevents the robot from inventing weird, jerky, robot-like movements that look unnatural.

4. The "Training Camp" (Sim-to-Real Transfer)

Before the robot plays against a real human, it trains in a video game simulation. But video games aren't perfect. To make sure the robot doesn't get confused when it steps into the real world, the researchers made the simulation chaotic:

  • They made the robot's joints slippery or sticky randomly.
  • They made the tennis ball heavier or lighter.
  • They added "static" to the robot's vision so it couldn't see the ball perfectly.

This is like training a swimmer in a pool with choppy water, heavy weights, and foggy goggles. When they finally get to the calm, clear ocean (the real world), they feel like they are swimming in slow motion because they are so well-prepared.

The Result

When they put this system on a Unitree G1 robot (a real, walking robot), the results were surprising:

  • The robot could rally with a human player (hit the ball back and forth) multiple times.
  • It could hit balls moving at high speeds (over 30 mph).
  • It moved naturally, looking like an athlete rather than a stiff machine.

Why This Matters

Previously, teaching robots sports required perfect data or impossible physics (like robots having infinite strength). LATENT proves that you can teach a robot to be an athlete using imperfect, messy, human data, as long as you give the robot a smart way to fill in the gaps and correct its own mistakes.

It's the difference between trying to copy a painting pixel-by-pixel (which fails if the original is blurry) versus teaching an artist the technique of painting, letting them practice, and then letting them adjust their brushstrokes to fit the canvas.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →