← Latest papers
💻 computer science

SemanticMoments: Training-Free Motion Similarity via Third Moment Features

To address the failure of existing models to disentangle motion from appearance, the authors propose **SemanticMoments**, a training-free method that achieves superior motion-based video retrieval by computing higher-order temporal statistics over pre-trained semantic features.

Original authors: Saar Huberman, Kfir Goldberg, Or Patashnik, Sagie Benaim, Ron Mokady

Published 2026-02-11
📖 4 min read☕ Coffee break read

Original authors: Saar Huberman, Kfir Goldberg, Or Patashnik, Sagie Benaim, Ron Mokady

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are watching a video of a person drinking coffee. If I asked you to find another video that is "similar," most AI models today would show you another person in a kitchen, perhaps wearing similar clothes, or sitting at a similar table. They are looking at the "scenery" (the coffee cup, the kitchen, the person's shirt).

But what if I asked you to find a video that has the same motion? You wouldn't care if the person was a man or a woman, or if they were in a cafe or a park. You would look for the specific rhythm and pattern of the arm lifting the cup to the mouth.

This paper, "SemanticMoments," is about teaching AI to stop being "distracted by the scenery" and start focusing on the "dance" of the action.

The Problem: The "Photographic Memory" Bias

Current AI models are like people with incredible photographic memories but very little sense of rhythm.

Think of it like this: If you show an AI a video of a person playing a cello, the AI sees the "cello" and the "musician" and says, "Aha! Music!" It doesn't actually need to see the bow moving to know what's happening; it just recognizes the objects. This is called Appearance Bias. Because of this, if you ask the AI to find "the motion of playing a cello," it might just give you a thousand pictures of different cellos, even if the people in them are sitting perfectly still.

On the other hand, older methods tried to look only at "pixels moving" (like watching a grainy black-and-white security camera). These models see the movement, but they have no idea what is moving. They can't tell the difference between a person waving their hand and a tree swaying in the wind.

The Solution: SemanticMoments (The "Rhythm Tracker")

The researchers created a new way to describe video called SemanticMoments. Instead of just looking at a single average of the video, they look at the "Statistical Moments" of the features.

To understand this, let's use a Metaphor: The Rollercoaster.

Imagine you are describing a rollercoaster ride to a friend:

  1. The First Moment (The Average): This is like saying, "The ride stays at a medium height." It tells you the general location, but nothing about the excitement. (This is what most AI does now).
  2. The Second Moment (The Variance): This is like saying, "The ride goes up and down a lot!" It captures the energy and the intensity of the movement.
  3. The Third Moment (The Skewness): This is like saying, "The ride has sudden, sharp drops rather than smooth hills." It captures the shape and the direction of the movement.

By combining these three "moments," the researchers create a mathematical "fingerprint" of the motion. It doesn't care if the rollercoaster is red or blue, or if it's in Florida or France; it only cares about the pattern of the ups and downs.

How they proved it works

The researchers did two things to test their "Rhythm Tracker":

  1. The "SimMotion-Synthetic" Test: They used AI to create "trick" videos. They made videos where the motion was identical (e.g., a person walking), but they changed everything else—the art style, the clothes, the camera angle, and even the person. They found that while other AIs got confused by the new clothes or styles, SemanticMoments correctly identified that the "walk" was the same.
  2. The "SimMotion-Real" Test: They tested it on real-world videos. Even when the videos were messy and uncoordinated, SemanticMoments was much better at finding the "soul" of the movement than the existing industry standards.

Why does this matter?

This isn't just about finding videos. This technology is a building block for:

  • Better AI Video Creators: Helping tools like Sora or Runway understand exactly how a specific movement should look.
  • Robotics: Helping a robot understand the difference between "picking up a glass" and "crushing a glass" by focusing on the dynamics of the hand.
  • Search Engines: Allowing you to search for "a person jumping in slow motion" and actually getting the jump, not just a picture of a person in a gym.

In short: SemanticMoments teaches AI to stop looking at the actors and start watching the dance.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →