← Latest papers
💻 computer science

Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence

Tora3 is a trajectory-guided audio-video generation framework that enhances physical coherence and motion-sound synchronization by using object trajectories as a shared kinematic prior to jointly control visual motion and acoustic events, supported by a newly curated large-scale dataset called PAV.

Original authors: Junchao Liao, Zhenghao Zhang, Xiangyu Meng, Litao Li, Ziying Zhang, Siyu Zhu, Long Qin, Weizhi Wang

Published 2026-08-26
📖 6 min read🧠 Deep dive

Original authors: Junchao Liao, Zhenghao Zhang, Xiangyu Meng, Litao Li, Ziying Zhang, Siyu Zhu, Long Qin, Weizhi Wang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine watching a video of a hammer striking a nail. In a perfect world, the sound of the impact happens at the exact moment the metal meets the wood, and the sharpness of the noise matches how hard the blow was. For years, computers have struggled to create videos where the movement and the sound feel like they belong to the same physical reality. While artificial intelligence has become very good at generating beautiful moving pictures or creating realistic soundscapes from text, combining the two has often resulted in a disconnect. The objects might move in ways that defy gravity, or the sounds might arrive a fraction of a second too late, or they might be too quiet for a loud crash. The core problem has been that the computer generating the video and the computer generating the audio were working in isolation, each guessing what the other should do without a shared understanding of how the object is actually moving through space.

A team of researchers at Alibaba Group and Fudan University has addressed this gap with a new system called Tora3. Their approach is built on a simple but powerful idea: use the path an object takes as a shared blueprint for both the video and the audio. Instead of letting the computer guess how a car should drive or a ball should bounce, the system first defines the exact trajectory, or the line the object follows, and then uses that line to dictate both the visual motion and the accompanying sound. This method treats the path not just as a guide for where the object goes, but as a source of physical information that tells the system when to make a sound and how loud it should be. By anchoring both the sight and the sound to the same mathematical description of movement, the researchers have created videos where the physics feel real and the audio is perfectly synchronized with the action.

The researchers found that previous methods often failed because they tried to generate the video and audio separately and then hoped they would match up, or they used the movement path only as a visual guide for the video while ignoring its potential to guide the sound. Tora3 changes this by using the trajectory as a "kinematic prior," which is a technical way of saying it uses the path as a fundamental rule that both the video and audio generators must follow. When the system sees an object moving along a specific line, it calculates the speed and acceleration at every point. If an object is speeding up, the system knows the sound should get louder or change in tone. If an object hits a surface, the system knows exactly when to trigger a sharp impact sound. This shared reliance on the path ensures that the visual and auditory elements are not just happening at the same time, but are reacting to the same physical forces.

To make this work, the team developed a specific way to feed this path information into the computer's brain. For the video part, instead of building a complex new machine to understand movement, they simply took the visual information from the first frame of the video and moved it along the path of the object as the video progresses. This is a direct and efficient way to tell the computer, "This object is here, and it is moving to there," without needing extra layers of processing that might confuse the system. For the audio, the system looks at the path and calculates how fast the object is moving and how quickly it is speeding up or slowing down. These calculations are then used to shape the sound, ensuring that a fast-moving object produces a different sound than a slow one, and that a sudden stop creates a distinct noise.

The team also created a new method for blending these instructions into the final video. They realized that some parts of the video need to stick strictly to the path to ensure the object moves correctly, while other parts, like the background, should remain flexible to look natural. Their solution was to use two different rules for the video generation process: one rule that tightly binds the object to its path, and another rule that allows the rest of the scene to flow naturally. This hybrid approach prevents the video from looking stiff or unnatural while still ensuring the main action follows the intended trajectory.

To train this system, the researchers gathered a massive collection of 460,000 video clips that focused on movement. They used advanced tools to automatically find the objects in these videos and map out their paths, creating a dataset rich in motion patterns. This allowed the system to learn from a wide variety of real-world movements, from a car driving down a road to a seal slapping water. When tested against other leading systems, Tora3 produced videos that were not only visually sharper but also had a much stronger connection between the movement and the sound. The system was better at matching the timing of sounds to the moment of contact and at adjusting the volume of the sound to match the intensity of the movement.

The results suggest that giving artificial intelligence a shared physical rulebook, in this case the path of an object, is a powerful way to improve the realism of generated content. The system does not just make things look good; it makes them feel like they exist in a world governed by consistent laws of physics. While the technology is still being refined, the success of this approach highlights a promising direction for the future of digital media, where the line between what we see and what we hear becomes indistinguishable from reality. The researchers plan to continue this work by exploring how other physical factors, such as the material of an object or the way sound travels through a room, can be added to this shared framework to create even more convincing experiences.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →