Viewport-Aware Immersive Video Content Identification Using Transformation-Robust Keypoints
This paper proposes a deep-learning-based method for identifying transformed immersive video content by utilizing viewport-aware preprocessing and transformation-robust feature points, achieving a 97.5% recognition rate and improved computational efficiency without relying on metadata.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine trying to find a specific book in a library where every copy has been stretched, squashed, rotated, or had its pages torn out, and the titles have been erased. This is the daily reality for computers trying to identify immersive video content. Unlike standard flat videos, immersive videos capture a full sphere of a scene, allowing viewers to look in any direction. To store and send these videos, computers flatten the sphere onto a rectangular image, a process that inevitably distorts the picture, stretching the top and bottom while keeping the middle relatively normal. When a user watches such a video, they only see a small window, or "viewport," of that flattened image. If someone then takes that video, crops it, changes its resolution, or rotates the view, a computer trying to recognize it faces a nightmare of visual confusion. Traditional methods, which rely on finding specific patterns in an image, often fail because the patterns they were trained to find have been warped beyond recognition or hidden behind the distortion.
This challenge has become critical as immersive media spreads across the internet. Content creators, copyright holders, and platform managers need a way to know if a video they see is a copy of something they own, even if that copy has been heavily altered. If a video is cropped to show only a corner of the original scene, or if the projection format is changed entirely, standard identification tools often give up. They cannot tell the difference between a new video and a modified version of an old one. Without a reliable way to identify these transformed copies, protecting intellectual property and managing massive libraries of 360-degree content becomes nearly impossible.
To solve this, researchers at Soongsil University in Seoul developed a new system designed specifically to navigate the quirks of immersive video. Instead of trying to match the entire distorted image at once, their method breaks the problem down into manageable pieces. First, the system selects only the most important moments in a video, ignoring the boring or repetitive parts to save time. Then, it takes the flattened, distorted video frame and mentally slices it into twelve smaller, rectangular windows, each looking at a different part of the sphere. This step is crucial because it removes the extreme stretching found at the edges of the flattened image, presenting the computer with clearer, more natural-looking views of the scene.
The system then acts like a careful editor, discarding the windows that are too blurry, empty, or distorted to be useful. It focuses only on the windows that contain clear structures or recognizable objects. From these selected windows, the computer extracts "key points"—distinctive visual features like the corner of a building or the edge of a tree. However, the researchers realized that not all key points are created equal. Some points might look clear in one view but vanish or shift wildly when the video is rotated or cropped. To handle this, they trained a computer model to predict which points would remain stable and recognizable even after the video undergoes heavy transformation. The system learns to trust only the points that have proven themselves reliable across different conditions, ignoring the ones that are likely to cause confusion.
When a new video arrives to be identified, the system runs it through the same process: it slices the view, picks the best windows, and extracts only the most robust points. It then compares these points against a database of known videos. Because the system has already filtered out the unreliable points and focused on the most stable features, it can find matches even when the query video has been cropped, rotated, or compressed. The final step involves checking if the matched points fit together in a way that makes geometric sense on a sphere and if they appear in the correct order over time. This multi-layered verification ensures that the system isn't just finding a lucky coincidence but is actually identifying the same content.
The results of this approach were tested against eighteen different types of video alterations, ranging from simple brightness changes to complex projection conversions and severe cropping. The system successfully identified the correct original video in 97.5 percent of the cases. It made a mistake in identifying the content only 2 percent of the time and failed to find a match in just 0.5 percent of the cases. Perhaps just as importantly, the system was fast. It took an average of about 1.03 seconds to identify a video, a significant improvement over older methods that could take nearly six seconds and still fail to recognize the content.
The researchers also tested what would happen if they skipped their special steps. When they tried to match videos without breaking them into smaller windows or without filtering out the unstable points, the success rate plummeted to 59.3 percent, and the processing time jumped to over five seconds. This confirmed that their specific strategy of focusing on stable regions and robust features was the key to the success. The system performed best when videos were simply compressed or had their colors changed, but it did struggle slightly more when large portions of the video were cut away, as there was simply less visual evidence left to work with. Even in these difficult cases, however, the system remained far more effective than previous methods.
This work demonstrates that by understanding how immersive video is distorted and by being selective about which parts of the image to trust, computers can become much better at recognizing content. The method does not rely on file names or hidden metadata, which can be easily stripped away; instead, it relies entirely on the visual structure of the video itself. This makes it a powerful tool for protecting copyright and managing digital libraries in an era where 360-degree content is becoming increasingly common. The findings suggest that with the right approach to handling distortion and selecting features, the problem of identifying transformed immersive video is solvable, offering a reliable way to track and protect content as it moves across the digital landscape.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.