← Latest papers
💻 computer science

CLARA: Clip-Level Multimodal Alignment with VLM-Derived Rationales for Hateful Video Detection

This paper proposes CLARA, a novel clip-level multimodal framework that enhances hateful video detection by modeling videos as fine-grained sequences with adaptive alignment, contrastive learning for temporal dependencies, and VLM-derived rationales, achieving state-of-the-art performance across multiple datasets.

Original authors: Yuchen Zhang, Shuang Dai, Zeyu Fu, Yunfei Long, Ravi Shekhar, Haralambos Mouratidis

Published 2026-08-18
📖 5 min read🧠 Deep dive

Original authors: Yuchen Zhang, Shuang Dai, Zeyu Fu, Yunfei Long, Ravi Shekhar, Haralambos Mouratidis

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

On the screens of billions of people, a new kind of conversation is taking place. It happens in the rapid, scrolling streams of video platforms where a joke, a news clip, or a personal vlog can shift in an instant from harmless to harmful. Unlike a written post, where the meaning is fixed in words, a video carries a complex layering of sound, moving images, and spoken language. The danger lies in how these elements interact. A hateful message might not be shouted in a single sentence; instead, it can emerge slowly, building up through a sequence of frames, a change in tone, or a visual cue that only makes sense when viewed in the context of what came before. Detecting this kind of content is a massive challenge for the internet, because the most dangerous signals are often brief, hidden, and dependent on the specific moment they appear.

For years, computer scientists have tried to teach machines to spot this kind of harm. Early efforts focused on text, looking for specific angry words. Later, they moved to images and memes, trying to understand how pictures and captions work together. But videos present a unique problem. They are not just a collection of images or a transcript of speech; they are a flowing timeline where meaning changes over time. If a computer looks at a whole video as a single block, it often misses the small, critical moments where the hate actually happens. It is like trying to understand a story by reading only the first and last sentence; the crucial plot points in the middle get lost in the summary.

A team of researchers at the University of Essex and the University of Exeter has proposed a new way to solve this. They call their system CLARA. Instead of treating a video as one big, unchangeable object, CLARA breaks it down into smaller, meaningful pieces. Imagine a video as a long sentence; CLARA does not read the whole thing at once. Instead, it pauses at every natural break in the speech, creating a series of short clips that align with what is being said. By focusing on these small segments, the system can catch the fleeting moments where a hateful idea is introduced, even if that idea is buried in a minute of neutral content.

The researchers found that hateful videos often rely on a mix of clues. A speaker might use a calm voice while showing a disturbing image, or use a neutral phrase while the background music turns aggressive. To handle this, CLARA uses a flexible system that decides, for each short clip, which clues matter most. Sometimes the text is the key; other times, it is the sound or the visual scene. The system does not force a single rule for the entire video. Instead, it adapts, weighing the different types of information differently as the video progresses. This allows it to see how a hateful message evolves, rather than just looking for a static pattern.

To make sure the system understands the bigger picture, the researchers added a second layer of intelligence. They used a powerful artificial intelligence tool, trained to describe and reason about images and text, to generate a high-level summary of the video's intent. This summary acts as a guide, helping the system understand how the small clips fit together to form a larger, potentially harmful narrative. The system then combines these detailed clip-level observations with the high-level summary, creating a complete picture of the video's content. This approach ensures that the system does not just see isolated moments, but understands the flow of the argument or the story being told.

The team tested this new method on three different collections of hateful videos, drawn from various social media platforms around the world. The results were clear: the new system consistently outperformed the best existing methods. It was better at correctly identifying hateful videos without mistakenly flagging harmless ones. The researchers also ran tests to see what would happen if they removed parts of the system. When they took away the ability to break videos into clips, or when they removed the high-level guidance, the system's performance dropped. This confirmed that every part of their design was necessary. The ability to focus on small, specific moments while keeping an eye on the overall context was the key to success.

The study suggests that the future of detecting harmful content lies in this kind of detailed, time-aware analysis. By respecting the way humans actually experience video—piece by piece, moment by moment—machines can finally learn to see the subtle, evolving nature of hate speech. The researchers did not claim to have solved the problem of online hate forever, but they have provided a significantly sharper tool for the job. Their work shows that to understand the danger in a video, you cannot just look at the whole picture; you must watch the story unfold, one small clip at a time.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →