Let the Bullets Fly: Multimodal Fake News Detection with Temporal-Aligned Generative Danmaku
This paper introduces Genda, a temporal generative framework that simulates real-time Danmaku interactions to overcome latency issues, and leverages the resulting pseudo-streams in a novel DM-FEND model to achieve state-of-the-art multimodal fake news detection on both Chinese and English benchmarks.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the modern digital landscape, news often arrives not as a static article, but as a fast-moving stream of video, sound, and text. These short clips, ubiquitous on platforms like TikTok and Bilibili, compress complex stories into seconds, creating a fertile ground for misinformation to spread before it can be fact-checked. Traditional methods of spotting fake news often rely on analyzing the video or audio itself, looking for digital fingerprints of editing or manipulation. However, sophisticated fakes can look and sound perfect, leaving these forensic tools with nothing to find. A different kind of clue exists in the chaotic, real-time chatter of the audience. On many Asian video platforms, viewers do not just leave comments after a video ends; they send "Danmaku," or bullet comments, that fly across the screen in sync with the video's timeline. These comments offer a unique, second-by-second record of how people react to specific moments, creating a rich layer of social context that moves in lockstep with the story being told.
The challenge for researchers has been that this social chatter arrives too slowly to be useful for real-time detection. By the time enough people have watched a video and sent their bullet comments to create a meaningful pattern, the fake news has often already gone viral. To solve this, a team of researchers at Yangzhou University and Auburn University has developed a new approach that simulates this human reaction process. They created a system that can predict exactly when and how an audience would react to a video, even before any real people have seen it. This system, which they call Genda, acts like a time machine for social interaction. It does not wait for the internet to catch up; instead, it generates a realistic stream of simulated bullet comments that align perfectly with the video's timeline, predicting the intensity and content of user reactions as if a crowd were watching the clip right now.
The researchers built this simulation using two distinct parts. First, a "trigger" component analyzes the video to understand its content, emotional tone, and narrative flow. It predicts the moments where viewers are most likely to feel surprised, confused, or angry, and estimates how strong those reactions will be. Second, a "generator" uses those predictions to write the actual comments. It creates a diverse range of human-like responses, from mild curiosity to heated debate, ensuring that the simulated comments match the specific visual and audio cues of the video at that exact second. The result is a synthetic but highly accurate stream of social feedback that mirrors how a real crowd would behave, filling the gap between the video's release and the accumulation of real user data.
With this simulated social layer in place, the team introduced a new detection model called DM-FEND. This model treats the generated bullet comments not as noise, but as a guide. It uses the simulated reactions to help the computer understand the video, audio, and text more deeply. By aligning the different parts of the video with the predicted human responses, the model can spot inconsistencies that a standard detector might miss. For instance, if a video shows a calm scene but the simulated reactions predict a wave of confusion or skepticism, the system flags that moment as suspicious. The model learns to ignore irrelevant details and focus on the parts of the video where the story and the audience's likely reaction do not match, effectively using the "voice" of the crowd to find the truth.
When tested on real-world datasets of short videos, including both Chinese and English examples, this new method proved significantly more effective than existing techniques. The system achieved an accuracy of 87.01% on one major dataset and 83.43% on another, outperforming previous state-of-the-art models. The researchers found that the key to this success was the ability to model the timing of human reactions. By simulating how people engage with content over time, the system could detect subtle lies that unfold gradually, rather than just looking for static errors. The study suggests that by bridging the gap between the speed of news and the speed of human reaction, it is possible to build a more robust defense against misinformation, even in the earliest moments of a video's life.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.