A New Hybrid Intelligent Approach for Multimodal Detection of Suspected Disinformation on TikTok
This paper presents a hybrid intelligent framework that integrates deep learning-based multimodal feature analysis with fuzzy logic to detect suspected disinformation on TikTok by evaluating text, audio, and video cues, ultimately generating detailed reports on disinformation behaviors.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine TikTok as a massive, bustling digital town square where millions of people are shouting, singing, and telling stories. The problem is that some people in this square are telling lies that sound very convincing. The authors of this paper built a new "detective team" to spot these liars, but instead of just one detective, they created a hybrid squad that combines two very different skill sets.
Here is how their system works, explained simply:
The Two-Part Detective Squad
Think of the system as having two main characters working together:
1. The High-Tech Scanner (Deep Learning)
Imagine a super-fast robot camera and microphone that never blinks. This part of the system watches the video and listens to the audio. It doesn't just "see" a person; it measures everything with mathematical precision.
- It listens: It transcribes what is being said, checks if the speaker is pausing too much, if their voice is shaky, or if they are speaking too fast.
- It looks: It tracks the person's eyes, how often they blink, where they are looking, and what their facial expressions are (like a forced smile or a frown).
- It reads: It analyzes the text (captions or speech) to see if the words are vague, repetitive, or don't make logical sense.
2. The Human-like Judge (Fuzzy Logic)
Now, imagine a wise human judge who isn't looking for a simple "Yes/No" answer. Real life is messy; a person isn't just "lying" or "telling the truth." They might be a little suspicious or very suspicious.
- This judge takes all the raw numbers from the robot scanner and translates them into human language.
- Instead of saying "The blink rate is 4.2 per second," the judge says, "The speaker seems a bit nervous."
- Instead of "The text has 15% repetition," the judge says, "The story feels a bit repetitive."
The "Big-5" Personality Test
To decide if someone is a liar, the system compares the speaker's behavior against a psychological profile of a typical "disinformation spreader." The paper uses a famous psychological model called the Big-5, which measures five personality traits:
- Openness (Are they open to new ideas?)
- Conscientiousness (Are they organized and careful?)
- Extroversion (Are they outgoing?)
- Agreeableness (Are they friendly?)
- Neuroticism (Are they anxious or emotional?)
The system asks: "Does this person's behavior match the profile of someone who spreads fake news in political situations?" For example, in the political videos they tested, they found that liars often scored high on Neuroticism (acting anxious or emotional) and low on Agreeableness.
The Final Report: A "Suspicion Score"
After the robot gathers the data and the judge interprets it, the system doesn't just give a red or green light. It produces a detailed report, kind of like a medical diagnosis for a video.
- The Score: It gives a "Suspicion Level" ranging from Low to High.
- The Explanation: It explains why. For instance, it might say: "This video is rated 'High Suspicion' because the speaker showed high anxiety (Neuroticism), spoke with low empathy, and used vague language."
- The "Why" Matters: The paper emphasizes that this system is explainable. You don't just get a result; you get the reasoning behind it, so you can understand how the system reached its conclusion.
What Did They Test?
The researchers tested their "detective squad" in two ways:
The Specific Test: They looked at videos about specific topics like "Does aspartame cause cancer?" or "The Alec Baldwin shooting." They checked if the system could tell the difference between videos spreading lies, videos correcting lies, and videos debunking myths.
- Result: It worked very well, correctly identifying the liars in almost every case.
The Big Test (Scalability): They threw a huge net at the "Invasion of Ukraine" topic, analyzing over 5,000 videos from thousands of different users.
- Result: The system successfully sorted these videos into categories (Low, Medium, High suspicion). It found that about 15% of the videos were highly suspicious, while the rest ranged from low to medium risk.
The Limitations (What the System Can't Do Yet)
The authors are honest about what their tool cannot do:
- Deepfakes: It cannot tell if the video is a "deepfake" (a computer-generated fake video). It assumes the person on screen is real and analyzes their behavior.
- Language Barriers: If the video is in a language other than English (like Russian), the system has to translate it first, which can sometimes lose the original meaning or nuance.
- Obstructions: If a video has subtitles covering the person's face or eyes, the system can't see their facial expressions well, making it harder to judge.
The Bottom Line
This paper presents a new way to fight fake news on TikTok. It combines the speed of AI with the nuance of human psychology to not just say "This is fake," but to explain why it feels fake by looking at how the speaker acts, sounds, and speaks. It's like having a detective who can read a person's body language and speech patterns to spot a liar, and then write a clear report explaining exactly what tipped them off.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.