Modeling Engagement with Brand and Organizational TikTok Videos Using Machine-Assisted Theory-Ensemble Annotation
This paper demonstrates how multimodal large language models can overcome the scalability challenges of manually annotating complex audiovisual and theoretical variables in short-form video, enabling the analysis of 10,000 TikTok videos to identify which narrative and structural factors drive engagement and provide actionable insights for content creators.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to understand why some short videos on TikTok go viral while others get ignored. Usually, researchers just look at easy-to-count things: "How many followers does the account have?" or "How old is the video?" But this is like judging a book only by its cover size and publication date, ignoring the story inside.
This paper tries to read the "story" inside thousands of TikTok videos made by brands and organizations in Estonia. Here is a simple breakdown of what they did and what they found.
The Problem: Too Much Video, Too Little Time
Short videos are a mix of music, talking, text, editing, and trends. To understand them, you usually need a human to watch and write down notes about the "vibe," the story structure, or the emotional tone. But doing this for 10,000 videos by hand is impossible—it would take years.
The Solution: The "Super-Reader" Robot
The researchers used a powerful AI (a Large Language Model) to act as a "super-reader." They taught this AI to watch the videos and fill out a very specific, complex checklist based on old-school theories about storytelling, advertising, and psychology.
Think of the AI as a detective with a magnifying glass that can instantly spot 77 different clues in a video, such as:
- Is the speaker talking directly to you?
- Is the music happy or sad?
- Is there a clear "hero" and a "villain"?
- Is the video trying to sell something or just make you laugh?
The Experiment: The "Human vs. Robot" Check
Before trusting the robot, the researchers had three human students watch a small sample of videos and fill out the same checklist.
- The Good News: The robot was very good at spotting obvious things, like "Is there music?" or "Is the camera moving fast?" Humans and robots agreed on these easily.
- The Bad News: When it came to deep, abstract ideas (like "What is the hidden symbolic meaning?" or "What is the philosophical conflict?"), even the humans couldn't agree with each other. The robot struggled here too. It's like asking a group of people to agree on the "true meaning" of a dream; everyone sees something different.
The Findings: What Actually Gets Likes?
The researchers then used the robot's notes to predict which videos would get the most Likes and Comments. They compared the robot's detailed notes against the simple "follower count" baseline.
1. The "Packaging" Matters More Than the "Deep Meaning"
The robot found that the most successful videos weren't necessarily the ones with the deepest philosophical stories. Instead, success was driven by how the video was packaged:
- Audio: Videos with original spoken voices (people talking) got more likes than those with just music.
- Intent: Videos meant to be funny or entertaining performed better than those meant to be purely educational.
- Setting: Videos filmed in recognizable indoor commercial settings (like a shop or office) did well.
2. The "Call to Action" is the Key to Comments
If you want people to comment, the video needs to ask for it. Videos that used "engagement bait" (directly asking questions or prompting a reaction) got significantly more comments. Practical "how-to" videos also sparked more questions and information-seeking comments.
3. The "Familiarity" Factor
The study found a surprising result: Being similar to the average works.
- Videos that looked and felt like the "standard" successful videos in this group (using common audio, common humor, common settings) tended to get more likes than videos that tried to be weirdly unique.
- It's like a restaurant: If you want a steady stream of customers, it's often safer to serve the popular dish everyone expects than to serve a weird, experimental meal that confuses people.
The Bottom Line
This paper shows that we can use AI to turn "vibes" and "stories" into hard data. While the AI can't perfectly understand deep human philosophy, it is very good at spotting the structural ingredients of a video.
For a brand making a TikTok, the takeaway isn't "write a deep, complex story." The takeaway is: Use clear audio, keep it entertaining, set it in a familiar place, and ask people to comment. The "magic" isn't in the hidden meaning; it's in the visible, repeatable patterns that the AI helped identify.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.