← Latest papers
🤖 machine learning

YTClickbait21K: Human-Annotated Multimodal Dataset for YouTube Clickbait Detection Across Diverse Channels and Content Categories

The paper introduces YTClickbait21K, a large-scale, human-annotated multimodal dataset comprising over 21,000 YouTube videos with rigorous labeling and diverse metadata, designed to advance automated clickbait detection and content moderation research.

Original authors: Md. Minhazul Islam, Md. Tanbeer Jubaer, Amith Khandakar, Shovon Sarker, Sumaiya Rahman, Md. Masum Mia, Mohamed Arselene Ayari, Hamed Noori

Published 2026-06-16
📖 4 min read☕ Coffee break read

Original authors: Md. Minhazul Islam, Md. Tanbeer Jubaer, Amith Khandakar, Shovon Sarker, Sumaiya Rahman, Md. Masum Mia, Mohamed Arselene Ayari, Hamed Noori

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the internet as a massive, bustling marketplace. In this marketplace, there are thousands of video stalls (YouTube channels) selling everything from cooking tips to gaming highlights. But, just like in any crowded market, some vendors use flashy, misleading signs to trick you into stopping by. They might scream, "You won't believe what happens next!" or show a picture of a giant burger when the video is actually about a tiny sandwich. This is clickbait.

While we've gotten pretty good at spotting these tricks in text (like fake news headlines), spotting them in videos is much harder because it involves both the words and the pictures. Until now, researchers didn't have a big enough "training school" to teach computers how to spot these video tricks effectively.

This paper introduces YTClickbait21K, a massive new tool designed to fix that problem. Here is a simple breakdown of what they did:

1. The "Giant Library" of Videos

Think of the dataset as a giant library containing 21,238 videos.

  • Where did they get them? They didn't just pick random videos. They carefully selected 40 different "stalls" (channels) from 29 different countries. These channels cover everything from serious news and documentaries to fun entertainment, gaming, and education.
  • What's in the box? For every video, they saved the title, the description, the number of views/likes, and the thumbnail (the little picture you see before you click). They didn't download the actual video files (to save space), but they kept the links so researchers can look at them later.

2. The "Human Jury" System

You can't just ask a computer to decide what is clickbait; it needs to learn from humans first. To make sure the labels were accurate, the researchers didn't rely on just one person.

  • The Process: They hired 15 different people (annotators) to act as a jury.
  • The Rule: Every single video was looked at by three different people independently. They couldn't talk to each other while judging.
  • The Scorecard: Each person had a specific checklist to follow. They asked questions like:
    • Is the title lying to make you curious?
    • Does the thumbnail show something that isn't actually in the video?
    • Is the whole thing a mismatch between what it promises and what it delivers?
  • The Verdict: If at least two out of the three people agreed a video was clickbait, it got the "Clickbait" label. If they disagreed, the system used a voting method to decide the final answer.

3. Why This Library is Special

The authors explain that previous attempts to build these libraries had some flaws:

  • Too Small: Some had only a few hundred videos, which isn't enough to teach a smart computer.
  • Too Text-Focused: Many only looked at the words, ignoring the pictures (thumbnails), which are huge clues for video clickbait.
  • Lazy Labeling: Some used computers to guess the labels instead of real humans.
  • YTClickbait21K fixes this by being huge (over 21,000 videos), multimodal (words + pictures), and rigorously checked by real humans.

4. Did the Humans Agree?

The researchers wanted to know if the humans were actually on the same page. They ran a statistical test (called Cohen's Kappa) which is like a "consistency score."

  • The Result: The score was about 0.65. In the world of human judgment, this is considered "substantial agreement." It means that while clickbait can be tricky and subjective (like deciding if a joke is funny), the humans generally agreed on what was misleading and what wasn't.
  • Confidence: The humans were also very sure of their choices, giving themselves high confidence scores (almost 5 out of 5) most of the time.

5. What Can People Do With This?

The paper states that this dataset is now a public "benchmark."

  • For Researchers: It's a standard test set. If a computer scientist builds a new AI to detect clickbait, they can run their AI against this dataset to see how well it does compared to others.
  • For Developers: It helps build better tools to automatically flag misleading videos on platforms.

Summary

In short, the authors built a massive, high-quality training manual for computers. They collected thousands of real-world YouTube videos, had teams of humans carefully label them based on strict rules, and made the whole collection available for free. This allows researchers to finally teach computers how to spot the difference between a genuine video and a misleading one, using both the text and the pictures.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →