CHASM: Unveiling Covert Advertisements on Chinese Social Media
This paper introduces CHASM, a novel dataset of nearly 5,000 real-world covert advertisements from the Chinese social media platform Rednote, to demonstrate that current Multimodal Large Language Models struggle to detect these deceptive posts and to highlight the need for improved fine-tuning and defense mechanisms against this emerging threat.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you're scrolling through your favorite social media app, looking for inspiration on what to wear or what skincare to try. You see a post that looks like a genuine friend sharing their honest opinion: "I tried this new toner, and wow, my skin feels amazing! Here's a photo of my morning routine."
It feels authentic. It feels like a recommendation from a peer. But what if that "friend" is actually a paid advertiser in disguise, and the "honest opinion" is a carefully crafted script designed to sell you something?
This is the problem the paper CHASM tackles. It's like a detective story where the criminals are so good at wearing masks that even the smartest AI detectives are getting fooled.
Here is the story of the paper, broken down into simple parts:
1. The Villain: The "Wolf in Sheep's Clothing"
In the world of social media (specifically a popular Chinese app called Rednote, similar to Instagram or TikTok), there is a rising threat called Covert Advertisements.
- Traditional Ads: These are obvious. They say "AD" or "SPONSORED." They are like a billboard on the highway. You know they are trying to sell you something, so you can choose to ignore them.
- Covert Ads: These are the wolves in sheep's clothing. They look exactly like normal posts. They don't say "Buy this!" They say, "I love this!" or "Check out my outfit!" But hidden inside the picture, the caption, or the comments are secret links, brand names, or instructions on how to buy.
The goal of these ads is to trick you into thinking, "Oh, this is just a real person sharing a tip," so you lower your guard and buy the product.
2. The Detective Tool: The CHASM Dataset
The researchers realized that current AI models (Large Language Models) are terrible at spotting these disguised ads. They thought, "How can we teach the AI to spot the wolf?"
To do this, they built CHASM. Think of CHASM as a giant training gym for AI detectives.
- The Workout: They collected nearly 5,000 real posts from Rednote.
- The Mix: Some posts were real ads in disguise (the "bad guys"), and some were just regular people sharing their lives (the "good guys").
- The Challenge: They made sure the "good guys" looked very similar to the "bad guys." It's like a game of "Spot the Difference" where the differences are tiny.
- Privacy: They scrubbed all the real names and faces out of the photos so no real people were harmed, just like blurring faces in a police sketch.
3. The Big Test: Can AI See Through the Mask?
The researchers took the smartest AI models in the world (like GPT-4o, DeepSeek, and others) and threw them into the CHASM gym. They asked: "Can you tell which of these posts is a fake ad?"
The Result? The AI failed miserably.
- Even the "super-smart" AIs only got about 60% right. That's barely better than flipping a coin!
- The AI kept getting tricked. It thought real friends were selling things, and it missed the actual sales pitches hidden in the comments.
- The Analogy: It's like giving a master chef a blindfold and asking them to taste a dish and tell you if it has salt in it. The AI just couldn't taste the subtle "flavor" of a hidden ad.
4. The Training Montage: Fine-Tuning
Since the AI was failing, the researchers tried a new strategy: Fine-Tuning.
- Instead of just asking the AI to guess, they showed it the CHASM dataset and said, "Look at these examples. Here is why this one is an ad, and here is why that one isn't. Now, try again."
- The Result: This was like giving the detective a magnifying glass and a handbook. The AI's performance jumped up significantly. One model (Qwen2.5) went from a 42% success rate to 75%.
- The Catch: Even with training, the AI still struggled with the hardest cases. It missed subtle clues hidden in the comments section or failed to notice that a post was focusing on only one brand (a classic ad trick) instead of comparing different brands (a real user habit).
5. Why This Matters
Why should we care if an AI can spot a fake ad?
- Trust: If we can't tell what's real and what's a sales pitch, we stop trusting social media.
- Money: People are losing money on products they didn't know were being sold to them.
- The Law: In many countries, hiding the fact that you are being paid to promote something is actually illegal.
The Bottom Line
The paper concludes that current AI isn't ready to police social media ads on its own. It's too easily fooled by clever disguises.
However, the CHASM dataset is a huge step forward. It's like a new "training manual" that helps AI get better at spotting these tricks. The researchers hope that by sharing this data, platforms can build better defenses, ensuring that when you see a post, you know if it's a friend's advice or a salesman's pitch.
In short: The internet is full of wolves in sheep's clothing. This paper built a training camp to teach our AI dogs how to sniff them out, but the wolves are still very tricky, and we need to keep training.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.