Auditing Exposure to Harmful Content on TikTok using Multimodal Language Models: A Cross-National, Age-Stratified Study
This study demonstrates that multimodal large language models can provide a cost-effective and scalable method for auditing TikTok's exposure to harmful content across France, Italy, and Sweden, revealing that keyword-based searches significantly increase harmful content exposure compared to passive scrolling and that safety filters often under-count explicit harms.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Every day, millions of young people open an app to watch short videos, trusting that the screen will show them things they might enjoy. Behind the scenes, complex computer programs decide which videos appear next, learning from every tap, pause, and skip to build a personalized feed. While these systems are designed to be engaging, they can sometimes push users toward content that is dangerous or upsetting, such as material promoting self-harm, eating disorders, or violence. Because these algorithms operate as a "black box" inside the company's servers, independent researchers have struggled to see exactly what young users are seeing. It is difficult to check every video, and the rules for what counts as harmful can vary depending on the language and the country. Without a clear way to measure this exposure, it is hard to know if safety measures are actually working or if they are failing the very people they are meant to protect.
To solve this, a team of researchers from Politecnico di Milano in Italy set out to audit TikTok in three different European countries: France, Italy, and Sweden. They wanted to see what videos the app would show to users of different ages, from early teens to adults. Instead of hiring armies of people to watch thousands of videos, they turned to a new kind of artificial intelligence called a multimodal language model. These are advanced computer systems that can read text, listen to audio, and look at images all at once, much like a human would. The researchers first tested several of these AI systems on a small set of videos that had already been labeled by human experts. They found that one specific model, when shown eight still images taken from a video along with the text description, could identify harmful content with a level of accuracy that was good enough to be used on a massive scale. This approach allowed them to analyze nearly 37,000 videos for the cost of about fifty dollars, a task that would have been prohibitively expensive and slow using human reviewers alone.
The study used computer accounts that pretended to be real users with specific ages: thirteen, sixteen, nineteen, and forty. These accounts were set up in France, Italy, and Sweden to see how the app behaved in different places. The researchers let these accounts scroll through their main video feeds without doing anything else, recording what the algorithm chose to show them. They also performed a second test where the accounts actively searched for specific words related to harmful topics, such as "suicide" or "gambling," to see how the app responded to direct requests. The results revealed a stark difference between what users see when they passively scroll and what they find when they actively look. When the accounts simply scrolled, the amount of harmful content varied significantly by country. Italy showed the highest rates of harmful videos for every age group, with nearly half of the videos shown to a nineteen-year-old persona being flagged as harmful. In contrast, the rates in France and Sweden were lower, though still present.
However, the situation changed dramatically when the accounts started searching. When the researchers typed in keywords related to harm, the app immediately returned a flood of harmful content, with the rate jumping to between thirty-five and fifty-six percent. This was a massive increase compared to the passive scrolling, showing that the app is very good at finding this material when asked for it, even without any visible warnings or blocks at the moment of search. Interestingly, this spike in harmful content was temporary. Once the search was over and the accounts went back to scrolling, the feed returned to its normal, lower level of harmful content within minutes. This suggests that the algorithm does not permanently "poison" a user's feed after a single search, but it does provide immediate access to dangerous material when requested.
The study also found that the age of the user mattered less than the country they were in. In France and Sweden, the amount of harmful content seen by younger teens was lower than what adults saw, which is what safety rules are supposed to achieve. But in Italy, the youngest users saw just as much, or even more, harmful content than the adults, indicating that the safety filters were not working as intended for that specific region. The researchers noted that the artificial intelligence they used was not perfect; it missed some videos and occasionally refused to analyze others, particularly those with very explicit nudity. This means the numbers they reported are likely conservative estimates, and the real amount of harmful content could be slightly higher. Despite these limitations, the study proves that using advanced AI to audit social media platforms is a feasible and affordable way to monitor safety across different languages and countries. It highlights that while the app's search function is a major gateway to harmful content, the passive feed also varies wildly by location, with Italy showing the most significant exposure for young users. The findings suggest that platform companies need to look more closely at how they moderate search results and how they tailor their safety measures for different regions, rather than assuming a one-size-fits-all approach is protecting children everywhere.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.