← Latest papers
💬 NLP

Leveraging Social Media Data for COVID-19 Studies

This chapter explores the use of social media data during the COVID-19 pandemic by analyzing linguistic, visual, and emotional indicators, reviewing machine learning and NLP methodologies, and outlining future research directions for leveraging these platforms to disseminate reliable information and public awareness.

Original authors: Nur Hafieza Ismail, Nur Shazwani Kamarudin, Nurol Husna Che Rose

Published 2026-06-10
📖 5 min read🧠 Deep dive

Original authors: Nur Hafieza Ismail, Nur Shazwani Kamarudin, Nurol Husna Che Rose

Original paper dedicated to the public domain under CC0 1.0 (http://creativecommons.org/publicdomain/zero/1.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the world during the COVID-19 pandemic as a massive, chaotic storm. In the middle of this storm, social media became the only lighthouse many people had to see the way. This paper is like a map that explores how researchers used that lighthouse (social media data) to understand what was happening inside the storm, specifically focusing on how people felt, what they were saying, and how they were behaving.

Here is a simple breakdown of what the paper says, using everyday analogies:

1. The Problem: A Flood of Noise

When the virus hit, everyone was scared and confused. Social media (like Twitter, Facebook, and Reddit) became the main place people went for news. But it was like walking into a crowded room where everyone is shouting at once. Some people were shouting helpful facts, but others were shouting fake news, rumors, and scary stories.

The paper explains that this "noise" was dangerous. Just like a rumor in a small town can cause a panic, fake news on social media caused real fear and anxiety. In some tragic cases, people even made life-ending decisions because they believed false information. The authors argue that we can't just turn off the radio; instead, we need to learn how to tune in to the right stations.

2. The Tool: The Social Media "Scanner"

The paper describes a process called "Social Media Analysis." Think of this as a high-tech scanner that looks at the massive ocean of posts people make every day.

  • The Volume: It's huge. The paper notes that on Twitter alone, 350,000 new messages are created every minute. That's like a waterfall of text.
  • The Steps: Researchers don't just read everything. They follow a recipe:
    1. Gather: Collect the data (the tweets or posts).
    2. Clean: Wash the data (remove useless words like "the" or "and").
    3. Pick: Choose the important parts (like specific words or images).
    4. Dig: Use computers to find patterns (Data Mining).
    5. Check: Test if their findings are accurate.

3. What They Looked At: Three Types of Clues

The researchers looked at three different kinds of "clues" left by people online:

  • The Words (Linguistic Data): This is like listening to the tone of voice in a crowded room. Researchers used computers to read millions of tweets to see if people sounded stressed, sad, or angry. They found that when new virus cases went up, people's posts sounded more stressed. They also noticed that when governments told people to stay home, the conversations shifted from global news to how people were coping with their daily lives.
  • The Pictures (Visual Data): Sometimes words aren't enough. Researchers also looked at photos. They used "smart cameras" (AI) to count how many people were in a photo or if they were wearing masks. For example, they analyzed millions of photos from six US cities to see if people were actually keeping their distance or gathering in large groups during protests.
  • The Mix (Combined Data): The paper suggests that looking at words and pictures together gives the best picture. It's like trying to understand a movie by only reading the script; you miss the visual emotion. By combining text and images, researchers could get a fuller sense of how people were feeling about the pandemic.

4. The Engine: Machine Learning

How do you read 13 million tweets in a day? You can't. That's where Machine Learning comes in. Think of Machine Learning as a super-fast, tireless robot assistant.

  • The Job: You teach the robot what "anxiety" or "fake news" looks like, and then it scans the data to find it.
  • The Results:
    • One study used this robot to find tweets about stress and found that the stress levels matched the number of new virus cases.
    • Another study taught the robot to spot "fake news" about COVID-19. They even created a special training manual (a dataset called COVIDLIES) so other researchers could teach their robots too.
    • Another team used the robot to read YouTube comments to detect anxiety. They found that a specific type of robot (called "Random Forest") was the best at guessing if someone was feeling anxious, getting it right more than 83% of the time.

5. The Big Takeaway

The paper concludes that social media is a double-edged sword. It can spread panic and fake news faster than a wildfire, but if we use the right tools (like the scanners and robots mentioned above), it can also be a powerful way to understand how people are feeling.

By listening to what people are saying and seeing what they are posting, governments and health officials can get a "real-time" snapshot of the public's mood. This helps them understand if people are scared, if they are following the rules, and if they need more support. The paper emphasizes that while we have a lot of data, we need to use it wisely to help people navigate the storm, rather than just letting the noise overwhelm us.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →