← Latest papers
💬 NLP

Assessing Capabilities of Large Language Models in Social Media Analytics: A Multi-task Quest

This paper presents the first comprehensive evaluation of modern large language models across three core social media analytics tasks—authorship verification, post generation, and user attribute inference—using a newly collected Twitter dataset to mitigate data bias and establish reproducible benchmarks.

Original authors: Ramtin Davoudi, Kartik Thakkar, Nazanin Donyapour, Tyler Derr, Hamid Karimi

Published 2026-04-22
📖 4 min read☕ Coffee break read

Original authors: Ramtin Davoudi, Kartik Thakkar, Nazanin Donyapour, Tyler Derr, Hamid Karimi

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you walk into a massive, chaotic digital town square (like Twitter/X) where millions of people are shouting, whispering, sharing memes, and arguing. For years, computers have tried to make sense of this noise, but they often got confused by the slang, the short sentences, and the fact that people change their minds (and their writing styles) quickly.

Recently, super-smart AI computers called Large Language Models (LLMs)—think of them as "digital geniuses" who have read almost everything on the internet—have arrived. The big question is: Can these digital geniuses actually understand and mimic the messy, real world of social media, or are they just fancy parrots?

This paper is a massive "report card" where the authors put seven of the smartest AI models (including GPT-4, Gemini, and Llama) through a three-part challenge to see how well they handle social media.

Here is the breakdown of their three challenges, explained with simple analogies:

Challenge 1: The "Handwriting Test" (Authorship Verification)

The Goal: Can the AI look at a tweet and say, "Yes, this was definitely written by User A," or "No, this is a fake"?

  • The Analogy: Imagine you are a detective trying to identify a suspect. You have a notebook of their old writings (the "training" data). Then, someone hands you a new note. Can you tell if it's their handwriting or a forgery?
  • The Twist: The authors didn't just test the AI on old notes. They gave it brand new notes written after the AI stopped learning (like giving a student a test on a textbook chapter they haven't read yet). This checks if the AI is actually smart or just memorizing answers.
  • The Result: GPT-4 was the star detective. It was incredibly good at spotting the "handwriting," even on new notes. Other models were okay, but some (like the older BERT) were easily fooled.

Challenge 2: The "Impersonator Contest" (Post Generation)

The Goal: Can the AI write a tweet that sounds exactly like a specific real person?

  • The Analogy: Imagine you hire a ghostwriter to write a tweet for you. You give them your diary, your bio, and a list of your friends. They write a new tweet. You then ask your friends: "Did I write this?"
  • The Test: The researchers asked real people to look at AI-generated tweets and guess if they were written by themselves or a robot. They also used math to measure how similar the words were.
  • The Result: This was a mixed bag.
    • Gemini and Llama were the best at tricking humans. People often thought, "Wow, that sounds just like me!"
    • DeepSeek wrote tweets that were very similar in topic but sometimes felt a bit robotic or less "human."
    • GPT-4o was a solid all-rounder but didn't win the "most human" award.
    • Key Takeaway: Just because an AI uses the right words doesn't mean it captures the soul of the user.

Challenge 3: The "Psychic Profile" (User Attribute Inference)

The Goal: Can the AI look at a user's tweets and guess their job and hobbies without ever seeing their profile?

  • The Analogy: Imagine you are a psychic. You are handed a stack of 50 random postcards from a stranger. Based only on what they wrote, can you guess: "This person is a dentist who loves sci-fi movies"?
  • The Test: The AI had to guess the user's job (using official government job categories) and their interests (using a standard list of hobbies).
  • The Result: Gemini was the best psychic. It guessed jobs and hobbies with high accuracy. Llama, however, struggled significantly, often guessing wildly wrong jobs. This showed that even smart AIs can get confused if the clues are subtle.

The Big Picture: What Did We Learn?

  1. No Single Winner: There isn't one "best" AI for everything. GPT-4 is the best detective, Gemini is the best psychic, and Llama is the best actor (impersonator).
  2. The "New Data" Problem: Many AIs are great at repeating what they already know, but they struggle when faced with brand-new trends or posts written after their "learning cutoff."
  3. Human vs. Machine: Sometimes, the AI that looks the most "perfect" mathematically isn't the one that feels the most human. The best AI for social media needs to balance being smart with sounding natural.

In short: These AI models are powerful tools that can help us understand social media better, but they aren't perfect replacements for human intuition yet. They are like very talented interns who need a human boss to double-check their work, especially when it comes to the messy, ever-changing world of online conversation.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →