← Latest papers
💻 computer science

From Codebooks to VLMs: Evaluating Automated Visual Discourse Analysis for Climate Change on Social Media

This paper evaluates the effectiveness of various vision-language models for automated visual discourse analysis of climate change on social media, demonstrating that while Gemini-3.1-flash-lite achieves the highest instance-level accuracy, even moderate-performing models can reliably recover population-level trends, thereby offering a scalable alternative to manual annotation for large-scale research.

Original authors: Katharina Prasse, Steffen Jung, Isaac Bravo, Stefanie Walter, Patrick Knab, Christian Bartelt, Margret Keuper

Published 2026-04-24
📖 5 min read🧠 Deep dive

Original authors: Katharina Prasse, Steffen Jung, Isaac Bravo, Stefanie Walter, Patrick Knab, Christian Bartelt, Margret Keuper

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to understand a massive, chaotic conversation happening in a giant, noisy town square. This town square is social media, and the topic everyone is shouting about is climate change.

For years, researchers have been trying to listen to this crowd. But they've mostly only been reading the words people are shouting. They've ignored the pictures people are holding up—photos of polar bears, memes about heatwaves, infographics about rising sea levels, and protest signs. These pictures tell a huge part of the story, but there are millions of them. It's impossible for a team of human experts to look at every single one; they would need a lifetime and a fortune to do it.

This paper is like a team of scientists asking: "Can we teach a super-smart robot to look at these millions of pictures for us, and will it actually understand what it's seeing?"

Here is the breakdown of their adventure, using some simple analogies:

1. The Problem: The "Needle in a Haystack"

The researchers had two piles of pictures:

  • The "Gold Standard" Pile (ClimateCT): A small, carefully curated box of 1,000 pictures that human experts had already labeled perfectly. Think of this as a teacher's answer key.
  • The "Wild West" Pile (ClimateTV): A massive, messy mountain of over 1.2 million pictures scraped from X (Twitter). This is the real world—full of weird angles, blurry photos, memes, and confusing images.

2. The Tools: The Robot Brains

They didn't just use one robot; they brought in a whole team of Vision-Language Models (VLMs).

  • Think of these models as super-intelligent interns. They have read billions of books and seen billions of photos.
  • Some are "Open-Source" interns (free to use, like Qwen or Gemma).
  • Some are "Big Tech" interns (expensive, like Google's Gemini or OpenAI's GPT).
  • They also tested older, simpler robots (like CLIP) that are good at spotting objects but bad at understanding context.

3. The Test: The "Codebook" Challenge

The researchers gave these robots a Codebook. In the old days, humans used codebooks to manually sort pictures into categories like:

  • Animals: (Is there a polar bear? A pet? A farm animal?)
  • Consequences: (Is this a flood? A wildfire? Or just a sad face?)
  • Action: (Is someone protesting? Installing solar panels?)
  • Setting: (Is this in a city, a forest, or outer space?)
  • Type: (Is this a photo, a cartoon, or a meme?)

The robots had to look at the pictures and sort them into these buckets without being specifically trained on climate change beforehand. They had to figure it out just by reading the instructions (prompts).

4. The Big Discoveries (The "Aha!" Moments)

🏆 The Winner: The "Flash" Robot

They found that Gemini-3.1-flash-lite was the clear champion. It was like the smartest intern in the room. It didn't need to be the biggest or most expensive model to win; it just needed to be the right tool for the job. It beat the other robots, even the huge 30-billion-parameter ones, which sometimes got confused by their own complexity.

📉 The "Good Enough" Secret

Here is the most surprising finding: The robots didn't have to be perfect to be useful.

  • If you ask a robot to label one specific picture, it might get it wrong 30% of the time.
  • BUT, if you ask it to label 1 million pictures, the mistakes cancel each other out. The robot gets the overall trend right.
  • Analogy: Imagine trying to guess the average height of everyone in a stadium. If you ask a slightly clumsy robot to measure every person, it might get some wrong. But if you ask it to measure the whole crowd, the average it calculates will be incredibly accurate. For social scientists, knowing the trend (e.g., "more people are posting about floods than polar bears") is often more important than knowing if one specific photo was labeled correctly.

🧠 The "Over-Thinker" Trap

The researchers tried telling the robots to "think step-by-step" (Chain of Thought), like a human solving a math problem.

  • Result: This actually made the robots worse at this specific task.
  • Analogy: It's like asking a chef, "Before you chop this onion, please write a 5-paragraph essay on the history of onions." By the time they finish writing, they've forgotten how to chop. For visual tasks, sometimes you just need to "see" and "act," not over-analyze.

📝 The "Prompt" Matters

How you ask the robot matters.

  • For simple things (like "Is this a photo or a drawing?"), a short instruction works best.
  • For complex things (like "Is this a drought or a flood?"), the robot needed a longer, detailed explanation with examples.
  • Lesson: You can't use the same instruction for every job. You have to tailor the request.

5. The Conclusion: Why This Matters

This paper is a bridge. It tells social scientists, "You don't need to hire 1,000 humans to analyze climate images anymore."

  • The Old Way: Manual, slow, expensive, and limited to small samples.
  • The New Way: Automated, fast, cheap, and capable of analyzing the entire conversation.

While the robots aren't perfect (they still get confused between similar things, like "sea level rise" vs. "floods"), they are reliable enough to spot the big picture. They can tell us if the world is shifting from talking about "distant polar bears" to "local floods," which is exactly what policymakers and scientists need to know to fight climate change effectively.

In short: We finally have a robot assistant that can read the visual news of the world, and it's ready to help us understand how we are talking about our planet's future.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →