← Latest papers
💬 NLP

Beyond speculation: Measuring the growing presence of LLM-generated texts in multilingual disinformation

This study provides the first empirical evidence of a growing presence of LLM-generated content in real-world multilingual disinformation datasets following ChatGPT's release, documenting specific patterns across languages, platforms, and time periods to bridge the scholarly debate on the actual scale of this threat.

Original authors: Dominik Macko, Aashish Anantha Ramakrishnan, Jason Samuel Lucas, Robert Moro, Ivan Srba, Adaku Uchendu, Dongwon Lee

Published 2026-02-05
📖 4 min read☕ Coffee break read

Original authors: Dominik Macko, Aashish Anantha Ramakrishnan, Jason Samuel Lucas, Robert Moro, Ivan Srba, Adaku Uchendu, Dongwon Lee

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Picture: Is the "Robot Pen" Writing Lies?

Imagine a world where anyone can pick up a super-smart, instant-writing pen (a Large Language Model, or LLM) that can write stories, news, and posts in almost any language. Since the release of tools like ChatGPT, people have been worried: Is this pen being used to write fake news and lies (disinformation)?

For a long time, experts were arguing in circles. Some said, "Don't worry, the internet is too messy for robots to take over." Others said, "Be careful, there are hidden corners where robots are already causing trouble."

This paper is the first to actually count how many of these "robot-written" lies are out there in the real world. They didn't just guess; they went into the digital archives of fact-checked news and social media to find the evidence.

The Detective Tools: How They Found the Robots

The researchers built two "digital lie detectors." Think of these as two different security guards at a club:

  1. Guard A (Gemma_GenAI): Trained to spot robot writing based on general patterns.
  2. Guard B (Gemma_MultiDomain): Trained specifically on social media and news to spot robot writing.

They tested these guards on known lists of "human vs. robot" writing to make sure they were good at their jobs. They found that when both guards agreed a text was written by a robot, they were right about 93% to 99% of the time.

Once they trusted their guards, they sent them to patrol four different "neighborhoods" (datasets) of real-world content, including fact-checked claims, election tweets, and war-related news.

What They Found: The Robot Invasion is Real (But Uneven)

The study found that robot-written disinformation is definitely present, but it's not a flood; it's more like a slow leak that is getting bigger.

1. The "Before and After" ChatGPT Effect
Imagine a timeline. Before late 2022 (when ChatGPT became popular), the amount of robot-written lies was very low.

  • The Analogy: Think of it like a quiet pond. In 2021, only a few ripples (about 1%) were made by robots.
  • The Change: By 2023, after ChatGPT arrived, the ripples grew. The study found that in 2023, between 1.5% and 15% of the fact-checked lies in their dataset were likely written or edited by robots. The most confident estimate puts it at around 1.85%.
  • The Trend: The number of robot-written lies is climbing every year.

2. The "Targeted" Attack
The robots aren't writing lies in every language or on every platform equally. It's like a burglar who only targets specific houses.

  • Languages: While English and Spanish have the most robot lies simply because there are more people speaking them, the percentage of lies is highest in Polish (4.7%) and French (4.2%).
  • Platforms: The robots are most active on Instagram (1.5%) and Facebook (1.3%), and less active on Twitter/X (0.64%).
  • Specific Events: During the 2024 US election, the amount of robot-written content in tweets doubled from January to November, peaking right before the election.

3. The "Fake News" Datasets
The researchers also looked at datasets specifically labeled as "Fake News."

  • They found that 2.6% to 3.3% of the articles labeled as "fake" were actually written by robots.
  • In a dataset about the war in Gaza, they found that over 10% of the French-language posts were robot-generated, compared to much lower numbers in other languages.

Why This Matters

The paper concludes that the fear of robots writing lies is not just speculation; it is happening right now.

  • The "Bad" News: Robots are being used to create disinformation, and they are getting better at it. This is happening in specific languages and during specific times (like elections), which suggests someone is using them strategically.
  • The "Good" News: The amount of robot lies is still a small fraction of the total internet (around 2% in 2024). However, because the internet is so huge, even 2% is a lot of fake content.
  • The Warning: As robots get smarter, they might start "polluting" the water. If future AI models are trained on data that already contains robot-written lies, those new models might start believing the lies are true and spreading them even faster.

The Bottom Line

The authors are saying: "Stop guessing and start measuring." We now have proof that robots are writing lies in the real world, especially in Polish, French, and during election seasons. We need better tools to catch them and better systems to tell us when we are reading a robot's story so we don't get fooled.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →