← Latest papers
🔬 physics

TeraGram: A Structured Longitudinal Dataset of the Telegram Messenger

This paper introduces TeraGram, a massive longitudinal dataset comprising over 5.9 billion public Telegram messages from 2015 to 2025, which serves as a unique, algorithm-free resource for studying engagement patterns, network evolution, and community formation across diverse languages and communities.

Original authors: Anastasia Golovin, Sebastian B. Mohr, Arne I. Gottwald, Ulrik Hvid, Srushhti Trivedi, Joao Pinheiro Neto, Andreas C. Schneider, Viola Priesemann

Published 2026-05-18
📖 4 min read☕ Coffee break read

Original authors: Anastasia Golovin, Sebastian B. Mohr, Arne I. Gottwald, Ulrik Hvid, Srushhti Trivedi, Joao Pinheiro Neto, Andreas C. Schneider, Viola Priesemann

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the internet as a giant, noisy city. Most of the major squares in this city (like Facebook or Twitter/X) are run by invisible, super-smart bouncers. These bouncers decide which speeches you hear, which ads you see, and which conversations get amplified, all based on secret rules designed to keep you glued to your screen.

TeraGram is a massive new map of a different kind of city square: Telegram.

Here is the simple breakdown of what this paper is about, using everyday analogies:

1. The "No-Bouncer" City

Unlike other social media platforms, Telegram is like a town square where there are no bouncers deciding what you see. If someone posts a message, it appears in a simple, chronological line (like a real-time news ticker). There are no hidden algorithms pushing specific content to the top.

The researchers built TeraGram to study this unique environment. They wanted to see how people talk, share, and form communities when no one is secretly manipulating the conversation.

2. A Time Capsule of 5.9 Billion Messages

The team didn't just take a snapshot; they built a time machine.

  • The Size: They collected 5.9 billion messages. If you printed them out, they would fill a library the size of a small country.
  • The Time: The data stretches back to 2015 and goes up to 2025. It's a decade-long diary of public conversations.
  • The Scope: They didn't just look at one topic. They gathered data from 712,000 different channels and groups.

3. How They Collected It: The "Snowball" Method

Since Telegram doesn't give researchers a direct download button for everything, the team used a clever trick called "Snowball Crawling."

  • The Analogy: Imagine you want to map a forest. You start with one big tree (a popular channel). You look at the branches (messages) and see where they point to other trees (channels). You follow those links, find new trees, and follow their branches.
  • The Result: Like a rolling snowball getting bigger as it moves down a hill, their data collection grew exponentially, eventually capturing the most influential "hubs" of the Telegram network.

4. What's Inside the Box?

The dataset is like a giant, organized warehouse. It's not just a pile of text; it's structured so computers can easily read it. Inside, you can find:

  • The Messages: The actual words people wrote (though the full text is locked away for privacy, researchers can ask to see it).
  • The Reactions: How many people "liked" or reacted to a post.
  • The Polls: Millions of questions and answers people voted on.
  • The Connections: Who forwarded a message to whom, creating a map of how information travels.
  • The Languages: While English is there, the map is dominated by Russian (56%) and Farsi (15%), reflecting where Telegram is most popular as a daily tool. English is smaller (6%) but still huge in absolute numbers.

5. What Did They Find? (The "Vibe Check")

The researchers did a quick tour of the data to see what the "vibe" of these different language groups was:

  • Russian & Farsi: These groups talked about a mix of everything: politics, but also books, fashion, art, music, and daily life. It felt like a diverse, mainstream town square.
  • English: This group was different. While they also talked about sports and news, they had a much higher concentration of conspiracy theories and extremist topics (like "climate change hoaxes" or "antisemitic narratives").
  • The "Trust" Test: When they checked the websites people shared in English chats, they found that 60% of the links were to unreliable or fake news sources. Compare that to Twitter, where reliable sources are much more common. It's as if the English Telegram square is filled with people shouting rumors from unverified pamphlets, while the Russian/Farsi squares are more like standard newsstands.

6. Why This Matters

This dataset is a scientific tool. It allows researchers to study human behavior without the "noise" of secret algorithms.

  • Privacy First: To protect people, the researchers scrubbed phone numbers and hid names. They only kept public data.
  • Open for Science: They made the data available (in a format called Parquet) so other scientists can use it to study how communities form, how rumors spread, and how people behave when they aren't being nudged by an algorithm.

In short: TeraGram is a massive, decade-long, organized record of a social media platform that runs on "time" rather than "algorithms." It gives scientists a rare, clean look at how humans actually talk to each other when no one is trying to sell them something or manipulate their feed.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →