A large-scale corpus of religious radio broadcast transcripts from webstream recordings in the United States
This paper introduces a large-scale corpus of over 700,000 transcribed and diarized English-language religious radio broadcast segments from July 2025, collected from 785 webstreams representing more than 2,000 US stations, to facilitate research on religious media content, social-political discourse, and speech processing.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the airwaves as a giant, invisible library where thousands of radio stations are constantly shouting stories, songs, and sermons into the void. For decades, scientists who study how people communicate have been able to read the books in the "News" and "Talk Show" sections of this library, but the "Religious" section has been a mystery. Why? Because while millions of people tune in to faith-based radio every day, no one had ever built a giant, searchable index of what was actually being said. It's like trying to understand a whole city's culture by only looking at a few street signs, rather than reading the actual conversations happening in the cafes. To solve this, researchers needed a way to turn hours of fuzzy radio static into clear, organized text, and then use smart computer programs to figure out what those texts were actually about. This is the challenge of "content-level analysis": moving from just knowing a station exists to understanding exactly what it says, how it says it, and who is saying it.
Enter the Religious Radio Corpus, a massive new digital treasure map created by a team at the Pew Research Center. Think of this project as a super-powered, automated robot librarian that spent an entire month (July 2025) listening to 785 different live radio streams across the United States. Instead of just recording the audio, this robot team captured 15-minute chunks of sound every 15 minutes, 24 hours a day. By the time they were done, they had gathered over 700,000 recordings—more than 172,000 hours of raw audio! But the real magic happened next: the team used advanced AI to turn all that audio into text, identifying who was speaking and when. They even used a super-smart computer brain (a Large Language Model) to read those transcripts and tag them with labels like "Sermon," "Caller Interaction," or "Politics."
The result is a dataset containing over 60 million lines of text, organized into three neat layers: a list of the radio streams, a log of every recording, and the actual spoken words. This isn't just a pile of data; it's a tool that lets researchers finally answer big questions. For instance, they can now see how religious radio changes from the South to the West, or how often politicians get mentioned in sermons versus how often they talk about family advice. The team checked their work carefully, finding that their computer transcription was surprisingly accurate—making only about 5 mistakes for every 100 words, which is close to how well two different humans would agree if they both transcribed the same tape. While the computer sometimes struggled with noisy phone calls or background music, the overall picture is clear and reliable.
This paper doesn't just say "we have data"; it proves that we can now study religious broadcasting on a scale never before possible. It rules out the idea that we have to rely on small surveys or guess what's on the airwaves. Instead, it offers a complete, searchable record of what is actually being broadcast. The authors are confident that this data is solid enough to help scholars understand the "electronic church" as a national phenomenon, rather than just looking at one station at a time. Whether you are a student of religion, a political scientist, or a computer expert trying to teach machines how to understand human speech, this corpus opens the door to a whole new way of listening to the world.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.