A new data descriptor on arXiv cs.CL details a corpus of transcribed English-language religious radio broadcasts. This dataset was compiled from live webstreams over a one-month period in July 2025. The project addresses the previous lack of large-scale transcript data for content-level analysis of religious radio in the United States.
Key Points
- The corpus contains over 60 million diarized lines of speech.
- Recordings were captured from 785 distinct webstreams.
- These streams rebroadcast signals from more than two thousand AM and FM stations.
- Over 700,000 recordings were collected.
- Each recording consists of a fifteen-minute segment.
- An automated pipeline handled transcription and speaker diarization.
- A large language model was used to segment and label content by programming format and topic.
Context
According to the arXiv data descriptor, religious radio is a widespread form of mass communication in the United States that has been understudied due to the absence of large-scale transcript data. The corpus was created by recording fifteen-minute segments on a rolling schedule from 785 distinct streams during July 2025. The collected data was processed using an automated pipeline for transcription and speaker diarization, and a large language model for content segmentation and topic labeling. The corpus is structured as linked tables, including stream metadata, recording metadata, and transcript lines.
Why It Matters
This corpus provides a resource for researchers to conduct descriptive studies of religious broadcasting, analyze discussions of social and political issues within religious media, and advance speech-processing research in a domain that has been underrepresented.
What To Do
- Review the arXiv paper to understand the full methodology of data collection and processing.
- Examine the corpus structure, particularly the linked tables of metadata and transcript lines.
- Consider how the dataset's scale and diarized speech can support new analytical approaches to religious media content.
- Note the use of a large language model for topic and format labeling, and consider its implications for content analysis.
