← AI PulseJul 30, 2026

Deep · research · Single-source brief

Corpus of Religious Radio Broadcast Transcripts Released

A new dataset comprising over 60 million diarized lines of speech from English-language religious radio broadcasts recorded in July 2025 is now available.

By Illumora Editorial

Source · Jul 30, 2026, 4:00 AM · On Illumora · Jul 30, 2026, 4:02 AM

Media from the primary source — shown here so you can stay on Illumora.

Rewritten from one allowlisted primary — not independent enterprise reporting. Lanes →

Brief drafted by Illumora’s editorial model from the linked primary source. Ops desk reviews flagged pieces. How we write →

Read the source →arXiv cs.CL — A large-scale corpus of religious radio broadcast transcripts from webstream recordings in the United States
Save

A new data descriptor on arXiv cs.CL details a corpus of transcribed English-language religious radio broadcasts. This dataset was compiled from live webstreams over a one-month period in July 2025. The project addresses the previous lack of large-scale transcript data for content-level analysis of religious radio in the United States.

Key Points

  • The corpus contains over 60 million diarized lines of speech.
  • Recordings were captured from 785 distinct webstreams.
  • These streams rebroadcast signals from more than two thousand AM and FM stations.
  • Over 700,000 recordings were collected.
  • Each recording consists of a fifteen-minute segment.
  • An automated pipeline handled transcription and speaker diarization.
  • A large language model was used to segment and label content by programming format and topic.

Context

According to the arXiv data descriptor, religious radio is a widespread form of mass communication in the United States that has been understudied due to the absence of large-scale transcript data. The corpus was created by recording fifteen-minute segments on a rolling schedule from 785 distinct streams during July 2025. The collected data was processed using an automated pipeline for transcription and speaker diarization, and a large language model for content segmentation and topic labeling. The corpus is structured as linked tables, including stream metadata, recording metadata, and transcript lines.

Why It Matters

This corpus provides a resource for researchers to conduct descriptive studies of religious broadcasting, analyze discussions of social and political issues within religious media, and advance speech-processing research in a domain that has been underrepresented.

What To Do

  • Review the arXiv paper to understand the full methodology of data collection and processing.
  • Examine the corpus structure, particularly the linked tables of metadata and transcript lines.
  • Consider how the dataset's scale and diarized speech can support new analytical approaches to religious media content.
  • Note the use of a large language model for topic and format labeling, and consider its implications for content analysis.