← AI PulseAug 20, 2026

Deep · news · Multi-source brief

xAI Introduces Speech to Text API with Real-time and Batch Transcription

xAI has launched a Speech to Text API that supports both batch file uploads and real-time WebSocket streaming for audio transcription.

By Illumora Editorial

Source · Aug 20, 2026, 3:13 PM · On Illumora · Aug 20, 2026, 4:13 PM

Media from the primary source — shown here so you can stay on Illumora.

Synthesized from multiple allowlisted primaries on the same event. Lanes →

Brief drafted by Illumora’s editorial model from the linked primary source. Ops desk reviews flagged pieces. How we write →

Read the source →xAI Docs — Speech to Text | SpaceXAI Docs
Save

xAI has released a new Speech to Text API, enabling developers to transcribe audio into text. This API supports two primary modes of operation: batch file uploads for transcribing audio files with a single API call and real-time audio streaming via WebSocket.

The Speech to Text API offers support for 12 audio formats, includes word-level timestamps, multichannel transcription, and text formatting. This capability is part of a broader suite of voice features from xAI, which also includes Speech to Speech and Text to Speech functionalities.

Key Points

  • The xAI Speech to Text API transcribes audio to text using batch file upload or real-time WebSocket streaming.
  • The API supports 12 audio formats for transcription.
  • Features include word-level timestamps, multichannel transcription, and text formatting.
  • The Speech to Speech API enables real-time voice conversations over WebSocket.
  • The Speech to Speech API is billed by minute of audio and a flat fee per text input message.
  • The Speech to Speech API supports function calling with web search, X search, collections, MCP, and custom functions.

Context

According to xAI documentation, the Speech to Text API is designed for converting spoken audio into written text. This complements the Speech to Speech API, which facilitates real-time voice conversations and supports various function calling capabilities, including web search and interactions with user-defined collections. The Speech to Speech API is priced starting at $0.05 per minute of audio or $3.00 per hour, with text input messages billed at $0.004 per message for conversation.item.create events. xAI's Collections service allows users to upload and search through documents, providing persistent storage and semantic search capabilities for integrating enterprise knowledge bases.

Why It Matters

For builders, the introduction of a dedicated Speech to Text API from xAI expands the options for integrating audio processing into applications. The availability of both batch and real-time transcription, alongside detailed features like word-level timestamps, offers flexibility for various use cases, from post-processing recorded media to enabling live voice interactions. Understanding the pricing structure and capabilities of related voice APIs, such as Speech to Speech, is crucial for designing cost-effective and feature-rich conversational AI systems.

What To Do

  • Review the xAI documentation for the Speech to Text API to understand specific implementation details and supported audio formats.
  • Compare the features of the Speech to Text API with the Speech to Speech API to determine the most suitable tool for real-time conversational applications.
  • Note the pricing model for the Speech to Speech API, which includes charges for both audio duration and text input messages, when planning for voice-enabled features.
  • Explore the Collections documentation to understand how custom knowledge bases can be integrated with xAI's voice APIs for enhanced functionality.

Keep Exploring

/atlas/**grok**-family