← AI PulseJul 23, 2026

Deep · news · Multi-source brief

xAI Introduces Speech to Text API with Batch and Real-time Transcription

xAI has launched a Speech to Text API that supports both batch file uploads and real-time WebSocket streaming for audio transcription.

By Illumora Editorial · Jul 23, 2026

Synthesized from multiple allowlisted primaries on the same event. Lanes →

Brief drafted by Illumora’s editorial model from the linked primary source. Ops desk reviews flagged pieces. How we write →

Read the source →xAI Docs — Speech to Text | SpaceXAI Docs
Save

xAI has released a new Speech to Text API, enabling developers to transcribe audio into text. This API offers two primary methods for transcription: batch processing of audio files via a REST endpoint and real-time, low-latency transcription using a streaming WebSocket endpoint. The service supports 12 audio formats, including WAV, MP3, WebM, OGG, and M4A, and provides features such as word-level timestamps, multichannel transcription, and text formatting.

Key Points

  • The Speech to Text API allows transcription of audio files with a single API call or real-time streaming over WebSocket.
  • The API supports 12 audio formats, including WAV, MP3, WebM, OGG, and M4A.
  • Features include word-level timestamps, multichannel transcription, and text formatting.
  • Pricing for REST-based batch transcription is $0.10 per hour, while streaming transcription costs $0.20 per hour.
  • Rate limits are 10 requests per second for REST and 10 concurrent sessions per team for streaming.
  • The API includes keyterm prompting for domain-specific vocabulary and Smart Turn end-of-turn detection for streaming.
  • The service is available in the us-east-1 region, with pricing details last updated on July 21, 2026.

Context

xAI's Speech to Text API is part of a broader suite of model capabilities that includes text generation, image generation, and video generation, according to xAI documentation. The API is designed to transcribe speech in multiple languages, with the language parameter enabling specific formatting for numbers, currencies, and units. The model, referred to as grok-4.5, also supports multimodal inputs, allowing images to be considered in generating responses, and offers video generation from text prompts.

Why It Matters

This API provides developers with tools for integrating audio transcription into applications, offering both cost-effective batch processing and low-latency real-time capabilities. The inclusion of features like keyterm prompting and Smart Turn detection can enhance the accuracy and utility of transcriptions for specific use cases.

What To Do

  • Open the xAI documentation for the Speech to Text API to review the full list of supported audio formats and language formatting options.
  • Compare the pricing structure for REST and streaming transcription to determine the most suitable option for specific project needs.
  • Test the API's keyterm prompting feature with domain-specific vocabulary to evaluate its impact on transcription accuracy.
  • Note the rate limits for both REST and streaming endpoints to plan for application scaling.

Keep Exploring

/atlas/grok-family