← AI PulseJul 23, 2026

Wire · news · Multi-source brief

xAI Introduces Speech to Text API with Batch and Streaming Options

xAI has launched a Speech to Text API that supports both file-based batch transcription and real-time, low-latency streaming transcription.

By Illumora Editorial · Jul 23, 2026

Synthesized from multiple allowlisted primaries on the same event. Lanes →

Brief drafted by Illumora’s editorial model from the linked primary source. Ops desk reviews flagged pieces. How we write →

Read the source →xAI Docs — Speech to Text | SpaceXAI Docs
Save

xAI has released a Speech to Text API, which transcribes audio into text. This API offers a REST endpoint for batch transcription of files and a streaming endpoint for real-time, low-latency transcription. The service was last updated on July 21, 2026, according to xAI documentation.

Key Points

  • The Speech to Text API supports transcription of multiple audio formats, including WAV, MP3, WebM, OGG, and M4A.
  • Pricing for the REST endpoint is $0.10 per hour, while the streaming endpoint costs $0.20 per hour.
  • The API operates in the us-east-1 region.
  • Capabilities include keyterm prompting for domain-specific vocabulary and Smart Turn end-of-turn detection for streaming.
  • The streaming service provides real-time interim results.
  • The grok-4.5 model is available for use with the Speech to Text API.

Context

According to xAI, the Speech to Text API is part of its broader voice capabilities. The company also offers reasoning models, such as grok-4.5, which can process information step-by-step and support a reasoning_effort parameter. Streaming outputs are supported across various models with text output capabilities, including chat and image understanding, and utilize Server-Sent Events (SSE) for real-time feedback.

Why It Matters

The introduction of xAI's Speech to Text API provides developers with tools for integrating audio transcription into applications, offering both batch processing for large files and real-time capabilities for interactive experiences. The pricing structure and support for multiple audio formats and languages can influence development choices for voice-enabled features.

What To Do

  • Review the xAI Speech to Text Guide for getting started with the API.
  • Compare the pricing details for REST and streaming transcription to align with project budgets.
  • Test the API's capabilities in the Playground to evaluate performance with specific audio formats and languages.
  • Note the rate limits for both REST (10 requests per second) and streaming (10 requests per second, 100 concurrent sessions per team).

Keep Exploring

/atlas/grok-family /techniques/multimodal-grounding