← AI PulseAug 20, 2026

Wire · news · Single-source brief

Prompt Caching Billing in Grok-4.6

xAI's documentation indicates that cached token counts are visible in API responses for billing purposes within the Grok-4.6 Chat Completions API.

By Illumora Editorial

Source · Aug 20, 2026, 3:13 PM · On Illumora · Aug 20, 2026, 3:38 PM

Media from the primary source — shown here so you can stay on Illumora.

Rewritten from one allowlisted primary — not independent enterprise reporting. Lanes →

Brief drafted by Illumora’s editorial model from the linked primary source. Ops desk reviews flagged pieces. How we write →

Read the source →xAI Docs — Usage & Pricing | SpaceXAI Docs
Save

The Grok-4.6 Chat Completions API now provides visibility into cached token counts. This information is available in API responses, specifically within the usage.prompt_tokens_details.cached_tokens field.

Key Points

  • Cached tokens are reported in API responses.
  • The specific field for cached tokens is usage.prompt_tokens_details.cached_tokens.
  • This applies to the Grok-4.6 Chat Completions API.

Context

According to xAI's documentation, understanding prompt caching billing involves reviewing these cached token counts. The information is part of the broader usage and pricing details for their advanced API features.

Why It Matters

Builders can use this detailed token reporting to monitor and understand the cost implications of prompt caching when utilizing the Grok-4.6 Chat Completions API.

What To Do

  • Review API responses from the Grok-4.6 Chat Completions API for the usage.prompt_tokens_details.cached_tokens field.
  • Compare reported cached token counts against billing statements.
  • Note how prompt caching impacts overall token usage and costs.

Keep Exploring

/atlas/**grok**-family