← AI PulseAug 24, 2026

Wire · analysis · Single-source brief

NVIDIA Vera Rubin and Blackwell Set New Agentic AI Performance-per-Watt Standards

NVIDIA's upcoming Vera Rubin NVL72 achieved up to 30x higher AI-factory throughput per megawatt than the GB300 NVL72 on agentic AI workloads, according to preview results using the SemiAnalysis AgentX benchmark.

By Illumora Editorial

Source · Aug 24, 2026, 3:00 PM · On Illumora · Aug 24, 2026, 3:02 PM

Media from the primary source — shown here so you can stay on Illumora.

Rewritten from one allowlisted primary — not independent enterprise reporting. Lanes →

Brief drafted by Illumora’s editorial model from the linked primary source. Ops desk reviews flagged pieces. How we write →

Read the source →NVIDIA Developer Blog — NVIDIA Vera Rubin and Blackwell Set a New Standard for Agentic AI Performance per Watt | NVIDIA Technical Blog
Save

NVIDIA has released preview results for its Vera Rubin NVL72 and Blackwell GB300 NVL72 systems, highlighting their performance on agentic AI inference workloads. These results, measured using the SemiAnalysis AgentX benchmark, indicate significant efficiency gains for multi-step AI agent operations.

AI agents have expanded inference from single-turn interactions to multi-step workflows that involve reasoning, tool invocation, subagent coordination, and growing context over time. This shift is reflected in real-world usage, where average prompt tokens per request have grown approximately fourfold, and single agentic requests can consume 15 times the tokens of standard chat interactions.

Key Points

  • NVIDIA Vera Rubin NVL72 achieved up to 30x higher AI-factory throughput per megawatt than GB300 NVL72 on AgentX workloads.
  • GB300 NVL72 extended its multi-generational throughput-per-megawatt advantage over H200 NVL8, delivering up to 80x gains for large Mixture-of-Experts (MoE) models such as Kimi K3 2.8T.
  • AgentX is an open-source benchmark from SemiAnalysis's InferenceX suite, designed to evaluate agentic AI inference by replaying production-style coding agent sessions.
  • The benchmark captures long-context prefill, KV-cache reuse, tool-call gaps, and dynamic concurrency, which are critical aspects of agentic workloads.
  • Efficiency gains are attributed to system-level optimizations, including MoE serving runtimes (SGLang, TensorRT-LLM, vLLM), DeepGEMM-based kernels, mixed-precision formats (MXFP4, MXFP8), the NVIDIA Dynamo session-aware serving stack, and the high-bandwidth NVIDIA NVLink scale-up fabric connecting 72 GPUs.
  • AgentX measures serving performance across prerecorded Claude Code sessions with interleaved reasoning and tool use.

Context

According to NVIDIA, properly characterizing hardware performance for agentic AI presents new challenges due to the variable, stateful, and long nature of agentic sessions. These sessions chain model calls, tool use, and growing context rather than following fixed prompt-and-response patterns. The AgentX benchmark addresses this by measuring how efficiently accelerators serve the request patterns produced by real coding agents, focusing on throughput per provisioned megawatt while maintaining acceptable user experience.

Why It Matters

The reported efficiency gains for Vera Rubin and Blackwell systems indicate a potential for substantial increases in agentic inference capacity for AI factories. This could impact the cost and scalability of deploying complex AI agents that require multi-step reasoning and tool use, offering builders more efficient hardware options for demanding workloads.

What To Do

  • Note the performance metrics of Vera Rubin NVL72 and GB300 NVL72 when planning future AI infrastructure investments for agentic workloads.
  • Examine the SemiAnalysis AgentX benchmark to understand its methodology for evaluating agentic AI inference.
  • Consider the impact of system-level optimizations, such as MoE serving runtimes and mixed-precision formats, on overall efficiency for agentic AI applications.
  • Watch for further details on the NVIDIA Rubin GPU architecture and its design for agentic AI performance per megawatt.

Keep Exploring

/atlas/claude-family