NVIDIA has introduced DSX MaxLPS, a suite of chip, thermal, system, and software technologies designed to maximize AI factory throughput within a fixed power budget. This system addresses the shift in focus from the number of GPUs in a data center to the AI output delivered per megawatt.
Key Points
- NVIDIA DSX MaxLPS integrates dynamic power allocation, performance-per-watt optimizations, and 45 C thermal site design.
- It aims to reclaim stranded capacity lost to static rack provisioning.
- Dynamic Power Software (DPS), currently in Developer Preview, enables fine-grained, real-time power redistribution across racks and GPUs.
- MaxLPS has been validated on NVIDIA Vera Rubin NVL72 and GB200 NVL72 systems.
- It can enable up to 40% more GPU capacity within the same facility envelope.
- Performance per watt improvements range from 1.3 to 1.5x, contingent on optimized infrastructure design, workload profiling, and early site-level engagement for 45 C liquid cooling.
- In a representative power-budget view, approximately 60% of delivered site power is allocated to compute for AI output.
Context
According to NVIDIA, AI factories are power-constrained industrial systems where the key metric for efficiency is application-level performance per watt for AI inference workloads. Traditional data center design often uses static rack provisioning, allocating maximum power per rack to meet worst-case peak demand. This approach can leave portions of the maximum power unused, as real workloads have varying power needs.
NVIDIA states that static provisioning treats each rack as an isolated power island, preventing unused power from being reallocated to other racks. DSX MaxLPS addresses this by optimizing three layers: land, utility power, and the physical shell holding infrastructure. DPS continuously monitors and optimizes power utilization, comparing allocated power against actual consumption and making unused headroom available to other units within the same managed group.
Why It Matters
This development is relevant for builders and operators of AI factories, as it offers a method to increase the efficiency and capacity of their compute infrastructure. By optimizing power utilization, organizations can achieve higher AI output from existing power budgets and facility footprints, potentially reducing the need for costly expansions or new facility construction.
What To Do
- Note that Dynamic Power Software (DPS) is currently in Developer Preview.
- Consider the implications of 45 C liquid cooling for infrastructure design.
- Evaluate the potential for 1.3-1.5x performance per watt improvements in your own optimized infrastructure designs.
- Watch for further guidance on integrating MaxLPS with NVIDIA Vera Rubin NVL72 and GB200 NVL72 systems.
