Every unused watt inside a modern data center represents infrastructure capacity left on the table. Traditional AI facilities are provisioned for the worst-case scenario when every single graphics processing unit reaches absolute peak power simultaneously.
This conservative, static approach creates a protective electrical buffer, but it leaves valuable hardware underused during standard training and inference operations. According to recent reports published by NVIDIA, a new software and architectural framework called DSX MaxLPS solves this inefficiency by implementing policy-governed power sharing across participating nodes.
Understanding how data centers manage electricity highlights the scale of this problem. An AI factory operates within a rigid hierarchy of electrical limits, ranging from utility connections and substations down to individual racks, server nodes, and chips.
Operators must maintain strict compliance within each managed boundary while simultaneously trying to maximize application throughput and keep latency low. Static power planning assumes that every node requires its maximum specified wattage at all times. In reality, AI workloads fluctuate significantly.
Training runs move constantly through compute phases, communication steps, synchronization hurdles, and checkpointing intervals. Inference tasks alternate between prefill stages, decode cycles, memory-bound routines, and network activity.
Because workloads vary, a massive gap opens up between reserved peak power and actual real-time consumption. Under standard static allocations, if one server node uses less than its allocated budget, that leftover power remains stranded. It cannot be dynamically lent to an adjacent node that might need extra electrical headroom to process heavier workloads.
The aggregate facility stays well below its permitted power ceiling, yet additional GPU capacity is prevented from coming online because of rigid per-node reservation rules.
Key Takeaways
- Traditional static power allocation leaves massive amounts of stranded power on the table in AI data centers.
- NVIDIA DSX MaxLPS utilizes policy-governed power sharing to dynamically redistribute power between nodes based on real-time telemetry.
- Commercial evaluations with Nscale in Iceland demonstrated a 49.2 percent increase in throughput per provisioned watt using identical power budgets.
- Operators can scale up their operational GPU capacity significantly without waiting years for new utility grid connections or electrical substations.
Inside the DSX MaxLPS Control Loop
To eliminate stranded power, NVIDIA DSX MaxLPS combines advanced chip telemetry, system architecture, thermal monitoring, and control software. Dynamic Power Software acts as the central orchestration layer, allowing operators to pool and redistribute power dynamically while strictly honoring facility safety boundaries. The architecture relies on five core technical components:
- Topology and resource groups: Operators map their physical infrastructure and group server nodes into managed clusters with a shared, aggregate power budget.
- Telemetry engines: The system continuously collects real-time power metrics from individual GPUs, nodes, racks, and entire groups at intervals fast enough to spot emerging power spikes.
- Policy engines: Rule sets defined by the facility operator establish strict node boundaries, group ceilings, allocation priorities, and automated responses to emergency or maintenance events.
- Allocation and control loops: When certain compute resources consume less than their allocated share, the software instantly adjusts GPU power limits elsewhere so other systems can utilize the headroom.
- Validation and enforcement: The system cross-references real-time consumption against approved aggregate budgets, automatically scaling back allocations if usage approaches a safety threshold.
This mechanism is entirely about coordinated allocation rather than artificially increasing the physical power supply of the data center. By intelligently shifting power allowances where they are needed most, operators can execute significantly more compute work underneath the exact same hard electrical limit.
Evaluating Performance Gains with Real-World Workloads
To measure the effectiveness of this framework, NVIDIA partnered with Nscale to evaluate the software inside a commercial data center located at the Verne campus in Keflavík, Iceland, which runs entirely on renewable energy. The evaluation tested Kimi K2.5 workloads running on NVIDIA GB300 NVL72 systems equipped with Blackwell Ultra GPUs. The test utilized FP4 precision, NVIDIA Dynamo, and TensorRT LLM with an 8K input sequence length and a 1K output sequence length. The workload combined high-throughput and low-latency inference instances simultaneously.
The baseline static configuration utilized 35 four-GPU nodes, totaling 140 GPUs running two high-throughput instances and one low-latency instance. By contrast, the DSX MaxLPS configuration expanded the fleet to 48 four-GPU nodes, totaling 192 GPUs, successfully adding a third high-throughput instance while keeping the low-latency instance active. Both setups operated under the exact same fixed power budget of 264.4 kilowatts.
| Metric | Static Baseline | DSX MaxLPS | Change |
|---|---|---|---|
| Managed GPUs | 140 | 192 | +37.1% |
| Aggregate Throughput | 1,084,503 tokens/s | 1,618,443 tokens/s | +49.2% |
| Mean GPU Power | 97.0 kW | 131.8 kW | +35.9% |
| Total Measured Power | 166.2 kW | 198.9 kW | +19.7% |
| Power-Budget Utilization | 62.9% | 75.2% | +12.3 percentage points |
| Throughput per Provisioned Watt | 4.10 tokens/s/W | 6.12 tokens/s/W | +49.2% |
The empirical results demonstrate a massive leap in efficiency. Because the provisioned power budget remained fixed at 264.4 kilowatts, dividing the normalized aggregate throughput by this static denominator revealed a 49.2 percent increase in throughput per provisioned watt, rising from 4.10 tokens per second per watt up to 6.12. Furthermore, individual per-instance performance metrics remained virtually identical, proving that expanding the overall GPU fleet did not degrade the output quality or speed of existing workloads.
Broader Implications for AI Infrastructure Management
As foundational models grow larger and inference demands soar, data center operators face severe constraints related to land availability, power generation, and electrical shell limits. Building new electrical substations can take years, making efficiency innovations critical for scaling artificial intelligence infrastructure.
Frameworks like NVIDIA DSX MaxLPS show that data centers do not necessarily have to wait for new grid connections to expand their operational capacity. By moving away from rigid static provisioning toward intelligent, policy-governed power sharing, operators can safely pack more hardware into existing footprints, maximize the utility of renewable energy sources, and drastically lower the capital expenditure required to scale enterprise AI factories.
Join our community by subscribing to our Weekly Newsletter to stay updated on the latest AI updates and technologies, including the tips and how-to guides.
(Also, follow us on Instagram (@tid_technology) for more updates in your feed and our WhatsApp Channel to get daily news straight to your Messaging App).
