NVIDIA and Nscale conducted a joint evaluation of DSX MaxLPS, a policy-driven dynamic power allocation system, at Nscale’s Verne campus data center in Keflavík, Iceland, which runs entirely on renewable energy. The test ran Kimi K2.5 workloads on NVIDIA Blackwell Ultra GPUs in NVL72 systems to see how much GPU capacity could be deployed inside the same approved provisioned power budget of 264.4 kW. Compared with a static baseline that managed 140 GPUs, DSX MaxLPS allowed 192 managed GPUs (+37.1%) and raised aggregate throughput from 1,084,503 to 1,618,443 tokens/s (+49.2%).
Why static provisioning leaves capacity unused
AI facilities are constrained by a hierarchy of electrical limits: utility connection, substations, distribution gear, racks, nodes and GPUs. Static provisioning typically reserves enough power for every node to hit its peak simultaneously. Because AI workloads are variable — training phases, inference prefill/decode, memory- and network-bound periods, and idle intervals — reserved peak power often exceeds actual consumption on many nodes at any given time. With per-node reservations, unused capacity in one reservation can’t be reallocated to others, so the site can operate below its aggregate limit while useful GPU capacity remains offline.
What DSX MaxLPS does
DSX MaxLPS monitors real consumption and dynamically reallocates available power across participating resources according to operator policies, preserving the aggregate budget and enforced boundaries. This is coordinated allocation within existing site limits, not an increase to the site’s power supply.
Elements of the control loop
- Topology and resource groups: operators map infrastructure and group nodes with an aggregate power budget.
- Telemetry: GPU, node, rack and group-level power telemetry are collected at intervals sufficient to detect headroom and rising power events.
- Policy: operator-defined rules set node and group limits, allocation priorities, reserve requirements and responses to maintenance or emergencies.
- Allocation and control: when some resources draw less than their allocation, the software adjusts GPU power limits so other resources can use the freed capacity.
- Validation and enforcement: measured power is compared to the approved group budget and allocations are adjusted when consumption approaches a limit.
How the method was evaluated
Nscale deployed MaxLPS software in its data center and collected telemetry while NVIDIA ran workloads to measure control behavior and workload trade-offs. Test details:
- GPUs: NVIDIA Blackwell Ultra
- Workload: Kimi K2.5 (FP4), NVIDIA Dynamo, NVIDIA TensorRT LLM
- Sequence lengths: 8K input, 1K output
- Workload mix: combination of high-throughput and low-latency inference instances
Static baseline: 35 four-GPU nodes (140 GPUs) running two 52-GPU high-throughput instances and one 36-GPU low-latency instance. DSX MaxLPS configuration: 48 four-GPU nodes (192 GPUs) that retained the 36-GPU low-latency instance and added a third 52-GPU high-throughput instance.
Jobs ran across four racks, with each distributed workload confined to a single rack in both configurations to control for cross-rack differences. The team measured normalized aggregate and per-instance throughput, time to first token, end-to-end latency, interactivity, and GPU/CPU/rack power telemetry to ensure throughput gains did not mask latency, stability, or compliance regressions.
Measured results (comparison)
- Provisioned power budget: 264.4 kW (unchanged)
- Managed GPUs: 140 → 192 (+37.1%)
- Aggregate throughput: 1,084,503 → 1,618,443 tokens/s (+49.2%)
- High-throughput output per instance: 59,153 → 59,220 tokens/s (+0.1%)
- Low-latency output per instance: 2,265 → 2,265 tokens/s (0%)
- Mean GPU power: 97.0 kW → 131.8 kW (+35.9%)
- Total measured power: 166.2 kW → 198.9 kW (+19.7%)
- Power-budget utilization: 62.9% → 75.2% (+12.3 percentage points)
- Throughput per provisioned watt: 4.10 tokens/s/W → 6.12 tokens/s/W (+49.2%)
Because the provisioned-power denominator (264.4 kW) was constant, the 49.2% increase in throughput per provisioned watt matches the aggregate-throughput gain. Per-instance output remained effectively unchanged at the reported precision, indicating the larger managed fleet increased aggregate throughput without materially reducing existing instances’ throughput.
On latency, median and P75 values stayed within 5% of baseline, but P99 time to first token increased by 17% from the baseline 15.7 seconds — highlighting the importance of evaluating tail latency alongside capacity and throughput.
Trade-offs revealed
- Workload mix matters: DSX MaxLPS reclaims unused headroom, so the amount of available reallocatable capacity depends on whether workloads peak at different times (complementary profiles) or simultaneously.
- Tail latency can worsen: stable median latency does not guarantee unchanged P99 behavior; production acceptance criteria should include tail latency requirements.
- Telemetry quality is critical: missing, delayed or mis-mapped measurements can undermine fleet-level decisions. Site-level telemetry was used in this evaluation to verify rack-level measurements; operators must confirm telemetry and topology before deployment.
Validation recommendations for operators
Follow a staged validation with explicit boundaries and acceptance criteria:
- Define the managed boundary: map utility, distribution, rack, node and GPU topology; set resource-group budget, reserve and escalation behavior; identify the enforceable measurement.
- Establish a representative baseline: run the production-like workload mix under static provisioning and measure performance and power long enough to capture variability and repeatability.
- Introduce policies conservatively: start with limits near the validated baseline and confirm telemetry, topology, control response and budget compliance before adding nodes.
- Add capacity and test each stage: incrementally increase population and compare aggregate and per-instance performance; verify behavior during peak demand, operating transitions, telemetry failure and reduced power availability.
- Set production operating limits: approve a configuration only when it meets throughput and latency objectives, stays within the managed budget, preserves required reserves and behaves predictably under faults and transitions.
Operational planning
Dynamic allocation is an operational capability, but the site must be planned to host the enabled capacity. Electrical distribution, cooling, network fabric, floor space and rack positions should be sized for the validated lifecycle target even if fewer racks are populated initially.
Note: this evaluation reports measurements on GB300 NVL72 hardware. For future NVIDIA Vera Rubin NVL72 deployments, DSX MaxLPS can be combined with performance-per-watt techniques and infrastructure designed for 45°C liquid-cooling inlet operation; any Vera Rubin capacity projections should be considered separate from this GB300 NVL72 evaluation.
Conclusion
DSX MaxLPS provides a framework for turning workload variability into managed capacity inside a fixed power budget. Operators must define boundaries, measure representative behavior, tune policies to service objectives and prove compliance under normal and adverse conditions. Read the NVIDIA DSX MaxLPS and NVIDIA Dynamic Power Software documentation, and follow the validation sequence described here to establish production operating limits for your AI facility.



