Nvidia DSX sets new token-per-watt record in live test
The gist
Nvidia's DSX platform just smashed the tokens-per-watt record by turning power into a schedulable resourcedelivering up to 49% more AI throughput without drawing a single extra watt.
What to know
- A live test in Santa Clara showed DSX MaxLPS slashed power draw from 4 MW to 3 MW while keeping top-priority AI inference humming.
- In a fixed 264.4 kW cluster test, DSX MaxLPS ran 192 GPUs versus a 140-GPU baseline, boosting aggregate throughput by 49.2% to 1.6 million tokens per second and 6.12 tokens/s/W.
- Nvidia is rolling out DSX as a full-stack certification frameworkcombining 800 VDC power, 3 MW liquid cooling, modular blocks, and DSX Ready standards to turn contracted megawatts into up to 1.4x more tokens.
Power Scheduling, Not Just Saving
DSX transforms power into a dynamic, software-controlled asset, letting facilities reallocate real-time energy headroom to priority AI workloads and boost compute density without raising utility caps.
DSX works by treating power as a schedulable resource rather than a fixed per-server entitlement, using software to move available headroom to the workloads that matter most while staying inside the site’s approved envelope. Nvidia’s September descriptions say DSX MaxLPS monitors GPU and rack-level power use and shifts headroom across systems by workload type, while Emerald AI’s Conductor showed the grid-aware layer above it in Santa Clara, where Silicon Valley Power sent a signal to an AI facility to cut electricity use during high demand and the software reduced the site’s draw from four megawatts to three while keeping higher-priority inference running.
That mechanism matters because it turns stranded electrical capacity into usable compute density without asking utilities for a bigger cap. Nvidia’s case study says Lambda, in a five-rack, 19-node cluster, ran 19 nodes within the same facility power budget as 16 nodes at full power and still recorded a 24% increase in cluster-wide token throughput, while Huawei described the same operating logic at the power-system layer, saying it improves adaptability to weak power grids and large load fluctuations through grid-forming control and energy storage synergy, including its 1000 V FusionSolar inverter line and 1000 V LUTERRA platform.
Throughput Gains Without Extra Watts
DSX MaxLPS enables clusters to run up to 40% more GPUs and achieve double-digit throughput gains—all while strictly holding the power budget steady.
The clearest evidence that DSX is more than a theoretical optimization is that recent validations measured higher output while holding power constraints in place. Wccftech summarized one real-cluster result in its headline — “Nvidia DSX MaxLPS Lifts Lambda’s Cluster to 5 M Tokens/sec, Cutting Power Use 23%” — while NVIDIA Developer detailed a joint NVIDIA-Nscale test running Kimi K2.5 on NVIDIA GB300 NVL72 systems at Verne’s Keflavík campus under a fixed 264.4 kW provisioned budget, where DSX MaxLPS increased normalized aggregate throughput by 49.2%. NVIDIA says DSX MaxLPS uses “policy-governed power sharing to dynamically allocate power across participating resources,” enabling customers to “deploy up to 40% more GPUs within the same approved power budget.”
What makes that Nscale result persuasive is that the gains were achieved inside the same approved envelope rather than by simply drawing more electricity. NVIDIA Developer said DSX MaxLPS uses “policy-governed power sharing to dynamically allocate power across participating resources,” and in the evaluation it ran 192 GPUs versus a 140-GPU static baseline within the same 264.4 kW budget; Quantum Zeitgeist reported throughput rose to 1,618,443 tokens per second from 1,084,503, power-budget use improved from 62.9% to 75.2%, and throughput per provisioned watt climbed from 4.10 to 6.12 tokens/s/W. That 192-versus-140 comparison also illustrates NVIDIA’s claim that customers can “deploy up to 40% more GPUs within the same approved power budget.”
Hardware Moves as Fast as Software
DSX’s modular stack and rapid partner rollouts are reshaping data center design, as AI hardware generations drive infrastructure to scale up power and cooling in lockstep.
DSX is emerging as a build pattern for AI factories because the physical plant now changes almost as fast as the compute roadmap. Futurise argues that data-center planning “takes time,” even as operators have “moved into the Blackwell generation… more like 150 kilowatts per single rack” and then toward “Vera Rubin… more like 4 or 500 kilowatts per rack,” forcing redesigns; they add that Rubin can “run on warm water at 45 degrees,” eliminating chillers in some cases, and warn, “we’re saying 1 megawatt now, but it could be 2 megawatts tomorrow in that same rack.”
That is why NVIDIA’s September partner rollout reads less like software distribution than infrastructure assembly. In PR Newswire’s Delta announcement, NVIDIA says “Delta's expertise in 800 VDC power delivery and liquid cooling will help customers deploy AI factories based on NVIDIA DSX faster,” while Delta’s prefabricated modular design “integrates 800 VDC In-Row Power and 3 MW of liquid-cooling capacity into infrastructure blocks assembled and tested in the factory,” alongside “energy storage systems (ESS), solid-state transformers (SST) and medium-voltage infrastructure” at the facility layer and “800 VDC power delivery with row-to-chip liquid cooling” across the IT stack.
Certification Becomes the New Benchmark
Nvidia’s DSX Ready program shifts the industry focus to validating end-to-end power-to-token efficiency, making facility certification—not just hardware upgrades—the key to AI factory performance.
NVIDIA is clearly trying to turn DSX from a tuning technique into a qualification regime. The Tech Buzz reported that the new DSX Ready program creates a compatibility checklist for third-party power and cooling products and is meant to reduce deployment risk, while Techzine Global shows how NVIDIA is bundling that logic into “DSX MaxLPS, part of its DSX platform for AI factories,” after describing rack-level power optimization with 224 kilowatts of peak draw running at 164 kilowatts in its most efficient mode, framing infrastructure as something that must be validated end to end rather than assembled through trial and error once projects reach production scale.
That shifts the market contest toward certifying whether facilities can reliably convert contracted megawatts into token output, not merely install more accelerators. Techzine Global says NVIDIA now breaks losses into “about 20 percent” in the building, “another 10 percent” in poorly configured racks, and “10 percent to failures and restarts,” while also warning, “You sign a contract for 100 megawatts, but you never actually get that”; at the “AI Infra Summit… in mid-September,” NVIDIA claimed “MaxLPS delivers up to 1.4 times more tokens per megawatt,” and in a technical explanation cited “up to 40 percent more GPU capacity within the same power,” so against that backdrop, CIO’s claim that liquid cooling can push PUE to 1.15 versus 1.6 for air, and OCP work on common 800V DC interfaces and safety standards, look like the emerging certification battleground.

