Nvidia rubin’s AI edge holds—but hardware costs dominate debate
The gist
Nvidia’s Vera Rubin NVL72 still dominates AI benchmarks, but sky-high hardware costs are fueling a fierce debate over whether to buy now or wait for software to close the gap.
What to know
- MLPerf v6.1 offers the first neutral, peer-reviewed showdown between Rubin, GB300, and optimized legacy systems—showing up to 36% performance lifts from software tuning alone.
- Rubin NVL72 tops the charts with up to 3.7x faster throughput than GB300 on Qwen3-VL, but that lead shrinks to 1.7–1.8x in offline benchmark runs.
- Rubin’s performance comes at a price—about 70% of a $4M rack goes to GPUs, and each gigawatt of Rubin hardware costs an eye-watering $40B, far outpacing its rivals.
Benchmarking Goes Neutral
MLPerf v6.1’s peer-reviewed, vendor-agnostic results finally let buyers compare Rubin, GB300, and legacy systems on equal terms, exposing software-driven gains and real-world relevance for procurement decisions.
MLPerf v6.1 matters because it is the first time buyers can compare Rubin, GB300, and software-tuned installed systems on the same neutral benchmark tape, rather than stitching together vendor claims from different test conditions. Tech Times notes that the distinction matters because v6.1 delivers the first peer-reviewed, independently verified performance data for NVIDIA’s Vera Rubin NVL72 platform, with results submitted to and verified by MLCommons, a neutral consortium with 130 members and bylaws designed to prevent any single vendor from controlling the outcome; NVIDIA and Nebius both submitted Rubin results, strengthening cross-checkability.
Just as important, v6.1 changes what is being measured: the new End-to-End RAG and Edge Agentic Inference categories better match production workloads, including multi-component pipelines on constrained hardware, which makes cross-platform and cross-stack comparisons more procurement-relevant. Tech Times calls the most useful finding the proof that software optimization alone delivered 8 to 36 percent throughput gains in five months on already-purchased hardware, a gap now measurable from a neutral, consortium-governed source for the first time; that matters more because v6.1 drew submissions from 30 organizations and 486 datacenter and edge results, up from 24 organizations in v6.0.
Rubin’s Edge: Interactivity
Nvidia’s Rubin NVL72 dominates interactive, latency-sensitive AI tasks with up to 3.7x the throughput of GB300—while its lead shrinks for batch workloads, highlighting a nuanced, workload-dependent advantage.
The only provided evidence indicates Rubin’s NVL72 can outperform GB300’s NVL72 on MLPerf inference, establishing a hardware-led performance edge that can vary by workload. Nvidia’s own headline makes that plain — “NVIDIA’s Vera Rubin NVL72 Beats GB300 NVL72 with Up to 3.7× Faster MLPerf Inference” — and shattered.io says MLCommons’ September 16, 2026 MLPerf Inference v6.1 release included Nvidia’s peer-reviewed preview submission showing Vera Rubin NVL72 at up to 3.7x Qwen3-VL throughput and up to 2.5x DeepSeek-R1 throughput versus GB300 NVL72.
But the same results also show that Rubin’s lead is not a flat multiplier across every inference mode: shattered.io says the biggest gains appear in interactive, latency-sensitive scenarios, while “the offline gap (where a system can batch requests without a latency ceiling) is the narrowest of the six rows, closer to 1.7x-1.8x than the headline 3.7x figure.” That pattern fits the benchmark mix, because DeepSeek-R1 is described as “reasoning-focused” and “measures raw token-generation speed,” and the analysis adds that “Reasoning models like DeepSeek-R1 generate far more intermediate tokens per answer than earlier” models, making tight-latency serving the place where Rubin’s hardware edge shows up most clearly.
Hardware Drives Rubin’s Price
With 70% of a $4M rack cost tied to GPUs and a staggering $40 billion per gigawatt, Rubin’s performance comes from hardware heft—making supply chain and scaling the new battleground.
The reason Rubin cannot be reduced to a software story is visible in the rack economics: Yahoo Finance reported that in the NVIDIA VR200 NVL72 rack-scale AI supercomputer, “70% of the roughly $4 million bill of materials goes to one line item: the Rubin GPUs themselves,” citing an HSBC estimated bill of materials. Huang’s own cost-intensity framing points the same way, saying on a fiscal Q2 2027 call, “For Hopper, we were at about 18 billion per gigawatt. For Grace Blackwell, we're about 25 billion per gigawatt. And for Vera Rubin, it's about 40 billion per gigawatt.”
That hardware weight also shows up in how Rubin is supplied and scaled: AD HOC NEWS noted that Nvidia’s first official Rubin performance data was explicitly benchmarked against the preceding GB300 NVL72 platform, with MLPerf Inference v6.1 showing “throughput up to 3.7 times higher than the preceding GB300 NVL72 system” on one model, tying the gain to a new physical platform rather than tuning alone. The same report said UBS raised its 2027 AI GPU output forecast to 8.8 million units, including roughly 200,000 Rubin CPX chips, because of expected CoWoS expansion at TSMC, while Nvidia has guided for 70 percent fiscal 2028 revenue growth but says supply-chain delivery capacity remains the primary constraint.
Full-Stack Design as Differentiator
Nvidia claims Rubin’s co-designed hardware and software ecosystem delivers unmatched efficiency and throughput, arguing that only a holistic, rack-level approach justifies the steep investment over legacy upgrades.
Nvidia’s counterargument is that Rubin is not just a tuned successor but a full-stack system for agentic inference. SemiAnalysis says “Rubin is the first platform co-designed across six products for the agentic era: Rubin GPU, Vera CPU, NVLink 6 Switch, ConnectX-9, BlueField-4, and Spectrum-6,” and Nvidia adds rack-level claims that the new rack uses NVLink 6 and delivers “260 terabytes per second of bandwidth.” The point is that the gain comes from coordinated design, not software uplift alone, and Nvidia also frames the economics as performance per total cost of ownership, with “the Y-axis is total tokens per $1 TCO.”
That is why Nvidia argues for buying now even if older systems still improve through software tuning. At GTC 2026, Jensen said “VR NVL72 achieved 3x performance per MW compared to Blackwell on O(1-3 Trillion) parameter model around 200 TPS,” while SemiAnalysis says Rubin on prelease software is already seeing up to 7x better token throughput per megawatt and over 2x more profit per gigawatt than Blackwell. Nvidia also highlights “10 times higher inference per watt at onetenth the cost per token,” and the broader claim is that Rubin’s efficiency gains justify procurement before software on older platforms fully matures.


