NeoClouds level up: Nvidia, Meta, and Microsoft drive AI infrastructure from GPU brawn to software brains

Tech Investments

The gist

AI infrastructure is leveling up fast as Nvidia, Meta, and Microsoft pivot from GPU stockpiles to smarter, software-driven cloud ecosystems that squeeze more value out of every watt.

What to know

  • Nvidia’s Ruben CPX, launching end of 2026, headlines a new breed of AI hardware built for seamless NeoCloud integration—exemplified by Nebius’s $19.4B deal with Microsoft.
  • NeoCloud providers like Nebius are ditching raw GPU sales for value-added AI platforms, racking up mega-contracts—like a $27B compute pact with Meta—to deliver exclusive, monetizable GPU clusters.
  • Operational efficiency is the new arms race: with inference workloads surging 79% CAGR, innovations like HPE’s Private Cloud AI and vCluster’s virtualization tech are redefining how AI compute is managed, secured, and scaled.

Nvidia’s NeoCloud Power Play

Nvidia’s Ruben CPX launch and deepening NeoCloud alliances are cementing its dominance by focusing on integrated hardware-software stacks and operationally efficient infrastructure, not just chip specs.

Nvidia’s upcoming Ruben CPX product, set to launch at the end of 2026, exemplifies the company’s strategic pivot towards enhancing AI task efficiency through hardware that integrates seamlessly with existing server architectures or operates as discrete units. This innovation aligns with the rise of NeoCloud providers like Nebius, which emerged from Yandex’s post-invasion restructuring and quickly gained credibility by securing a massive $19.4 billion infrastructure deal with Microsoft, underscoring the growing importance of these agile cloud players in expanding GPU capacity without the risks of traditional data center builds.

By early 2026, Nvidia was doubling down on partnerships with NeoCloud providers such as Nebius and CoreWeave to distribute its next-generation Ruben chips, capitalizing on surging demand for AI inference and training. These NeoClouds predominantly rely on Nvidia’s GPUs, constrained by the dominant CUDA ecosystem and Nvidia’s TSMC production advantage, which limits diversification but reinforces Nvidia’s hardware hegemony within the AI cloud space.

Rather than competing on hardware variety, NeoCloud providers differentiate themselves through sophisticated software layers and innovative data center infrastructure designs focused on operational efficiency, such as optimized cooling and networking. This strategic emphasis reflects a maturation of the AI cloud ecosystem where performance gains increasingly derive from how Nvidia’s hardware is deployed and managed, rather than from alternative chip architectures.

Sources
Bloomberg TechThe Information

Mega-Deals Reshape AI Cloud

Nebius and peers are racing to convert massive hyperscaler contracts into revenue by shifting from raw GPU sales to differentiated AI platforms—while facing huge capital and execution risks.

NeoCloud providers like Nebius have secured unprecedented multi-billion dollar long-term contracts with hyperscalers such as Meta and Microsoft, reflecting overwhelming demand for AI compute capacity. For instance, Nebius’ deals include a $3 billion five-year contract with Meta and a $17.4-$19.4 billion agreement with Microsoft, underscoring the strategic importance of large-scale partnerships. However, this surge in demand has exposed capacity as the primary bottleneck to revenue growth, prompting Nebius to aggressively expand its infrastructure to 2.5 gigawatts contracted and up to 1 gigawatt of connected data centers by the end of 2026, highlighting the capital-intensive nature of scaling AI infrastructure.

While Nebius and peers initially focused on selling raw GPU capacity at discounted rates to attract hyperscaler clients, they are now pivoting towards offering differentiated, value-added software and services to broaden their market appeal and improve margins. Nebius’ launch of enterprise-ready platforms like Aether and the Nebius Token Factory exemplifies this strategic shift from commoditized hardware sales to integrated AI cloud solutions, aiming to diversify their customer base beyond hyperscalers and increase revenue per gigawatt beyond the current $9-10 billion benchmark, which still lags behind major cloud providers.

Despite holding a robust $20 billion-plus backlog from hyperscalers, Nebius faces significant financial and operational challenges, including a $16-20 billion capital expenditure plan for 2026 against only $3.7 billion in cash, exposing substantial funding and execution risks. NVIDIA’s strategic $2 billion investment, increasing its stake to 8.3%, not only alleviates immediate funding pressures but also deepens the partnership through co-engineering and preferential access to next-generation Rubin platform GPUs. This alliance signals a market evolution from raw capacity sales to integrated AI infrastructure solutions, with Nebius under intense scrutiny to convert backlog into revenue and achieve operational targets like 800MW-1GW online capacity and a 40% adjusted EBITDA margin.

Nebius’ landmark $27 billion AI compute partnership with Meta, structured as a $12 billion dedicated capacity commitment plus a $15 billion optional capacity tranche, exemplifies the shift towards delivering exclusive, low-latency, and secure GPU clusters tailored for hyperscaler needs. This deal not only provides Meta with physically isolated GPUs ensuring performance and security but also enables Nebius to monetize optional capacity at higher retail margins by third-party customers, with Meta acting as a guaranteed backstop buyer. Scheduled to begin delivery in early 2027, this partnership marks one of the largest commercial deployments of NVIDIA’s Vera Rubin architecture and reflects Nebius’s dual market positioning: anchoring hyperscaler demand while building a competitive AI cloud platform to rival AWS, Oracle, Azure, and Google Cloud.

Sources
Tech InvestmentsFOMO研究院電子報Global Equity Briefing

Efficiency Overtakes GPU Counts

AI infrastructure leaders are prioritizing software-driven orchestration and real-world utilization metrics like tokens per watt, marking a decisive shift away from simply amassing GPU hardware.

By mid-2026, the AI infrastructure narrative has decisively shifted from a singular obsession with GPU quantity to a nuanced focus on operational efficiency, driven by software innovations such as workload orchestration, virtualization, and advanced management stacks. Industry leaders like Virtuozzo CEO Kurt Daniel emphasize that organizations now prioritize extracting maximum usable AI output per watt, reflecting growing concerns over power constraints and the diminishing returns of simply adding more GPUs. This evolution acknowledges that raw FLOPS no longer suffice as the benchmark; instead, metrics like tokens per watt and utilization rates have become critical to sustainable AI deployment.

The rise of inference workloads, growing at a staggering 79% CAGR and projected to comprise two-thirds of AI compute by 2026, has catalyzed a transformation in infrastructure design and operational priorities. Companies like QumulusAI, which inked $124 million deals tied to Nvidia Blackwell deployments, highlight how inference economics hinge on efficient GPU utilization rather than sheer capacity. Hyperbolic CEO Jasper Zhang underscores that idle GPU capacity is the market’s costliest problem, prompting a shift toward balancing performance, reliability, and cost to support continuous, production-grade AI services distinct from finite training jobs.

Software-driven orchestration has emerged as the linchpin for operational efficiency, enabling dynamic scaling, intelligent request routing, and workload-specific infrastructure tuning that hardware alone cannot achieve. Lambda’s focus on cloud software facilitating retail pricing and maximizing GPU utilization exemplifies this trend, as does Deloitte’s projection that 80-85% of inference workloads will require global distribution within two years. This shift demands sophisticated service level objectives—covering latency, throughput, and canary deployments—and the segregation of inference compute from training clusters to handle heterogeneous, unpredictable demand patterns effectively.

Looking ahead, AI infrastructure platforms are evolving into more automated, secure, and deeply integrated ecosystems that address not only efficiency but also compliance and sovereignty concerns. Virtuozzo’s software stack, which optimizes hardware utilization through Linux OS enhancements, orchestration, and automation, exemplifies this trend, while CloudPe’s GPU cloud services now contribute up to 25% of its revenue amid surging generative AI demand. This maturation signals a market moving beyond the GPU boom into a cloud-centric era where operational excellence and software sophistication define competitive advantage.

Sources

Enterprise AI Gets Secure and Smart

HPE and Nvidia’s new stack brings rigorous governance, advanced security, and unified management to enterprise AI, setting a new bar for scalable, production-grade agentic systems.

In June 2026, HPE and NVIDIA jointly launched the HPE Private Cloud AI platform, a turnkey solution designed to securely transition agentic AI from experimental stages into governed enterprise production. Integrating NVIDIA Vera CPUs within the HPE ProLiant Compute DL394 Gen12 server and leveraging the NVIDIA Agent Toolkit for behavior monitoring and policy enforcement, this full-stack architecture enables enterprises to scale autonomous AI agents with centralized governance and operational efficiency, as emphasized by HPE CEO Antonio Neri and NVIDIA CEO Jensen Huang.

Security and governance are paramount in this platform, with HPE incorporating technologies such as Zerto Software for continuous data protection, NVIDIA Confidential Computing for cryptographic attestation, and NVIDIA BlueField with DOCA to enforce zero-trust principles and runtime threat detection. Complemented by the HPE Alletra Storage MP X10000’s claim of up to 20x faster token response times, these innovations collectively support metadata-driven data preparation critical for large-scale AI deployments, ensuring both robust protection and high performance.

Building on this foundation, HPE’s GreenLake Intelligence framework and Morpheus Software upgrades introduced in mid-2026 unify AI operations by delivering centralized management, orchestration, and governance across infrastructure and applications. Features like self-service provisioning, integrated observability via OpsRamp Operations Copilot, and Morpheus Orchestration Copilot facilitate efficient, scalable AI deployment while strategic ecosystem enhancements—including deeper Citrix DaaS integration and expanded air-gapped private cloud deployments—further bolster secure and flexible infrastructure options for enterprises.

Addressing operational efficiency and security in GPU-intensive AI workloads, vCluster’s Kubernetes virtualization technology emerged as a game-changer by enabling isolated cluster experiences on shared infrastructure. By supporting both fully shared clusters and virtualized control planes with dedicated worker nodes, vCluster reduces costly cluster sprawl and underutilized GPUs while enhancing security through stronger isolation than traditional namespaces. Its recent bare metal provisioning automation accelerates transforming GPU racks into managed services, a critical advantage for NeoCloud and sovereign cloud providers, and its developer-friendly Kubernetes-native tooling streamlines rapid ephemeral cluster creation and integrated monitoring, delivering significant FinOps and operational benefits.

Sources

Part of these trends

Get the stories behind the trends

Deep-dive reporting and the weekly brief, in your inbox.