Rafay, Nvidia tokenize AI cloud infrastructure

The gist

Rafay Systems and NVIDIA are rewriting the AI cloud playbook—ditching traditional GPU sales for a token-metered, multi-tenant model that’s turbocharging global, governed AI infrastructure.

What to know

  • By mid-2026, Rafay, NVIDIA, Cisco, and Dell pivoted from hardware sales to token-based, self-service AI cloud services spanning private, hybrid, and sovereign clouds.
  • NVIDIA’s July 2026 revenue-sharing model powered Sharon AI and Firmus Technologies to deploy 210,000 Grace Blackwell GPUs for nonstop inference workloads across Australia and Indonesia.
  • The UAE’s Ras Al Khaimah data center launched a Sovereign AI Compute platform with NVIDIA GPUs, setting a new bar for data residency, compliance, and national AI ambitions.

Tokenized AI Clouds Rise

Rafay, NVIDIA, and partners redefined AI infrastructure by shifting from hardware sales to governed, token-metered cloud services, unlocking new recurring revenue streams and operational agility.

By mid-2026, the AI infrastructure market was undergoing a fundamental shift from traditional GPU hardware sales and rentals toward token-metered, multi-tenant AI cloud services. Rafay Systems spearheaded this transition with its Elevate AI infrastructure ecosystem, partnering with industry giants like Cisco, Dell Technologies, Unisys, and NVIDIA to deliver governed, self-service GPU compute environments across private, hybrid, and sovereign clouds. This evolution enabled new monetization strategies that moved beyond one-time hardware transactions to continuous, usage-based billing models, reflecting a broader industry trend toward full-stack AI service delivery.

In early July 2026, NVIDIA catalyzed this market transformation by unveiling a revenue-sharing and credit-support model designed to alleviate the capital bottleneck for AI cloud providers. This innovative approach allowed NVIDIA to share in recurring cloud revenue while still earning traditional hardware sales, effectively positioning itself as a long-term infrastructure partner rather than a mere hardware vendor. Early adopters such as Sharon AI and Firmus Technologies demonstrated the model’s scale and viability by deploying a combined total of approximately 210,000 Grace Blackwell GPUs, enabling expansive AI cloud infrastructure deployments that support the shift from one-time model training to continuous inference workloads.

The collaboration between Aolani and Rafay in late July 2026 marked a pivotal moment in operationalizing AI cloud services, moving beyond raw GPU deployment to production-grade, multi-tenant AI platforms. By integrating NVIDIA’s DSX OS on GB200 NVL72 hardware with Rafay’s orchestration and multi-tenancy software, the joint platform enabled secure, self-service provisioning of Kubernetes clusters, AI workspaces, and inference environments. This deployment exemplified the industry’s first wave of governed AI cloud services that dramatically reduce manual integration efforts and accelerate time-to-value for customers.

Concurrently, the AI infrastructure market witnessed a surge in demand from a diverse range of startups, particularly those underserved by existing providers, for flexible, on-demand compute access without long-term contracts or complex hardware setups. Hybrid cloud models emerged as the new standard, allowing companies to maintain private clouds while leveraging interconnected, multi-tenant services. Providers like San Francisco Compute capitalized on this trend by emphasizing trusted, technically recognized teams to attract well-funded AI startups seeking reliable compute resources, echoing the early AWS adoption model where ease of access and flexibility were paramount.

Sources
PR NewswireNVIDIA BlogPR Newswire - Business TechnologySiliconANGLE theCUBE

UAE’s Sovereign AI Surge

The UAE’s Ras Al Khaimah data center and Sovereign AI Compute platform set a new standard for national data residency, compliance, and AI service delivery, fueling a regional race for AI autonomy.

By mid-2026, the UAE emerged as a regional pioneer in sovereign AI infrastructure, launching its first live sovereign AI data center in Ras Al Khaimah through a partnership between Innovation City and IOPn’s Siada. Equipped with NVIDIA B200 GPUs, this facility was designed to meet stringent regional demands for data residency and sovereignty amid heightened GCC regulatory scrutiny, targeting sectors such as fintech, digital health, and government. This initiative underscored the strategic importance of keeping data and compute within national jurisdiction to comply with governance and sovereignty requirements, a trend echoed globally as enterprises increasingly favor private AI infrastructure for control, latency, and compliance benefits, as highlighted by Dell’s Venkat Sitaram.

Building on this momentum, the UAE further advanced its sovereign AI ambitions with the July 2026 launch of the Sovereign AI Compute platform, a collaboration between e& UAE and Core42 that integrates Core42’s Sovereign AI Cloud with national-scale digital infrastructure. This platform offers managed AI services with no upfront capital expenditure or egress fees, featuring bundled GPU resources and carrier-grade connectivity to accelerate the country’s National Strategy for Artificial Intelligence 2031. Executives like Esam Mahmoud and Jaafar Al Hashmi emphasized that this initiative not only simplifies and secures AI deployment at scale but also positions the UAE as a leader in fostering an AI-native economy and agentic AI capabilities.

The rise of sovereign AI clouds is part of a broader shift toward hybrid and distributed AI infrastructure models that balance private and public resources to optimize workload management, governance, and cost control. Core42’s expansion of its US data center capacity alongside its regional UAE deployments exemplifies this hybrid approach, blending sovereign regional data centers with international cloud presence. Meanwhile, regional providers like Rafay Systems are capitalizing on the growing demand for local AI infrastructure that guarantees data and compute sovereignty by delivering enterprise-grade features such as quotas, policies, auditability, and security controls—transforming themselves into critical AI utilities within their markets.

Operational agility and rapid deployment capabilities have become crucial competitive differentiators for sovereign and neocloud AI providers facing tight governance and compliance mandates. As Rafay Systems’ Budhani notes, the company’s most significant investment has been in its 'delivery muscle,' enabling customers to quickly operationalize AI services while meeting local regulatory requirements. This emphasis on software and service delivery accelerates time to revenue and underscores that beyond hardware and infrastructure, the ability to provide seamless, compliant, and scalable AI consumption experiences is key to success in the evolving sovereign AI cloud landscape.

Sources

Rafay’s Orchestration Breakthrough

Rafay’s Managed Model Control Protocol and production-grade AI platforms embed governance, security, and agentic operations directly into enterprise AI workflows, accelerating operational readiness.

By mid-2026, Rafay Systems had firmly positioned itself as a pivotal player in AI infrastructure orchestration, expanding its Elevate AI ecosystem to deliver governed, multi-tenant, self-service GPU compute environments across private, hybrid, and sovereign clouds. This evolution from simple GPU rental to token-metered AI service delivery reflects a strategic shift toward full-stack operational readiness, enabling providers and enterprises to monetize AI workloads with enhanced governance, security, and operational control, as evidenced by Rafay’s NVIDIA AI Cloud-Ready platform integrations.

Rafay’s introduction of the Managed Model Control Protocol (MCP) Server marked a significant operational innovation by securely embedding AI assistants into infrastructure workflows without data export or custom integrations. Initially targeting Kubernetes environments for fleet intelligence, cost attribution, and incident diagnosis, the MCP Server leverages role-based access controls to provide real-time operational context and governance, thereby reducing manual integration efforts and laying the groundwork for agentic AI operations under strict enterprise governance.

The July 2026 collaboration between Aolani and Rafay to launch a production-ready AI platform on NVIDIA’s GB200 NVL72 hardware with DSX OS integration exemplifies the maturation of AI infrastructure orchestration. This platform enables secure, centralized provisioning of Kubernetes clusters, VMs, AI workspaces, and inference environments through a unified interface, dramatically reducing manual overhead and accelerating enterprise time-to-value. As CEOs Nicholas Chia and Haseeb Budhani emphasized, this initiative signals a strategic move from raw GPU deployment toward operationally ready, governed AI services at scale.

Rafay’s focus on the operational software layer addresses the critical need to transform costly AI hardware into secure, reliable, and profitable cloud services by integrating orchestration, networking, security, multitenancy, and developer experience into a cohesive platform. By engaging customers early—during infrastructure purchasing rather than post-deployment—Rafay accelerates time to revenue and supports diverse consumption models including bare metal, Kubernetes, serverless, and token-based offerings. This comprehensive approach is particularly vital in the rise of sovereign AI clouds, where local infrastructure must meet stringent governance, auditability, and security standards shaped by hyperscale cloud expectations.

Sources

Channel Partners Shift Gears

Channel partners are moving beyond hardware sales to offer integrated AI services and lifecycle management, driving sustainable profitability through hybrid, governed cloud models.

By mid-2026, Rafay Systems had strategically expanded its Elevate AI infrastructure ecosystem through partnerships with industry giants like Cisco, Dell Technologies, Unisys, and NVIDIA, signaling a decisive shift from mere GPU rental models to comprehensive, token-metered AI service delivery. This evolution not only integrates hardware with governed, multi-tenant, self-service GPU compute environments across private, hybrid, and sovereign clouds but also fosters sustainable revenue growth for channel partners via incentives, joint selling, and technical enablement, positioning Rafay as a leader in operationalizing AI cloud infrastructure.

Dell’s Venkat Sitaram underscores the necessity for channel partners to transcend initial AI hardware sales by cultivating adjacent services such as cyber resilience, data protection, and lifecycle modernization, thereby securing long-term, sustainable profitability. This paradigm shift reflects the broader evolution of AI infrastructure engagements into multi-year partnerships that leverage hybrid cloud and distributed architectures—including edge deployments—to optimize workload management, governance, and cost control.

The growing prominence of private AI infrastructure, particularly in regulated sectors like BFSI and government, is driven by its advantages in customer control, lower latency, economic efficiency, and compliance with stringent governance and data sovereignty requirements. As Sitaram highlights, these verticals are key demand drivers encouraging partners to integrate deep infrastructure expertise with broader AI services, thereby enhancing their value proposition and ensuring sustainable profitability in a complex, evolving market.

Sources

Serverless GPUs Meet Privacy

Token-metered, serverless GPU models from NVIDIA, Protopia, and Rafay are converging with privacy-enhancing multi-tenancy, transforming idle infrastructure into secure, revenue-generating AI utilities.

In July 2026, NVIDIA revolutionized AI cloud infrastructure financing by introducing a token-metered, revenue-sharing model that significantly lowers upfront GPU costs for providers while enabling NVIDIA to earn recurring revenue from usage. This model underpins large-scale, serverless GPU provisioning exemplified by Sharon AI and Firmus Technologies deploying a combined 210,000 Grace Blackwell GPUs across Australia and Indonesia, targeting continuous inference workloads that demand flexible, on-demand compute access. By shifting from traditional reserved capacity to token-based metering, this approach aligns with privacy-enhanced multi-tenancy principles, allowing multiple enterprises to securely share infrastructure without sacrificing performance or control.

Building on NVIDIA’s foundation, Protopia advanced serverless GPU infrastructure by enabling token-metered, shared GPU slices with enhanced security layers, allowing providers to monetize GPU usage per second while offering enterprises more cost-effective and flexible access. As one analyst explained, this model eliminates the inefficiencies of pre-allocated but idle GPUs, addressing enterprise pain points around high capital expenditure and underutilized reserved capacity. Crucially, Protopia’s innovations intertwine privacy with compute, recognizing that robust security is not just a compliance checkbox but a critical enabler for broader enterprise adoption of shared GPU resources.

By late July 2026, Protopia AI and Rafay Systems integrated token-metered, serverless GPU access with privacy-enhancing multi-tenancy technologies, such as Rafay’s Token Factory and Protopia’s Stained Glass, which transforms raw inputs into secure, stochastic representations. This synergy maximizes GPU utilization by drastically reducing idle capacity and converts governance and security from cost centers into revenue drivers for infrastructure providers. As the GPU-as-a-service market expands, this fusion of tokenomics, privacy, and multi-tenancy is emerging as a critical framework for scalable, cost-efficient, and secure AI compute access across diverse enterprise and regional providers.

Sources

Cisco Bets Big on AI

Cisco’s expanded Rafay partnership and bullish revenue forecasts highlight a strategic pivot toward AI-powered networking and governance—though concentrated hyperscaler exposure keeps investor optimism in check.

Cisco's expanded partnership with Rafay Systems, announced at Cisco Live US 2026, significantly bolsters its AI networking and governance capabilities by integrating Rafay’s AI orchestration platform with Cisco’s Nexus One infrastructure. This strategic move not only deepens Cisco’s foothold in enterprise AI and security spending but also reinforces its central role in emerging neocloud and sovereign AI networks. However, investor sentiment remains cautiously optimistic due to Cisco's heavy reliance on a concentrated hyperscaler customer base, which poses risks if spending from these key clients slows, highlighting a delicate balance between AI-driven growth and concentrated market exposure.

Management’s projection of US$9.0 billion in fiscal 2026 AI infrastructure orders has driven Cisco to raise its full-year revenue guidance to approximately US$62.8–63.0 billion, underscoring AI as a pivotal growth engine. This optimism is echoed in bullish long-term forecasts, with some analysts anticipating revenues as high as US$81.3 billion and earnings near US$19.6 billion by 2029. These elevated expectations hinge on the sustained success of AI partnerships like Rafay, which serve as both a catalyst for growth and a litmus test for the company’s ability to maintain momentum in a rapidly evolving market.

The market responded positively to Cisco’s Rafay partnership, with shares climbing approximately 3.5% shortly after the announcement, reflecting investor approval of Cisco’s enhanced AI and cybersecurity offerings. This collaboration not only broadens Cisco’s presence in fast-growing technology sectors but also strengthens its appeal to enterprise clients focused on digital transformation. Nonetheless, technical analysis indicates that Cisco’s stock is currently overbought, suggesting a near-term consolidation phase between $114.11 and $119.51, which points to a tempered investor enthusiasm despite the strong fundamentals.

Sources

Part of these trends

Get the stories behind the trends

Deep-dive reporting and the weekly brief, in your inbox.