AI infrastructure race turns to utilization, cost

The gist

The AI infrastructure arms race has shifted from hoarding GPUs to a high-stakes battle over utilization, efficiency, and slashing operational costs.

What to know

  • Industry leaders like QumulusAI and Hyperbolic are building unified platforms to eliminate idle GPU capacity and optimize both training and inference on the same stack.
  • Neoclouds such as Zettabyte now offer bare-metal NVIDIA H100 access at $1.99/hour—about one-third the cost of hyperscalers—while legacy data centers and energy innovations drive up to 25% server energy savings.
  • Enterprises are choosing between GPU-as-a-service, self-hosting, and hybrid models as IBM and others roll out serverless GPU fleets, with the neocloud market set to grab 20% of the $267B AI cloud market by 2030.

From Idle GPUs to Efficiency

AI infrastructure leaders are shifting focus from raw GPU acquisition to maximizing utilization, tailoring stacks for real-time inference and dynamic workloads.

By mid-2026, AI infrastructure priorities have decisively shifted from the aggressive acquisition of GPU capacity toward maximizing utilization, efficiency, and operational cost-effectiveness, especially for continuous inference workloads. Industry leaders like QumulusAI’s Mike Maniscalco and Hyperbolic’s Jasper Zhang emphasize that idle GPU capacity represents a critical cost burden, prompting a focus on flexible infrastructure that supports both large-scale production inference and smaller-scale training or fine-tuning on the same platform. This evolution marks a fundamental departure from the training-centric buildouts of prior years, as highlighted by HyperFrame Research’s Steven Dickens, who underscores the necessity of differentiated CPU-to-GPU ratios, workload orchestration, and deployment strategies tailored to the distinct demands of production AI services.

The transition from separate training and inference stacks toward unified, workload-tuned AI infrastructure is driving operators to customize configurations across storage, networking, and tenancy models to meet diverse performance and latency requirements. QumulusAI adapts Nvidia reference architectures with options ranging from local NVMe to tiered external storage and varying network designs, balancing budget and latency constraints. This approach reflects a broader industry recognition that inference workloads—now projected by Deloitte to constitute two-thirds of AI compute by 2026 and growing at a 79% CAGR—demand global distribution, unpredictable scaling, and heterogeneous hardware setups that traditional training clusters cannot efficiently support.

Treating AI model deployment as a continuous, latency-sensitive service rather than a batch job has become a critical operational imperative, necessitating dedicated practices such as latency SLOs, cost tracking, canary deployments, and the physical separation of inference and training infrastructure. This shift is driven by the complexity of production AI environments that run dozens of models with varying precisions and scaling needs, making fixed hardware ownership increasingly inefficient. Cloud-based, flexible infrastructure models now dominate due to their ability to adapt to variable demand and reduce idle capacity, fundamentally altering the economics and operational paradigms established during the GPU training boom.

Innovations in high-density AI infrastructure, exemplified by ZTE’s SuperPOD with up to 16,000 GPUs, alongside architectural breakthroughs like the OEX midplane-free design and AI-native hardware accelerations such as DPU-enabled KV caches, are enabling unprecedented token generation efficiency and sustainable continuous inference operations. These system-level synergies—integrating advanced power supplies, liquid cooling, and intelligent compute-electricity coordination—reflect a holistic approach to utilization optimization that transcends isolated component upgrades. Moreover, pre-integration and pre-adaptation strategies have slashed deployment cycles from over a year to six months, accelerating ecosystem convergence and facilitating agile rollouts tailored to the demands of large-scale, long-context, and high-concurrency AI workloads.

Sources

Legacy Data Centers Reborn

Operators are slashing AI costs and energy use by retrofitting 1990s data centers with next-gen chips and autonomous optimization software.

By mid-2026, AI inference centers have significantly cut costs by repurposing legacy data center infrastructure from the 1990s, enabling operations without the need for new power capacity investments. This strategy leverages existing compliance and building frameworks, while integrating air-cooled SNOVA chips that reduce energy consumption compared to traditional liquid-cooled hardware. Together, these innovations not only lower the cost per AI token but also create a comprehensive operational efficiency model that maximizes the value of pre-existing assets.

ZTE’s AI factory concept exemplifies a holistic approach to cost optimization through multidimensional co-design, integrating computing, networking, storage, and energy systems to boost token generation efficiency. Their OEX architecture-based SuperPOD achieves high-density, energy-efficient deployments with up to 16,000 GPUs, utilizing a midplane-free, zero-cable design to minimize bottlenecks and latency. Complementing this, ZTE’s AI-native KV cache powered by DPU acceleration attains over 70% cache hit rates with microsecond latency, while energy innovations like 800V HVDC power and full-stack liquid cooling enable sustainable, lower-carbon AI operations.

The integration of autonomous AI software by QiO Technologies marks a pivotal shift in operational efficiencies, as their system autonomously optimizes data center energy consumption without human intervention. Validated by Intel, QiO’s solution delivers up to 25% reduction in server energy use, exemplifying how intelligent, real-time adjustments can dramatically lower operational costs and energy footprints. This autonomous approach complements ecosystem offloading strategies, where companies like Enthropping delegate optimization workloads externally, freeing internal compute resources for critical tasks and further driving down AI operational expenses.

ZTE’s pre-integration and pre-adaptation strategy accelerates AI infrastructure deployment by slashing product tuning cycles from over a year to under six months, enabling faster commercial rollouts without sacrificing flexibility. This rapid ecosystem convergence supports scalable, cost-effective AI factories that balance performance and sustainability, reinforcing the importance of streamlined development processes in achieving total cost of ownership (TCO) optimization.

Sources

GPUaaS vs. Self-Hosting Showdown

Enterprises are weighing flexible, usage-based GPU rentals against the control and predictability of owned hardware, with hybrid models gaining traction for complex AI needs.

Enterprises face a nuanced decision when choosing between renting GPU resources via GPU-as-a-service (GPUaaS) and owning or self-hosting GPUs, with each model presenting distinct trade-offs in cost, control, and performance. GPUaaS, exemplified by offerings like Zettabyte’s zCLOUD MaaS with NVIDIA H100 GPUs priced as low as $1.99 per hour, provides flexible, usage-based pricing that eliminates hefty upfront capital expenditures and infrastructure challenges such as power and cooling demands highlighted by Vasily Mazin. This model suits organizations with sporadic or experimental AI workloads, offering rapid deployment and operational agility through on-demand, reserved, or spot pricing options, but it also raises concerns about vendor lock-in and long-term cost sustainability.

Self-hosting GPUs, whether through dedicated servers, colocation, or renting hardware for private environments, remains the preferred approach for enterprises with substantial, latency-sensitive AI workloads requiring stringent data privacy and cost predictability. By 2026, reserved cloud GPU pricing has made long-term commitments more economical, with discounts up to 60% for three-year reservations, enabling continuous operation of large models like a 70B-parameter LLM at around $2,000 monthly on platforms such as CoreWeave. However, owning hardware entails significant upfront capital, ongoing operational expenses for power and cooling, and risks associated with secondary-market GPUs lacking warranties, underscoring the importance of engineering capacity and workload characteristics in deployment decisions.

Hybrid deployment models are emerging as a compelling compromise, offering enterprises the ability to rent raw GPU capacity or leverage managed model endpoints within a unified platform, as demonstrated by Zettabyte’s zSUITE. This approach balances the control and data residency advantages of self-hosting with the convenience and scalability of GPUaaS, enabling rapid cluster deployment in under six hours and ensuring high reliability with a 99.5% uptime SLA. Such flexibility is particularly valuable for organizations managing diverse workloads that require both low latency and operational simplicity.

Cost efficiency in deployment models also hinges on workload composition; while managed APIs may be cheaper for single large models due to reduced idle GPU waste, self-hosting becomes more cost-effective when running multiple smaller models that can share GPU resources and maintain high utilization. This insight underscores the strategic value of self-hosting for enterprises aiming to optimize GPU usage across embedding, reranking, and extraction tasks, thereby minimizing idle capacity and maximizing return on investment.

Sources

Neoclouds Disrupt Hyperscalers

Specialized neoclouds are capturing a growing share of the AI market by delivering bare-metal GPU access at a fraction of hyperscaler prices and scaling rapidly across regions.

Neoclouds have rapidly emerged as specialized AI infrastructure providers focused exclusively on renting NVIDIA GPU compute capacity, offering bare-metal access optimized for large-scale AI workloads. Services like Zettabyte’s zCLOUD MaaS exemplify this trend by providing cost-effective, flexible access to cutting-edge GPUs such as the NVIDIA H100 at $1.99 per GPU-hour, with rapid cluster deployment options and integrated access to popular open-source AI models via a unified API. This specialization allows neoclouds to reduce operational complexity and accelerate AI deployment across hyperscale and edge environments, positioning them as agile alternatives to traditional hyperscalers.

Despite the dominance of hyperscale cloud providers, neoclouds carve out a competitive niche by delivering vertically integrated, AI-native infrastructure that combines bare-metal GPU access, high-bandwidth InfiniBand networking, and optimized storage solutions to maximize performance and efficiency. Analysts like Hardeep Singh and industry leaders emphasize that neoclouds are architected for single-tenant, thousands-of-GPUs workloads with minimal hypervisor overhead, contrasting sharply with hyperscalers’ legacy multi-tenant VM models. This architectural focus, coupled with close partnerships with chip manufacturers such as NVIDIA, enables neoclouds to offer GPU capacity at roughly one-third the cost of hyperscalers, with reported savings of 60% to 70%.

The neocloud market is experiencing explosive growth and significant regional expansion, with revenues surpassing $25 billion globally in 2025 and forecasts predicting a 20% share of the $267 billion AI cloud market by 2030. However, McKinsey notes that only a small subset—around 10 to 15 neoclouds—operate at meaningful scale, reflecting an evolving but competitive landscape. Australia has emerged as a key proving ground for this expansion, highlighted by Sharon AI’s $373 million five-year deal to deploy over 2,000 NVIDIA Blackwell Ultra B300 GPUs inside NEXTDC data centers in Melbourne and Sydney, with plans to scale to more than 55,000 GPUs by mid-2027. This regional growth underscores neoclouds’ strategic focus on filling critical supply gaps amid tightening frontier GPU availability.

Sources

Workload-Tuned AI Infrastructure

Providers are abandoning one-size-fits-all GPU stacks in favor of infrastructure tailored to the specific demands of training and inference, with latency and geography now critical factors.

Tailoring AI infrastructure to workload-specific demands is essential, as a one-size-fits-all approach centered solely on the latest GPU hardware fails to address the nuanced requirements of diverse AI applications. Distinguishing between training and inference workloads is particularly critical; while training focuses on computational intensity, inference demands prioritization of user experience factors such as latency and data storage. As one expert emphasized, "Inference really is where the delivery point... we have to focus on the application, not necessarily what the language model is learning," underscoring the need for infrastructure that aligns closely with the end-use scenario.

Managed Service Providers (MSPs) encounter significant challenges in optimizing AI infrastructure due to factors like geographic placement and latency, which can directly influence customer retention. A cautionary example involved a provider losing clients because infrastructure was poorly located, leading to unacceptable latency and degraded application performance. This highlights the operational complexity MSPs face in balancing capacity planning, tuning, and regional compliance—challenges that GMI Cloud addresses through unified billing, multi-region deployment across North America, Europe, and Asia-Pacific, and robust SLAs guaranteeing 99.9% uptime with autoscaling and burst capacity.

GMI Cloud exemplifies how tailored AI infrastructure can unify training and inference on a single platform, streamlining workflows and eliminating the need to rebuild stacks between phases. Their offering includes three distinct capacity shapes—Bare Metal GPU for full hardware control during large-scale training, Managed GPU Clusters for distributed multi-node training, and Container Services for Kubernetes-based GPU orchestration—each designed to meet specific workload demands. For inference, GMI Cloud’s Prime Inference service leverages dedicated single-tenant GPUs with runtimes tuned per model to optimize latency and throughput, delivering up to 500,000 tokens per minute per GPU while mitigating cold start penalties and noisy neighbor effects through warm endpoints and isolation.

Sources

IBM’s Serverless AI Leap

IBM is integrating Nvidia’s B300 clusters and serverless fleet models to deliver compliant, high-performance AI compute for regulated industries—redefining hybrid cloud deployment.

IBM's deployment of Nvidia's HGX B300 clusters on IBM Cloud represents a leap forward in AI hardware integration, delivering up to 144 petaflops of FP4 precision compute tailored specifically for regulated industries such as healthcare, financial services, and government. By early 2026, IBM plans to expand this offering with a serverless fleet model that allows enterprises to consume GPU resources without the burden of infrastructure management, significantly enhancing flexibility and simplifying AI workload deployment.

The launch of IBM's B300 AI chips, developed in partnership with Nvidia and Together AI, marks a pivotal advancement toward industrial-scale AI computing. This collaboration underscores the industry's shift from small-scale AI setups to building reliable, high-speed, and scalable data centers, which remain a critical challenge. Moreover, vertical integration—where companies control both the hardware and AI models—emerges as a strategic economic advantage, enabling better cost management and operational efficiency at scale.

IBM's integration of HGX B300 nodes with Red Hat OpenShift and its watsonx platform creates a unified hybrid-cloud environment that supports regulated AI workloads while maintaining strict compliance by keeping sensitive data within controlled boundaries. This seamless orchestration of containerized AI workflows not only addresses regulatory demands but also exemplifies how integrated platforms are evolving to support complex, secure AI deployments across hybrid infrastructures.

Sources

Part of these trends

Get the stories behind the trends

Deep-dive reporting and the weekly brief, in your inbox.