Hybrid AI takes center stage as soaring memory costs push enterprises beyond the cloud

The gist

Hybrid AI is taking over as sky-high memory costs and chip shortages force enterprises to ditch cloud-first strategies and reinvent their AI infrastructure.

What to know

  • A global DRAM and HBM memory crunch has driven costs up to 125% and maxed out public cloud capacity, leaving enterprises scrambling for alternatives.
  • 87% of CIOs now plan to repatriate at least some AI workloads from public clouds to hybrid setups—combining on-prem, sovereign edge, and cloud for compliance, cost, and latency.
  • Dell leads the charge with its AI Factory platform and liquid-cooled PowerEdge servers, offering up to 87% savings over public cloud APIs and powering a new wave of deskside-to-datacenter AI.

Memory Crunch Reshapes AI

Skyrocketing DRAM and HBM shortages forced enterprises to rethink cloud dependence as hyperscalers prioritized their own AI, leaving customers to confront rising costs, latency issues, and limited capacity.

By early 2026, an acute shortage of DRAM and HBM memory chips, driven by insatiable AI demand, created severe supply-demand imbalances that forced enterprises to accelerate cloud migrations. AWS CEO Andy Jassy highlighted skyrocketing memory costs and limited availability, pushing firms toward hyperscalers with larger capacity pools. However, as hyperscalers prioritized their own explosive compute needs, public cloud capacity for external enterprises tightened, leaving many caught in a squeeze between soaring demand and constrained supply.

This capacity crunch at hyperscalers compelled enterprises to explore alternatives beyond pure cloud strategies. Options like neocloud providers offering bare metal services or costly, complex on-premises builds became necessary despite their operational challenges. Gartner forecasted DRAM prices soaring 125% through 2026, with Samsung already hiking server memory prices by up to 60%, underscoring the unsustainable economics of traditional hardware scaling amid a structural supply crisis expected to persist into 2027.

Rising memory costs and hyperscaler capacity limits also exposed performance and resilience shortcomings in public clouds, especially for latency-sensitive AI workloads. Enterprises found that production environments often underperformed compared to testing due to latency and recovery models that placed failover sites in distant regions, as Steve Spittal noted. These operational realities forced a reassessment of workload placement, emphasizing proximity and recovery design in infrastructure decisions.

By mid-2026, hyperscale cloud capacity constraints were driving a notable resurgence of on-premises infrastructure and hybrid models. Major enterprises, including a Fortune 100 firm in Texas, retained data centers they planned to shutter after being denied additional cloud capacity. Multi-cloud strategies emerged as pragmatic responses, with companies juggling workloads across providers like GCP and Azure, while regulatory factors such as French data sovereignty further cemented private data center retention. As Desai observed, the once-prevailing belief that all workloads would shift to hyperscalers was being decisively challenged.

Sources

Hybrid AI Becomes Standard

Hybrid AI architectures, anchored by sovereign edge and colocation, have overtaken cloud-first as enterprises seek regulatory compliance, data control, and cost efficiency in an era of unpredictable cloud economics.

By mid-2026, hybrid AI architectures have firmly established themselves as the preferred enterprise model, combining on-premises, private cloud, sovereign edge, and public cloud environments to optimize latency, data sovereignty, cost, and operational resilience. Dell’s leadership, with over 5,000 AI Factory customers running Nvidia GPU-equipped servers alongside cloud resources, exemplifies this deliberate three-tier approach—on-device, on-premises, and cloud—each tailored to specific workload demands. This balanced architecture addresses the growing enterprise imperative to keep sensitive data off public networks while leveraging cloud elasticity for burst capacity, reflecting a strategic shift away from cloud-first to hybrid-first AI infrastructure strategies.

The rise of sovereign edge computing and colocation facilities plays a pivotal role in hybrid AI infrastructures by enabling enterprises to meet stringent data residency and regulatory compliance requirements, such as GDPR, HIPAA, and emerging national AI laws. Gartner’s prediction that 20% of workloads will migrate from global public clouds to local or regional sovereign alternatives by the end of 2026 underscores this trend, while nearly 70% of CIOs now favor colocation for AI/ML production workloads due to its high connectivity and power density. This shift not only reduces latency for inference workloads near end-users but also mitigates operational risks and vendor lock-in, as seen in sectors like banking where hybrid cloud balances innovation with control over sensitive data.

Economic considerations and workload predictability are driving enterprises to repatriate AI workloads from costly public clouds to on-premises or private cloud environments, effectively creating 'AI factories' that tokenize workflows internally. Case studies highlight scenarios where research projects costing hundreds of dollars per cloud run become economically viable when shifted on-premises, while persistent AI agents supporting fraud detection or developer assistance benefit from reduced latency and enhanced data governance. This evolving hybrid model balances public cloud’s elasticity for experimentation with on-premises control for repeatable, sensitive, and high-utilization workloads, reflecting a maturation from pilot projects to intentional, scalable AI architectures.

Leading enterprises are advancing hybrid AI infrastructures by integrating sophisticated governance frameworks that are policy-driven, data-centric, and portable across environments, ensuring consistent controls over data lineage, access, retention, and sovereignty regardless of workload location. Industry leaders like Broadcom and HPE are pioneering hybrid cloud and resiliency platforms that embed AI-enabled operations with explainable AI and human oversight to mitigate risks amid rising cyber threats and geopolitical challenges. This holistic approach unifies observability, automation, financial management, and cyber resilience, enabling organizations to optimize AI deployment cycles, reduce total cost of ownership, and maintain strategic flexibility in a complex, multi-environment digital landscape.

Sources

Repatriation Driven by Regulation

Strict compliance and cost pressures are accelerating the migration of AI workloads back to private infrastructure, with regulated industries leading the shift to hybrid models that balance performance and control.

By mid-2026, a clear enterprise trend has emerged where AI workloads are increasingly repatriated from public clouds back to private clouds and on-premises data centers, driven primarily by regulatory compliance, data sovereignty, latency, and cost predictability concerns. Research from Vanson Bourne and Gartner highlights that 87% of CIOs plan partial or full migration away from public clouds, with 56% of organizations running or planning AI inference and production workloads on private cloud infrastructure, a shift particularly pronounced in regulated sectors like financial services, healthcare, and government. Dell’s case study exemplifies this movement with its scalable on-prem AI solutions that package compute, storage, and networking into portable units, enabling enterprises to meet stringent data sovereignty and privacy requirements while controlling total cost of ownership, as cloud AI workloads often prove expensive and slow for large-scale research tasks.

This repatriation is not a wholesale abandonment of the cloud but rather a strategic recalibration toward hybrid AI infrastructure models that balance the elasticity and scalability of public clouds with the control, compliance, and performance benefits of private and edge environments. IDC and Gartner analyses reveal that 64% of digital infrastructure leaders now operate hybrid environments, placing latency-sensitive inference AI workloads closer to end-users via edge or colocation facilities, while reserving large-scale generative AI training for hyperscale clouds. Enterprises like Goldman Sachs illustrate this hybrid approach by integrating on-premises AI agents with human workflows, underscoring the growing importance of AI factories within corporate data centers that demand advanced cooling, power, and security architectures.

Cost predictability and long-term financial efficiency have become critical drivers for repatriation, as enterprises grapple with unpredictable and escalating public cloud expenses. Surveys and case studies reveal that 83% of enterprises are actively considering or have already moved workloads on-premises or to private clouds, with organizations like 37signals and GEICO saving millions by shifting off AWS. Broadcom’s analysis supports this, showing modern private cloud deployments can reduce total cost of ownership by 40-50% for steady-state workloads, while Flexera and IDC data underscore that managing cloud spend remains the top challenge for 84% of organizations. Consequently, many enterprises prefer colocation or MSP-backed private platforms to balance cost control with operational simplicity.

Regulatory and geopolitical pressures, especially around data sovereignty and compliance with frameworks like GDPR, HIPAA, and the EU AI Act, are compelling enterprises—particularly in Europe, Asia Pacific, and regulated industries—to keep sensitive AI workloads on-premises or within private clouds. With 79% of businesses citing sovereignty as a key investment driver and 54% of IT leaders emphasizing residency requirements, organizations are increasingly wary of US-based public cloud providers due to concerns over cross-border data flows and auditability. This has led to a resurgence of private data centers and hybrid strategies that satisfy both innovation demands and stringent legal mandates, as exemplified by a Northern European government’s move to private cloud and a Paris-based enterprise’s strict data localization policies.

Sources

Dell’s Hybrid AI Blueprint

Dell’s AI Factory and liquid-cooled PowerEdge servers are redefining enterprise AI with a three-tier architecture and industry partnerships, slashing costs and enabling secure, scalable AI from deskside to datacenter.

By mid-2026, Dell Technologies has emerged as a frontrunner in hybrid AI infrastructure innovation, advancing its AI Factory platform to seamlessly integrate on-device, on-premises, and cloud layers for optimized workload placement. This three-tier hybrid architecture, exemplified by the deskside AI initiative using Nvidia's NemoClaw software, enables enterprises to achieve up to 87% cost savings over public cloud APIs and payback periods as short as three months, while enhancing data privacy and experimentation speed. Complementing this software evolution, Dell's refreshed PowerRack lineup—with PowerScale storage, high-speed networking, and advanced cooling solutions—provides a robust hardware foundation that parallels Nvidia's enterprise AI system showcases, underscoring a deliberate shift toward hybrid models that reduce latency and improve cost-effectiveness.

Dell's operational strategy extends beyond hardware and software to include comprehensive services and industry-specific blueprints, developed in partnership with firms like ServiceNow, Mistral, CrowdStrike, and Uneeq, integrated via Dell's Automation Platform. This ecosystem approach facilitates smoother transitions from AI infrastructure investments to production-ready solutions, particularly enhancing the success rates of deskside AI deployments. Moreover, Dell's AI Data Orchestration engine, enhanced with Nvidia NIMs and integrated with Omniverse and cuDF analytics, forms a cohesive AI Data Platform that supports both physical AI workflows and digital twin applications, ensuring data lineage and context critical for trustworthy agentic AI outputs.

Addressing the physical and economic constraints of dense AI workloads, Dell has pioneered liquid cooling innovations such as the PowerCool CDU C7000 and introduced eleven new PowerEdge servers optimized for AI, signaling a broader industry shift toward liquid cooling as a standard for operational efficiency. This hardware evolution complements Dell's 'deskside to data center' hybrid AI infrastructure strategy, which spans from local AI workstations running autonomous agents to liquid-cooled rack-scale systems designed for continuous enterprise reasoning. By enabling AI agents to run where data resides, Dell reduces costly cloud inference dependencies and helps enterprises navigate the rapidly emerging physical constraints on AI infrastructure, a challenge facilities teams had not fully anticipated.

Software-defined infrastructure and unified hybrid cloud management platforms are critical enablers for maximizing AI infrastructure efficiency amid supply constraints and rising costs. Techniques such as hypervisor-level compression, memory ballooning, and dynamic processor load balancing allow enterprises to run significantly more virtual machines per physical host, extending hardware life and unlocking stranded capacity through software-defined networking. By virtualizing traditionally hardware-dependent functions like load balancing and security, these platforms eliminate expensive, power-hungry appliances while maintaining or improving performance, empowering customers to intelligently select workload placement across cloud and on-premises environments and abstracting the complexity of agentic AI for seamless day-to-day use.

Sources

Hybrid AI: The New Default

Hybrid infrastructure is now the enterprise norm as agentic AI, liquid cooling, and orchestration tools converge to deliver scalable, resilient, and cost-effective AI beyond the limits of public cloud.

By mid-2026, enterprises are decisively moving beyond cloud-first AI strategies toward hybrid infrastructures that strategically place AI workloads closer to data sources to optimize performance, cost, and governance. Dell’s 'deskside to data center' approach exemplifies this shift by enabling on-premises agentic AI to reduce latency and bandwidth costs, while IDC predicts that by 2028, 75% of AI workloads will run on hybrid platforms. This evolution is underscored by the need for sophisticated orchestration layers and rigorous data lineage controls, as emphasized by Dell COO Jeff Clarke, to ensure trustworthy and scalable AI operations across distributed environments.

The surge in agentic AI workloads is catalyzing a fundamental redesign of data center architectures, with Dell projecting $50 billion in AI server revenue for FY2027 and investing heavily in liquid-cooled, ultra-dense servers like the PowerEdge M9825. This infrastructure evolution supports continuous, large-scale inference and persistent multi-agent workflows, addressing the economic and physical constraints of cloud-only AI. As Kaushik notes, enterprises are becoming savvy in hybrid workload placement, balancing cloud and on-premises deployments to harness the strengths of each environment effectively.

Hybrid AI models and infrastructure are rapidly normalizing, supported by innovations in containerization, Kubernetes, and sovereign cloud deployments that enable distributed AI workloads with enhanced governance and compliance. Industry leaders such as Broadcom and HPE are advancing hybrid cloud platforms that integrate AI-enabled operations with explainable AI and human oversight to mitigate risks, reflecting a strategic shift where hybrid is no longer transitional but an end state. Gartner’s forecast that 90% of organizations will adopt hybrid cloud by 2027 further cements this trajectory, as enterprises prioritize sovereignty, cyber resilience, and operational control across diverse environments.

Looking ahead, the hybrid AI infrastructure landscape is poised to embrace emerging physical AI and robotics, marking a transformative phase in AI applications. As agentic AI capabilities become as ubiquitous and user-friendly as everyday software tools, the integration of physical AI promises to redefine enterprise operations and user experiences. This anticipated evolution invites a retrospective perspective on today’s agentic AI excitement, highlighting the continuous innovation cycle that will shape AI’s strategic role in business and technology ecosystems.

Sources

Part of these trends

Get the stories behind the trends

Deep-dive reporting and the weekly brief, in your inbox.