AI workloads flee the public cloud: enterprises build private ‘AI factories’ amid soaring costs and security fears

The gist
Enterprises are stampeding away from public clouds, building private ‘AI factories’ to slash runaway costs and lock down security as AI workloads scale out of control.
What to know
- By early 2026, 56% of enterprises globally will run or plan to run AI inference workloads on private clouds, with public cloud use for these tasks plunging from 56% to 41% in just one year.
- Soaring, unpredictable token-based AI costs and a wave of security attacks—reported by 73% of enterprises—are forcing organizations to seek budget control and regulatory compliance on-premises.
- Vendors like Nvidia, Dell, and HPE are fueling this shift with new AI hardware and cooling tech, while the trend is hottest in Asia Pacific and Japan, where 82% are considering or already repatriating workloads.
AI's Public Cloud Cost Crunch
Runaway token-based AI expenses and unpredictable billing models are forcing enterprises to abandon public clouds in favor of on-premises AI 'factories' that deliver cost control and budget discipline at scale.
Escalating and unpredictable token-based AI usage costs in public clouds have become a significant budgeting headache for enterprises, especially at scale. For example, with 10,000 employees each incurring $200 monthly on services like Anthropic’s Claude, CFOs quickly face unforecasted expenses that strain financial planning. This economic pressure is driving organizations to reconsider their infrastructure strategies, favoring on-premises AI 'factories' that offer unlimited token usage and greater cost predictability, as vendors like Nvidia, Dell, and HPE respond with dedicated solutions enabling secure, internal experimentation before scaling with CFO-approved budgets.
The traditional cloud-first approach is unraveling under the weight of AI’s sustained, high-intensity compute demands, which public clouds are not economically optimized to handle. As one analysis puts it, treating the public cloud as the default no longer works; instead, enterprises must evaluate infrastructure based on workload-specific economics. Private or on-premises environments provide superior cost stability and performance consistency for AI inference workloads, fundamentally reshaping the economic logic of AI deployment and compelling organizations to adopt disciplined governance models to align infrastructure decisions with business priorities.
While token cost inflation is a short-term challenge rather than a market bubble burst, it is accelerating a shift toward economic rationalization and sustainable AI usage. Industry leaders like Chamath Palihapitiya emphasize a new phase of realism, where enterprises tighten AI token budgets and conduct rigorous cost-benefit analyses to ensure AI investments deliver tangible ROI. Case studies, such as Lindy’s migration from Anthropic models to more cost-effective alternatives, demonstrate that switching inference workloads from public clouds can save millions while boosting performance, underscoring the financial imperative behind workload repatriation.
The decision to repatriate AI workloads to private clouds often hinges on workload predictability and persistence. As AI inference transitions from experimental bursts to steady, repeatable processes, enterprises increasingly question the sustainability of renting token capacity from public clouds. This mirrors a longstanding infrastructure calculus: when utilization is constant, owning or controlling the stack becomes more cost-effective than renting. Public clouds retain their appeal for elastic, experimental workloads, but for continuous, high-value internal applications, private cloud ownership offers superior cost control and economic efficiency, aligning with enterprise demands for financial governance and operational stability.
Shadow AI Spurs Governance Shift
A surge in AI-related security breaches and shadow deployments is driving enterprises to private clouds, where stricter controls and data sovereignty frameworks address escalating compliance and risk challenges.
Security, governance, and data sovereignty have emerged as paramount concerns driving enterprises to shift AI inference workloads to private clouds, as highlighted by Broadcom and industry analyses. With 73% of enterprises reporting AI-related attacks, according to Paul Turner, and 37% citing new data protection requirements, private clouds offer the controlled environments necessary to mitigate risks posed by shadow AI—unregulated AI deployments outside formal IT oversight that nearly 80% of APJ IT leaders have encountered. This phenomenon complicates compliance with evolving regulations such as Singapore’s MAS guidelines and India’s data protection laws, underscoring the critical need for private cloud governance frameworks that can manage the complexity of AI’s extensive open source components and patch cycles.
Data sovereignty concerns extend beyond mere data location to encompass control over the entire AI operational environment, including the control plane, as Chris Wolf explains, emphasizing the necessity for enterprises to run AI workloads disconnected from the internet to maintain compliance and governance. This nuanced sovereignty requirement is especially acute in regulated sectors like financial services and healthcare, where AI sovereignty involves practical deployment mandates covering data classification, access control, and trusted private cloud operation, as noted by Craig McLellan. Moreover, 54% of IT leaders identify data residency and sovereignty as leading geopolitical factors influencing infrastructure choices, often elevating these issues to boardroom priorities amid complex cross-border governance challenges.
The operational challenges of shadow AI and fragmented governance, compounded by silos between business units and IT—as 93% of Singaporean organizations report—highlight infrastructure mismatches that private clouds are uniquely positioned to resolve. By consolidating control, auditability, and workflow integration, private clouds address the governance complexities introduced by persistent internal AI agents interacting across multiple systems, which demand enhanced identity, permissions, and policy enforcement controls. This comprehensive approach not only mitigates security gaps identified by experts like Prashanth Shenoy but also aligns with enterprise priorities where control and compliance often outweigh cost considerations in AI workload placement decisions.
The intricate interplay of data sovereignty, regulatory complexity, and AI workload governance is reshaping enterprise strategies, sometimes leading organizations to reconsider market presence in regions with prohibitive compliance costs, as one analyst observed regarding India. The necessity to localize AI models and human-in-the-loop agents in-country inflates operational expenses and complicates scaling, while the inherent difficulty in controlling AI output data—even with monitoring—raises persistent data leakage risks. These realities reinforce the private cloud’s role as a strategic enabler for enterprises seeking to balance innovation with stringent data protection, privacy, and security mandates in an increasingly fragmented geopolitical landscape.
Private AI Goes High-Tech
Breakthroughs in energy-efficient hardware, edge deployments, and cyber-resilient architectures are redefining private cloud AI, enabling enterprises to achieve hyperscale performance without public cloud trade-offs.
The evolving AI infrastructure landscape is moving beyond the traditional race for bigger GPUs toward a nuanced balance of distributed inference, energy efficiency, and resilience, as highlighted by the World Economic Forum. This shift prioritizes regional data centers, edge nodes, and on-device chips over hyperscale public clouds, enabling enterprises to deploy AI inference workloads closer to users and sensitive data to meet real-time demands and regulatory compliance. Innovations such as subsea data centers leveraging seawater cooling, photonic computing, and optical interconnects promise up to tenfold energy efficiency gains, addressing critical bottlenecks in power and cooling that have long constrained private cloud AI deployments.
Advancements in AI-specific hardware, exemplified by NVIDIA's Blackwell architecture available through OEMs like Dell, HPE, and Lenovo, have empowered enterprises to achieve petaflop-scale inference performance on-premises, effectively rivaling hyperscale cloud capabilities. This hardware evolution, combined with hybrid cloud models that strategically layer public cloud elasticity with private data centers optimized for AI workloads and edge deployments, reflects a sophisticated, distributed infrastructure approach. However, supporting such dense AI workloads—often exceeding 100 kilowatts per rack—requires significant re-engineering of data center power, cooling, and network architectures, including the adoption of liquid cooling and fully connected GPU networks using spine-leaf topologies and NVLink to maximize throughput and efficiency.
Security and operational considerations are driving enterprises to embed cyber resilience deeply into AI infrastructure, spanning hardware and software layers to protect training data, model weights, and inference outputs. Privacy-preserving architectures like federated learning and domestically governed quantum-secure networks align closely with private cloud and edge AI strategies, addressing heightened concerns around data sovereignty and governance. This focus is complemented by a 'two-speed' AI infrastructure strategy that pairs exascale training clusters with distributed inference capacity at the edge, enabling enterprises to balance massive compute needs with localized, secure, and latency-sensitive AI workloads.
Enterprises are increasingly adopting a pragmatic, utilization-driven approach to AI infrastructure, favoring ownership or dedicated capacity for steady, predictable workloads that handle sensitive data—such as continuous fraud detection or embedded developer assistants—while reserving public clouds for bursty, experimental, or elastic tasks. This targeted shift toward private and on-premises AI infrastructure is not a wholesale reversal but a strategic response where utilization, latency, governance, and workflow value converge. As a result, flexible data center designs that allow scalable power and cooling expansions are critical to accommodate the cautious, ROI-focused AI adoption pace typical of enterprises, who often proceed with a 'ready, aim, convene a committee' methodology to balance innovation with risk management.
Asia Leads AI Cloud Repatriation
Asia Pacific and Japan are at the forefront of the global shift to private cloud AI, propelled by regulatory pressures, security demands, and a growing need for operational control over sensitive workloads.
By early 2026, a clear enterprise trend has emerged where AI inference and production workloads are increasingly migrating from public to private cloud environments, signaling a shift from experimental pilots to scalable, secure deployments. Broadcom’s Private Cloud Outlook 2026 survey reveals that 56% of organizations globally now run or plan to run AI inference on private clouds, while public cloud usage for these workloads has dropped from 56% to 41% within a year. Prashanth Shenoy, Broadcom’s VP of Marketing, highlights that as AI scales, rising infrastructure costs, security vulnerabilities, and operational complexity are driving this preference for private cloud solutions that offer greater control and predictability.
Operational realities such as security, compliance, and cost predictability are the primary drivers behind the widespread repatriation of AI workloads to private clouds, with 83% of enterprises considering this move and half having already repatriated some workloads. This trend is especially pronounced in regulated sectors like financial services, healthcare, and the public sector, where data sovereignty and governance have become board-level priorities—54% of respondents cite data residency requirements as a key geopolitical factor influencing infrastructure decisions. Analyst Mauricio Sanchez notes that AI intensifies the trade-off between public and private clouds, with enterprises valuing the control and economic predictability that private clouds provide.
Regional dynamics further underscore this shift, with Asia Pacific and Japan showing even stronger momentum toward private cloud adoption for AI workloads. In these markets, 82% of organizations have considered repatriation and 54% have initiated it, driven by priorities around security, cost predictability, performance, and operational control, as noted by Broadcom’s Sylvain Cazard. Additionally, geopolitical concerns about data protection outside the US, where most public cloud providers are based, reinforce private cloud’s appeal as a more compliant and secure environment, according to ABI Research analyst Michela Menting.
Enterprises are increasingly adopting a hybrid cloud approach where public clouds remain the preferred platform for AI experimentation and bursty workloads, but private clouds dominate production AI inference due to their advantages in cost management, governance, and latency. This is exemplified by the rise of the 'private AI factory' concept—enterprise-owned or controlled AI infrastructure optimized for persistent, mission-critical workloads such as fraud detection and developer code assistance. As recurring token costs and data sensitivity grow, organizations like financial institutions and infrastructure teams are opting to control their AI stacks to ensure predictable utilization and tighter security, moving away from the cost unpredictability and governance challenges of public cloud rentals.






