On-prem AI surges as enterprises demand data sovereignty

The gist

Enterprises are ditching the cloud and bringing AI in-house as data sovereignty, compliance, and control become non-negotiable in the race for AI supremacy.

What to know

  • Dell, NVIDIA, and Intel have unleashed on-prem AI infrastructure—like Deskside Agentic AI, PowerRack, and SN50 RDU—giving enterprises scalable, sovereign AI options by early 2026.
  • Korea Telecom’s KT NPU LLM Station, launched August 2026, powers air-gapped, regulation-friendly AI using domestic chips and language models to comply with South Korea’s strict data laws.
  • Kasm Technologies and Intel now offer containerized private AI workspaces that run LLMs securely on Xeon 6 CPUs, letting regulated industries deploy AI at scale without sending data to the cloud.

On-Prem AI Goes Mainstream

Enterprises are abandoning cloud AI pilots for production-ready, on-premises infrastructure that delivers both data sovereignty and rapid scalability.

By early 2026, Dell spearheaded a pivotal shift in enterprise AI infrastructure with the unveiling of its Deskside Agentic AI and the PowerRack turnkey rack system, alongside the ultra-dense ObjectScale X7700 storage appliance boasting 45% greater density. These innovations, integrated within the expanded Dell AI Factory in partnership with NVIDIA, marked a strategic move from experimental cloud-based AI pilots to robust, production-scale on-premises deployments. This evolution emphasized enhanced data sovereignty and cost control, enabling enterprises to run autonomous AI agents locally and scale rapidly while maintaining tight control over proprietary data.

NVIDIA and Dell’s leadership underscored a fundamental realignment from cloud-centric AI toward on-premises intelligence, driven by the imperative to safeguard sensitive and proprietary data within enterprise walls. Jensen Huang articulated that intelligence must be generated 'at the point of context,' meaning AI agents operate where the secure, company-specific data and skills reside—on-premises rather than in distant cloud environments. This approach not only enhances data sovereignty but also aligns AI deployments with the unique operational realities of companies like Samsung and Lily, where localized AI agents are essential for future manufacturing and innovation.

By mid-2026, enterprises increasingly embraced the vision of building and managing their own local AI models and data centers as a trust and security imperative. Industry voices highlighted that creating proprietary AI models internally is the only way to ensure full containment and trustworthiness of sensitive information, rejecting reliance on third-party inferencing and model building due to pervasive concerns over data exposure. This shift is supported by technological advances in data center miniaturization and power efficiency, transforming what once required entire rooms of servers into compact, manageable units, thus making on-premises AI infrastructure both feasible and economically attractive.

Looking ahead, there is a growing expectation that capital markets and public shareholders will exert pressure on enterprises to adopt more cost-efficient AI data center solutions, framing this transition as a critical margin conversation. The drive for operational efficiency in powering local AI infrastructure is anticipated to catalyze innovation in energy management and hardware design, ultimately reducing expenses and enabling broader adoption of sovereign, secure on-prem AI deployments. This financial impetus complements the technical and strategic motivations, signaling a comprehensive evolution in enterprise AI infrastructure.

Sources
QCwireBloomberg PodcastsTo The Point - Cybersecurity

Intel’s Agentic AI Blueprint

Intel’s hybrid hardware stack and Project Terafab are set to redefine local AI performance and semiconductor capacity, signaling a new era in enterprise AI hardware.

In August 2026, Intel unveiled a cutting-edge blueprint for agentic AI tools developed in collaboration with SambaNova, integrating the SN50 RDU, Xeon 6 processors, and GPUs to deliver high-throughput, low-latency AI inference optimized for the x86 software stack. This design reflects a strategic push to enhance localized AI workloads by combining specialized hardware components that balance memory bandwidth and acceleration, with a targeted launch in the latter half of 2026. Concurrently, Intel's Project Terafab, bolstered by partnerships with SpaceX, xAI, and Tesla, aims to revolutionize semiconductor fabrication by producing over one terawatt of compute annually, signaling a significant leap in AI hardware capacity and reinforcing Intel's competitive edge in the AI semiconductor market.

Sources

Korea’s Air-Gapped AI Revolution

KT’s fully sovereign AI appliance, built on domestic chips and language models, sets a new standard for legally compliant, high-performance AI in regulated sectors.

In August 2026, Korea Telecom unveiled the KT NPU LLM Station, a groundbreaking sovereign AI appliance that integrates a domestically produced inference chip from Rebellions with a Korean-developed large language model, all housed within a single on-premises server. This innovation directly addresses South Korea's stringent mangjuri regulation, which mandates physical air-gapping of networks in sensitive sectors, effectively barring cloud-based AI services. By confining every byte of AI computation within the customer's facility on Korean-made silicon, KT's solution pioneers a new paradigm for sovereign AI deployments tailored to stringent legal and technical constraints.

The KT NPU LLM Station is engineered to meet the demanding needs of regulated industries such as government agencies, defense contractors, pharmaceutical companies, financial institutions, and manufacturers, all of which face legal requirements to isolate their networks from external connections. Leveraging Rebellions' ATOM-MAX neural processing unit, the appliance delivers exceptional computational power—128 teraflops FP16 and 512 TOPS INT8—while surpassing NVIDIA's L40S in energy efficiency measured by tokens-per-second per watt. This robust hardware configuration supports models with up to 70 billion parameters on a single server, ensuring high-performance AI capabilities within the confines of sovereign infrastructure.

Powering this sovereign AI appliance is KT's Mi:dm K 2.5 Pro, a 32-billion-parameter enterprise reasoning model designed specifically for complex document analysis and agentic workflows in regulated environments. Its expansive 128,000-token context window enables sophisticated retrieval-augmented generation tasks critical for document-heavy industries, reinforcing the appliance's suitability for sectors where compliance and data sovereignty are paramount. This combination of advanced hardware and tailored AI models exemplifies how regional sovereign AI solutions can reconcile cutting-edge capabilities with rigorous regulatory compliance.

Sources

Private AI Workspaces at Scale

Kasm and Intel are enabling regulated industries to deploy secure, containerized AI workspaces across entire workforces without sacrificing compliance or resource efficiency.

In August 2026, Kasm Technologies deepened its collaboration with Intel to launch containerized private AI workspaces powered by Intel Xeon 6 processors with Advanced Matrix Extensions (AMX). This innovative platform leverages Intel's OpenVINO toolkit to enable local large language model inference at interactive speeds without relying on GPUs or sending sensitive data outside the enterprise perimeter, effectively addressing critical data sovereignty concerns.

This joint solution is tailored to meet the stringent demands of regulated industries such as healthcare, finance, legal, defense, and government by enabling workforce-scale AI adoption while maintaining compliance and data sovereignty. Jaymes Davis, Chief Technical Evangelist at Kasm Technologies, emphasized that this architecture represents the long-awaited breakthrough for regulated sectors seeking secure, scalable AI deployment across their organizations.

Beyond security and compliance, the platform offers a unified control plane orchestrating AI workspaces across Intel's diverse compute portfolio, including CPUs with AMX, integrated NPUs, and discrete GPUs. Notably, Kasm 1.19 introduces SR-IOV bifurcation of Intel Arc Pro GPUs, allowing a single physical GPU to be shared as multiple isolated virtual functions, enhancing resource efficiency and scalability for enterprise deployments.

Economically, this containerized AI workspace solution has achieved cost parity with traditional per-seat AI subscriptions at around 40 users per node, making it a financially viable option for large-scale enterprise adoption. Early uptake from regulated industries underscores the platform’s appeal, as it balances cost-effectiveness with the critical need for data control and compliance in sensitive environments.

Sources

Part of these trends

Get the stories behind the trends

Deep-dive reporting and the weekly brief, in your inbox.