AI data centers turn to modular liquid cooling

The gist
AI data centers are racing to reinvent themselves with liquid cooling, modular designs, and unprecedented power density to meet explosive AI demand—rewriting the playbook for speed, scale, and sustainability.
What to know
- Rack power densities are skyrocketing from 10–20 kW to over 150 kW by 2026, forcing a shift from air to direct-to-chip liquid cooling solutions like those from NVIDIA.
- Prefabricated, modular data centers from Supermicro and Runware now deploy thousands of racks in days, slashing build times and sidestepping grid bottlenecks by tapping underused power infrastructure.
- NVIDIA’s Vera Rubin platform and edge-focused innovations like Infinium Edge’s EdgeSites are redefining performance by tightly integrating compute, cooling, and networking—while chasing ultra-low energy costs and decentralized, low-latency inference.
Power Surges Reshape Cooling
AI’s unpredictable, millisecond-scale power swings are forcing data centers to overhaul both power delivery and cooling, with native 800V DC and liquid cooling now essential to survive the new era of ultra-dense, high-variance workloads.
By late 2025, the volatile nature of AI workloads became a clear challenge for data center power delivery, with power loads swinging dramatically from 10% to 100% within milliseconds to minutes, far exceeding traditional CPU load patterns. This volatility strained conventional AC to DC power conversion methods, which introduced inefficiencies and harmonics, prompting a shift toward native 800-volt DC power delivery directly to racks to better handle the surging and fluctuating power densities.
The explosive growth in AI workloads drove rack power densities from the traditional 10–20 kW range to staggering levels of 50–150 kW and beyond by mid-2026, with bleeding-edge configurations like Nvidia's liquid-to-chip cooled super pods consuming up to 20 megawatts in a single 10,000 sq ft room. This leap rendered legacy air cooling systems inadequate, as they could not economically or efficiently dissipate the intense heat generated, especially given localized thermal spikes and hot spots caused by dense GPU clusters drawing 700–1,000 watts per chip.
As rack power densities crossed the 15–25 kW threshold, data center operators faced a critical inflection point: either overhaul mechanical cooling systems with advanced solutions like direct-to-chip liquid cooling, rear-door heat exchangers, or immersion cooling, or risk losing premium AI tenants. This necessity elevated cooling from a manageable operational expense to a fundamental architectural constraint, determining which facilities could compete in the lucrative AI workload segment, a reality underscored by enterprises and providers such as Aligned Data Centers who now emphasize scalable, on-demand liquid cooling to future-proof infrastructure without excessive upfront costs.
The transition from air to liquid cooling reflects a paradigm shift in data center design driven by AI’s unprecedented power density and thermal management demands. While air cooling remains viable for lower-density AI deployments, the superior heat transport capacity of liquids—capturing heat closer to the chip—enables more compact, efficient, and high-performance designs. Industry leaders like Nvidia and Microsoft are pioneering direct liquid-to-chip and microfluidic cooling technologies to meet these challenges, signaling a move toward integrated infrastructure approaches that unify power delivery, cooling, and water resources to sustain the rapidly evolving AI ecosystem.
Next-Gen Cooling Breakthroughs
Diamond-based heat sinks, co-packaged photonics, and hybrid liquid-air systems are slashing energy use and unlocking record GPU densities, pushing data centers toward near-perfect efficiency and million-dollar savings per server.
By mid-2025, the integration of co-packaged silicon photonics into AI networking infrastructure marked a pivotal advancement, dramatically reducing power consumption and component complexity. Moving optical engines inside switch packages enabled a 3.5x reduction in power usage and allowed data centers to connect three times more GPUs within the same power envelope, significantly enhancing scalability and resiliency in AI data centers.
Direct-to-chip liquid cooling has emerged as an indispensable technology to manage soaring rack power densities, which have surged from around 10kW to over 100kW, with projections exceeding 200kW in next-generation AI servers. Companies like BDC and NVIDIA have pioneered this approach since 2022, deploying tens of megawatts of capacity, while hybrid cooling strategies enable phased modernization of existing data centers by combining liquid cooling for AI workloads with air cooling for traditional systems.
Innovations in cooling materials and architectures have further revolutionized thermal management: Akash Systems’ diamond cooling technology uses lab-grown diamonds to passively reduce GPU temperatures by 10–15°C, pushing data center Power Usage Effectiveness (PUE) closer to 1.0 and potentially saving operators around one million dollars per server by eliminating complex cooling infrastructure. Meanwhile, Dell’s breakthrough parallel cooling system for AI workstations combines dual-pump flow with a novel Element 49 thermal interface material, enabling efficient, plug-and-play cooling of up to 1000-watt GPUs and scalable mini AI factories via CX8 interconnects.
Photonic networking and advanced semiconductor packaging are converging to unlock unprecedented AI compute density and efficiency. Photonic interconnects now enable ultra-high bandwidth communication across tens of thousands to millions of GPUs, accelerating foundational model training cycles from six months to as little as two and supporting AI models with hundreds of trillions of parameters. Complementing this, Samsung and SK hynix are embedding thermal management directly into next-generation memory designs, such as zHBM and HBM4E, reducing thermal resistance and boosting power efficiency, while Compute Express Link (CXL) memory pooling optimizes data flow to alleviate thermal bottlenecks.
Modular Data Centers Go Mainstream
Containerized, factory-built AI data centers are transforming deployment speed and flexibility, letting operators bypass grid delays and repurpose idle buildings for high-density compute in weeks instead of years.
The AI infrastructure landscape is rapidly evolving toward modular, prefabricated, and containerized data center designs that dramatically accelerate deployment timelines while enhancing scalability and local service delivery. Industry leaders like Supermicro and Runware are pioneering this shift with solutions such as Supermicro’s Data Center Building Block Solutions (DCBBS) and Runware’s Sonic Inference Pods—containerized 1MW AI data centers deployable in days rather than months or years. These modular systems integrate compute, cooling, and power into unified platforms managed via open standards, enabling plug-and-play installation that bypasses traditional multi-year construction cycles and utility interconnection delays. For example, Supermicro’s manufacturing capacity of up to 3,000 AI racks monthly and Runware’s plan to deploy 10,000 nodes across 160 sites by 2026 underscore the scale and speed of this transformation, meeting surging demand across hyperscale and edge environments.
This modular revolution extends beyond new builds to innovative repurposing of existing infrastructure, as exemplified by Infinium Edge’s EdgeSites™, which convert underutilized commercial and industrial buildings into AI-ready facilities using factory-built Vector ONE™ units. By leveraging existing behind-the-meter electrical capacity—often only 40% to 60% utilized—Infinium Edge circumvents utility interconnection bottlenecks and water-use constraints, slashing deployment times from years to mere months. This approach aligns with the broader industry trend toward distributed AI infrastructure, bringing high-density compute closer to end users and supporting inference workloads in diverse locations without the need for new construction or municipal water connections.
The construction methodologies underpinning modular data centers are also evolving to further compress timelines and improve reliability. Transitioning from traditional precast concrete shells, which can still take 18 to 20 months to deliver as seen with CloudHQ’s LC-2 facility, the industry is embracing pre-engineered metal buildings (PEMB) and steel modular designs. These factory-cut, labeled, and punched components enable rapid on-site assembly through bolting, reducing material use and supporting scalable, repeatable single-story halls optimized for AI workloads. This shift not only accelerates build times but also enhances cost predictability and scalability, embodying the ‘pay-as-you-grow’ model that is critical for meeting the explosive growth in AI compute demand.
Strategic partnerships and site acquisitions are key enablers of this modular acceleration, as demonstrated by CleanSpark’s 271-acre Austin County site acquisition, which leverages a meet-me fiber backbone and a partnership with a major ERCOT substation developer to fast-track energization of 207 megawatts by early 2027—years ahead of typical greenfield timelines. This collaboration facilitates behind-the-meter generation and modular scalability, illustrating how integrated infrastructure planning complements modular design to overcome traditional grid and permitting delays, thereby enabling rapid, large-scale AI data center deployment.
Rack-Scale AI Systems Dominate
NVIDIA, Supermicro, and ZTE are driving a shift from piecemeal hardware to fully integrated, liquid-cooled AI racks—locking in performance gains, reducing integration headaches, and creating formidable competitive moats.
NVIDIA's Vera Rubin platform epitomizes the evolution from discrete AI components to fully integrated rack-scale AI systems by unifying six major silicon domains—including Rubin GPUs, Vera CPUs, DPUs, and advanced networking—into a cohesive architecture that delivers unprecedented interconnect bandwidth and system-level optimization. This shift, marked by the NVL72 rack housing 72 Rubin GPUs and 36 Vera CPUs interconnected via NVLink 6 fabric with 260 TB/s bandwidth, redefines AI infrastructure by embedding power delivery, liquid cooling, and software tightly alongside hardware, thereby creating significant switching costs and competitive moats that extend beyond traditional GPU-centric designs.
Complementing NVIDIA's platform, industry leaders like Supermicro and Compal have advanced integrated rack-scale AI solutions that emphasize system-level co-design by unifying compute, power, cooling, and networking into modular, factory-assembled racks. Supermicro’s DCBBS portfolio, capable of delivering up to 10x throughput per watt compared to prior NVIDIA Blackwell systems, exemplifies this trend with plug-and-play AI racks supporting multiple standards and liquid cooling, enabling rapid deployment at scale—as CEO Charles Liang highlights, the goal is to 'accelerate AI cluster deployment, maximize compute density, and address persistent integration challenges.'
The move toward integrated rack-scale AI platforms also extends beyond NVIDIA and its direct partners, as exemplified by ZTE’s SuperPOD architecture, which integrates high-density GPUs with unified power, cooling, and networking to optimize token generation efficiency and energy consumption. ZTE’s approach reduces product adaptation cycles from over a year to six months through pre-integration, enabling scalable AI infrastructure from single racks with 128 GPUs to clusters of 16,000 GPUs, thereby demonstrating how system-level co-design and software-hardware synergy are critical for commercial AI factory deployments.
By mid-2026, the Vera Rubin platform's production ramp and cloud validation by providers like CoreWeave underscore a pivotal industry shift where rack- and pod-level metrics—such as intra-rack fabric bandwidth, DPU integration, and power/cooling management—supersede traditional GPU FLOPS as the primary performance constraints. CoreWeave’s DeepSeek-R1 benchmark revealed Vera Rubin's ability to generate roughly 10 times more tokens per megawatt than previous architectures, signaling a new era where AI infrastructure success hinges on holistic system co-design rather than isolated chip performance.
Inference Moves to the Edge
The pivot from training to real-time inference is triggering a wave of distributed, sovereign, and waterless AI pods that cut latency, maximize local power use, and slash operational costs by up to 90%.
As AI workloads pivot from large-scale model training to continuous, real-time inference across billions of devices, the industry faces a looming bottleneck in inference infrastructure characterized by stringent demands for sustained efficiency, low latency, and privacy preservation. Kneron’s 2026 warning underscores that traditional centralized data centers, optimized for periodic training tasks, are ill-suited for this shift, with energy demands from data centers projected to nearly double by 2030, intensifying power and cooling challenges. This evolution necessitates a fundamental rethinking of AI infrastructure, emphasizing distributed and edge deployments that operate efficiently at the data source to meet these new operational imperatives.
Emerging distributed AI inference models leverage existing power infrastructure by situating micro data centers near utility substations or repurposing commercial and industrial buildings with underutilized electrical capacity, as exemplified by Nvidia’s pilot network and Infinium Edge’s EdgeSites™. These approaches circumvent lengthy utility interconnection delays and reduce grid strain by tapping into spare behind-the-meter power, enabling rapid, modular deployments that scale from 1MW units to multi-megawatt sites. Infinium Edge’s proprietary immersion cooling technology achieves a remarkable power usage effectiveness (PUE) of 1.05 while occupying up to 70% less floor space, illustrating how energy-efficient, space-conscious designs are critical to unlocking cost-effective, low-latency AI inference closer to end users.
Sovereign AI infrastructure demands rapid, jurisdiction-specific deployments that integrate advanced energy management, modular cooling, and stringent regulatory compliance, a complexity highlighted by Craig Tavares and the World Economic Forum’s 2026 analysis. Runware’s Sonic Inference Pods epitomize this trend, delivering 1MW, waterless liquid-cooled AI data centers in 20-foot containers deployable within weeks across the US, Europe, and Asia-Pacific. These pods support multi-tenant architectures with robust encryption and dynamic asset sharing, addressing sovereign mandates while reducing inference costs by up to 90%. Their distributed network design enhances resilience by routing workloads to the nearest available pod, marking a decisive shift from monolithic hyperscale centers to nimble, edge-oriented sovereign AI factories.
Innovative edge AI infrastructure models are also emerging within urban environments by harnessing spare power in Class A commercial office buildings, as Perimeter Compute’s initiative demonstrates. By deploying GPUs in underutilized spaces such as basements, Perimeter not only accelerates AI inference speeds through proximity to users but also navigates local grid constraints via collaboration with utilities and revenue-sharing with building owners. This strategy exemplifies a broader industry move away from massive, land-intensive data centers toward distributed, modular deployments that balance performance, cost, and regulatory demands while mitigating community resistance and infrastructure bottlenecks.
Speed, Sustainability, and Scale
Solar heat batteries, hybrid cloud models, and a relentless focus on 'Time to Token' are redefining AI infrastructure, as the industry races to deliver exascale power and instant inference with minimal waste or delay.
By mid-2026, the AI infrastructure landscape is rapidly evolving to prioritize a delicate balance between speed, sustainability, and scalability. Companies like Exowatt are pioneering modular, factory-made solar heat battery systems—such as their exawatt P3 container-sized units—that leverage mature technologies like Fresnel lenses and rock-based heat batteries to deliver dispatchable renewable power at a targeted cost as low as $0.01 per kilowatt hour, enabling gigawatt-scale deployments that can be linearly scaled to meet any project size. This renewable energy backbone addresses the surging power demands forecasted to double global data center consumption to 200 gigawatts by 2030, while also tackling critical bottlenecks in power, cooling, and land availability highlighted by the World Economic Forum’s 2026 report.
The AI infrastructure race is shifting from a singular focus on raw GPU capacity to a nuanced 'two-speed' strategy that combines massive exascale clusters for training with distributed, latency-sensitive inference capacity deployed closer to users via regional data centers, edge nodes, and on-device chips. This approach, underscored by the World Economic Forum and exemplified by France’s Alice Recoque supercomputer slated for 2026, reflects a broader pivot toward hybrid cloud models where customers intelligently allocate workloads between on-premise and cloud environments. As Kaushik notes, this hybrid model will replace the current project-oriented mindset, making AI infrastructure as ubiquitous and seamlessly integrated as everyday productivity tools like Excel or PowerPoint.
Speed to market has emerged as the paramount challenge for AI infrastructure deployment, with the industry adopting 'Time to Token'—the duration from initial planning to first AI output—as the critical success metric. Modern deployments demand deep, partnership-based orchestration among power, cooling, and hardware vendors from day one, enabling converged infrastructure and factory-integrated data center blocks to slash deployment times by up to 85%. This accelerated timeline is vital as idle AI hardware racks worth millions impose severe financial penalties, shifting economic priorities toward maximizing utilization and operational cost-effectiveness, particularly for inference workloads where efficiency trumps raw compute power.
Future-proofing and flexibility are becoming indispensable as AI workloads grow unpredictably and infrastructure constraints tighten. Investments in advanced hybrid cooling solutions—combining liquid and air cooling to support extreme rack densities up to 600 kW—alongside modular, energy-efficient architectures like ZTE’s OEX SuperPOD, which integrates 128 GPUs per rack and supports scaling to 16,000 GPUs, exemplify this trend. These designs not only reduce product adaptation cycles from over a year to six months but also embed sustainability through innovations like 800V HVDC power supplies and full-stack liquid cooling, enabling AI factories to operate efficiently with lower carbon footprints amid rising power scarcity and longer development timelines.










