AI data centers hit boiling point: liquid cooling tech races to quench soaring power demands

Tech Xplore

The gist

AI data centers are in a cooling arms race, as surging compute power turns traditional air cooling into a relic and liquid cooling tech becomes make-or-break for next-gen infrastructure.

What to know

Cooling Crisis Redefines Design

AI’s surging power densities have turned cooling from a routine utility into the defining constraint for next-generation data center architecture and profitability.

By early 2026, the surge in AI workloads has catapulted rack power densities from traditional 10–20 kW levels to staggering figures between 50 and 130 kW, far exceeding the cooling capacity of legacy air systems. This exponential increase, exemplified by NVIDIA’s H100 and Grace Hopper chips drawing several kilowatts each, creates localized hot spots and bursty thermal spikes that air cooling—limited by air’s low heat capacity and practical chip TDP ceilings around 700 W—simply cannot manage. As a result, cooling has evolved from a mere operational expense into a critical architectural constraint that determines which data centers can attract premium AI tenants, forcing a fundamental redesign of facilities rather than incremental upgrades.

The inadequacy of traditional air cooling at high AI densities translates into significant financial and operational risks, including margin compression, stranded capacity, and heightened outage costs. Cooling already accounts for 30–40% of data center energy consumption, with water use reaching 17 billion gallons in 2023, and downtime losses running into six figures per event. Moreover, the scarcity in AI data centers has shifted from power delivery and fiber adjacency to effective heat rejection constrained by water availability, permitting, and reliability, underscoring cooling as the new bottleneck in AI infrastructure viability.

This cooling crisis has accelerated a decisive industry pivot from air to advanced liquid cooling architectures, including direct-to-chip and immersion cooling, which can handle densities up to 600 kW per rack as seen in hyperscale facilities like Vera Rubin and CPX. Liquid cooling not only surpasses air cooling’s physical limits but also enables energy efficiencies by eliminating chillers and enabling free cooling with warm water at around 45 °C—a shift hailed by NVIDIA CEO Jensen Huang as revolutionary for supercomputer-scale deployments. However, retrofitting legacy air-cooled data centers is costly and complex, often requiring complete redesigns of floor loading, piping, and electrical systems, pushing providers to either invest heavily or cede AI workloads to hyperscalers built for liquid cooling from the ground up.

Compounding the challenge, AI data centers now contend with unprecedented physical constraints as equipment weight per cabinet has surged from under 1,000 lbs to over 4,200 lbs, and power densities have reached half a megawatt per cabinet in some NVIDIA configurations. This necessitates not only advanced cooling but also sophisticated design tools like computational fluid dynamics (CFD) modeling to mitigate hot air recirculation and uneven airflow that plague traditional air-cooled setups. Such modeling has demonstrated potential cooling efficiency gains of around 15%, highlighting how precise airflow management is vital to reducing energy consumption, preventing hardware damage, and ensuring the operational reliability of these ultra-dense AI environments.

Sources
Global Data Center HubSEMIVISION @_@Global Data Center HubThe Data Center Frontier ShowData GravitySuper Data Science: ML & AI Podcast with Jon Krohn

Diamond Tech Disrupts Heat

Lab-grown diamond cooling and advanced liquid systems are slashing power needs and unlocking ultra-dense AI compute without massive infrastructure overhauls.

By early 2026, Akash Systems pioneered the application of lab-grown diamond cooling technology directly onto GPUs, achieving temperature reductions of 10 to 15 degrees Celsius and enabling significantly higher compute densities without the need for new power infrastructure. Felix Jackham, CEO of Akash, emphasized that this approach not only slashes cooling power consumption—pushing Power Usage Effectiveness (PUE) closer to the ideal 1.0—but also delivers substantial economic benefits, estimating savings of about one million dollars per server. Partnering with industry leaders like Supermicro, Nvidia, and AMD, Akash has integrated diamond cooling into ready-to-use servers, demonstrating a practical path to scaling AI compute efficiently within existing data center footprints.

The thermal demands of next-generation AI chips such as NVIDIA’s H100 and Grace Hopper, with TDPs reaching several kilowatts, have rendered traditional air cooling obsolete, catalyzing a rapid shift toward advanced liquid cooling methods. Innovations like direct-to-chip cooling, high-temperature liquid loops, and dual-loop architectures enable data centers to operate with server outlet water temperatures around 45°C, facilitating free cooling without energy-intensive chillers. This transition not only reduces operational costs and carbon footprints but also supports the extreme rack densities—up to 150 kW per rack—required by AI workloads, as highlighted by Frost & Sullivan’s 2026 whitepaper and Nvidia’s Rubin AI data center deployment.

Emerging liquid cooling technologies are pushing the envelope beyond conventional single-phase systems, with approaches such as two-phase cooling, microchannel architectures, and nuclear-inspired Adaptive Phase Cooling (APC) gaining traction. Ferveret, an MIT spinout, exemplifies this innovation by adapting subcooled boiling techniques from nuclear reactors to generate tiny, rapidly detaching bubbles that accelerate heat transfer without using water or toxic chemicals. Their APC technology promises a 15% boost in computational power efficiency and a 35% increase in AI model output when combined with power control, addressing the urgent need to curb data center electricity consumption projected to reach 17% of U.S. national usage by decade’s end.

The maturation of liquid cooling is reflected in broad industry adoption and infrastructure evolution, where turnkey solutions like JetCool’s SmartPlate and large-scale deployments by BDC demonstrate the scalability and operational benefits of direct-to-chip and immersion cooling. These systems reduce CPU and GPU temperatures by up to 11%, cut server power consumption by 30%, and enable AI data centers to handle power densities up to half a megawatt per cabinet, as seen in Nvidia’s GB300 super pod. However, this shift demands comprehensive redesigns of data center facilities—including enhanced power feeds, plumbing, and leak management—to accommodate the complexity and reliability requirements of liquid cooling at industrial scale.

Sources
SiliconANGLE theCUBESEMIVISION @_@PR Newswire - Consumer TechnologyThe Data Center Frontier ShowThe VergeTech Xplore

Water Use Hits a Tipping Point

Closed-loop and hybrid cooling breakthroughs are dramatically shrinking data center water footprints, enabling AI growth even in drought-prone regions.

By early 2026, the AI data center industry had widely embraced closed-loop water cooling systems, which recirculate water internally to absorb and dissipate heat without relying on external water sources or generating waste, thereby significantly reducing water consumption. Companies like OpenAI and Microsoft highlighted that these systems cut fresh water withdrawal dramatically, with Microsoft projecting savings of over 125 million liters per site annually. However, closed-loop cooling still involves some water use through initial fills, periodic top-offs, and indirect consumption via power generation, underscoring the complexity of the overall water footprint.

Innovations in cooling technology are pushing the boundaries of sustainability by combining water efficiency with energy savings. Oracle's direct-to-chip, non-evaporative closed-loop systems achieve near-zero community water usage while slashing server fan power consumption by up to 80%, and Nvidia's Rubin AI data center leverages a 75% water and 25% glycol coolant mix operating at 45°C to eliminate nearly all water usage and reduce power needs. These advances not only address environmental concerns but also enable data centers to be sited in arid or water-stressed regions, expanding location options beyond traditional water-abundant areas.

Hybrid and waterless cooling approaches further illustrate the industry's commitment to minimizing water use while maintaining energy efficiency. The Isambard AI data center's hybrid closed-loop system operates primarily in dry mode, using evaporative cooling only during high temperatures, resulting in annual water consumption comparable to that of just 21 typical UK homes. Meanwhile, companies like AAON and Ferveret are deploying waterless chiller and adaptive phase cooling systems, respectively, which eliminate water consumption altogether by leveraging cold climates or specialized liquid immersion techniques that accelerate heat transfer without traditional water-based methods.

Industry-wide efforts to balance sustainability with operational demands are exemplified by strategic investments and partnerships aimed at integrating energy-efficient, low-water-use cooling solutions. Ecolab's $4.75 billion acquisition of CoolIT Systems and collaboration with Nvidia to launch near-zero water footprint cooling platforms underscore the growing market demand for liquid cooling technologies that reduce both water and energy consumption in high-density AI data centers. Despite these advances, experts caution that challenges remain in fully assessing sustainability benefits due to factors like construction impacts, overall energy demands, and upfront costs.

Sources

Standards and Scale Drive Change

Unified energy frameworks and billion-dollar M&A moves are accelerating liquid cooling adoption, turning thermal management into a key battleground for hyperscalers.

By mid-2026, cooling has emerged from a peripheral facility concern to a strategic cornerstone of AI data center performance and sustainability, as highlighted by Frost & Sullivan's Monica Miches and Prem Shanmugam. This shift is driven by the escalating thermal demands of AI workloads, prompting widespread adoption of advanced liquid cooling architectures, direct-to-chip systems, and innovative thermal management technologies like two-phase cooling and microchannel designs. Investment decisions increasingly prioritize cooling not just for operational uptime but as an integral part of energy optimization and scalable infrastructure, positioning cooling innovation as a critical competitive differentiator among hyperscalers, cloud providers, and component suppliers.

The collaborative release of the AI Data Center Energy Performance Framework by ASHRAE, NEMA, and PNNL in June 2026 marks a pivotal industry milestone, unifying over a dozen standards into a dynamic, online resource that addresses energy sourcing, thermal management, electrical systems, and water use across the full data center lifecycle. This framework not only tackles the complex challenges of rising rack densities and power consumption but also fosters seamless integration between power distribution and cooling systems, as emphasized by NEMA's Debra Phillips. It supports operators in managing costs, enhancing reliability, and preparing for future AI infrastructure growth, reflecting a market-wide commitment to resilience, efficiency, and grid-interactive designs including microgrids and on-site energy storage.

Market dynamics in 2026 reveal rapid growth and consolidation within the AI data center cooling ecosystem, exemplified by the direct-to-chip coolant market's projected tenfold expansion from $133 million in 2025 to over $1.3 billion by 2032, fueled by hyperscale adoption of dense GPU configurations. Major strategic moves, such as Ecolab's $4.75 billion acquisition of CoolIT Systems and its partnership with NVIDIA to launch integrated, energy-efficient cooling solutions, underscore the sector's escalating investment momentum. This consolidation trend, alongside active M&A by players like Vertiv, is reshaping the landscape to accelerate technology adoption at scale, with hyperscalers and neocloud providers leading demand despite regulatory hurdles in the U.S.

The evolving AI data center landscape demands that energy performance frameworks remain living, adaptable resources to keep pace with rapid technological and grid changes. As articulated by PNNL's Bing Liu and echoed by industry leaders like Patrick Hughes, the Framework integrates electrical infrastructure, building systems, and energy research to bridge traditional silos and enable coordinated approaches to power distribution, cooling, and thermal management. This cross-industry collaboration supports both new builds and modernization of existing facilities, balancing innovation with reliability and long-term planning, while electrical manufacturers innovate with next-generation solutions such as 800V DC systems and liquid cooling for power delivery components to meet soaring power density requirements.

Sources

Operational Readiness Gets Smart

Digital simulation, turnkey liquid systems, and rigorous commissioning are making reliable, high-density AI cooling achievable at scale—before the first server powers on.

By mid-2026, operational readiness and reliability in AI data center cooling have evolved into a sophisticated interplay of advanced fluid management, precise system orchestration, and early-stage commissioning rigor. Closed-loop secondary fluid networks using sub-2.5 micron filtered glycol-based fluids, as detailed in the June 23 analysis, minimize contamination and fluid loss, while innovations like heat rejection systems with giant radiators have drastically reduced water consumption, aligning sustainability with operational stability. Complementing these hardware advances, ChemTreat’s Operational Readiness Framework emphasizes integrating cooling expertise early in the construction lifecycle to prevent corrosion and ensure measurable turnover standards, with Jacob Paugh underscoring that "long-term reliability is often decided before a system ever comes online."

Turnkey solutions such as JetCool’s liquid-cooled Dell XE7745 system exemplify the new standard for operational readiness by delivering factory-tested, rack-level infrastructure with unified warranties and end-to-end deployment services, significantly reducing operational risk and enabling rapid scaling of AI compute density without major facility overhauls. The integration of SmartPlate direct-to-chip cooling not only manages thermal loads up to 8 kW per server but also slashes power consumption by 30% and acoustic output by 23 dB, demonstrating how liquid cooling advances directly support resilience and efficiency in AI workloads. Leveraging Flex’s global manufacturing footprint, JetCool’s approach also anticipates the hybrid cloud future, simplifying high-density AI adoption across distributed environments.

The future of AI data center cooling increasingly hinges on digital innovation and simulation-driven design to accelerate deployment and enhance reliability. Modelon’s Data Center Library, launched in late June 2026, empowers infrastructure teams with AI-assisted simulation and cloud collaboration tools that model the entire cooling chain, enabling faster, energy-efficient designs that avoid costly overbuilding and reduce operating expenses. Similarly, Lehigh University’s use of validated computational fluid dynamics (CFD) models to optimize airflow and heat transfer illustrates how precision simulations eliminate guesswork, supporting resilient, scalable infrastructure that aligns with hybrid cloud demands and evolving AI workloads.

Operational readiness now also embraces system-level validation and monitoring innovations to ensure reliability under real-world stresses and diverse deployment scenarios. ASUS’s environmental chamber tests full racks under extreme temperatures and humidity, revealing critical dependencies unique to liquid cooling and enabling accelerated aging analyses that balance coolant temperature with component longevity. Meanwhile, IT Availability’s DeepCool Analytics™ platform introduces AI-driven monitoring and cybersecurity tools that enhance immersion cooling system resilience, reflecting a broader industry shift toward integrated digital management. These advances, coupled with industrialized, pre-engineered data center modules that cut deployment times by up to 85%, underscore a strategic pivot toward rapid, reliable, and hybrid-ready AI infrastructure.

Sources

Part of these trends

Get the stories behind the trends

Deep-dive reporting and the weekly brief, in your inbox.