Nvidia’s kyber delays expose AI buildout bottlenecks

The gist
The AI gold rush is screeching to a halt as bottlenecks in GPUs, advanced packaging, and memory send shockwaves through data center supply chains and investor strategies.
What to know
- Nvidia has locked up 60% of TSMC’s CoWoS advanced packaging capacity through 2027, but every constraint eased just shifts pressure downstream—leaving the whole ecosystem teetering.
- Delays and redesigns, like Nvidia’s 12-month setback for its Kyber rack and a forced Rubin Ultra GPU overhaul, expose nonlinear scaling headaches and open the door for rivals like AMD and Google.
- Scarcity mania is driving stocks like AXT Inc. and Lumentum up thousands of percent despite flat revenues, as investors chase the next bottleneck in critical AI materials and supply chains.
Supply Chain Domino Effect
Every fix in AI hardware—whether GPUs, memory, or power—simply shifts the bottleneck further downstream, exposing the entire data center ecosystem to new, unpredictable chokepoints.
By early 2026, AI data center infrastructure revealed a complex web of cascading bottlenecks spanning GPUs, memory, advanced packaging, and power systems, where easing one constraint merely transfers pressure downstream. This systemic fragility is exemplified by TSMC’s CoWoS advanced packaging capacity, sold out through 2026 with Nvidia locking up roughly 60% through 2027, underscoring how chokepoints in specialized components confer structural leverage akin to ASML’s role in semiconductor manufacturing. As one analysis put it, “Removing one constraint shifts pressure to the next layer — which is why aggregate GPU supply numbers mask the actual delivery reality.”
The supply chain crunch extends beyond silicon wafers to critical materials and components such as copper, aluminum, indium phosphide, laser parts, and High Bandwidth Memory (HBM), many booked out over two years. This multi-dimensional shortage is intensified by geopolitical disruptions—like the Strait of Hormuz closure impacting helium supply vital for chip fabrication—and by the rapid quintupled demand for HBM since 2023, forcing Nvidia to cut consumer graphics card production by up to 40% to prioritize AI hardware. Such constraints reveal a systemic bottleneck not just in manufacturing but in the entire AI compute ecosystem.
The evolution of AI infrastructure bottlenecks follows a sequential pattern from GPUs to storage (HBM), then optical interconnects, and finally power and cooling infrastructure, with each alleviation exposing the next choke point. This dynamic is reflected in Nvidia’s $2 billion strategic investments in laser suppliers Lumentum and Coherent, signaling laser components as a critical compute bottleneck, and in the projected 55-gigawatt US data center power shortage through 2028, accompanied by a $122 billion financing gap. Together, these overlapping constraints force architectural shifts—such as moving from copper to optical networking—and geographic limitations, fundamentally reshaping AI data center expansion strategies.
The AI compute demand surge far outpaces the hardware supply chain’s ability to scale, creating at least four overlapping bottlenecks that must be addressed simultaneously to unstick the system. This multi-layered fragility is evident in the fragmentation of silicon supply eroding Nvidia’s monopoly, the slow scaling of power inputs versus rapid compute growth, and the halving of new data center announcements in late 2025 due to local opposition and supply constraints. As one analysis summarized, “The bottleneck is no longer any one thing... solving any one of them doesn't actually unstick the system bottleneck.”
Packaging: The New Battleground
With Moore’s Law slowing, advanced packaging like TSMC’s CoWoS has become the fiercest arena for AI supremacy, as memory and substrate shortages cascade through the semiconductor supply chain.
As Moore’s Law reaches its physical and economic limits around 3nm and 2nm transistor nodes, the semiconductor industry has pivoted towards advanced packaging technologies like TSMC’s Chip on Wafer on Substrate (CoWoS) to sustain AI performance gains. This three-dimensional stacking approach addresses the critical memory access bottleneck by integrating High Bandwidth Memory (HBM) directly with GPUs, transforming chip packaging from a mere protective layer into a complex, performance-critical technology. However, CoWoS capacity is severely constrained, with TSMC’s output sold out through 2026 and Nvidia alone locking up about 60% of this capacity through 2027, underscoring the intense competition for advanced packaging resources amid the generative AI boom.
The surge in demand for HBM memory, which has quintupled since 2023, exacerbates supply shortages as only SK Hynix, Samsung, and Micron produce these specialized components, all largely sold out well into 2026. This scarcity extends beyond memory to critical packaging materials such as wafer-level substrates, printed circuit boards, copper-clad laminates, and multilayer ceramic capacitors (MLCCs), with lead times exceeding 20 weeks. The prioritization of AI server customers for MLCC production further tightens availability for automotive and consumer electronics sectors, illustrating how AI-driven demand cascades through the entire semiconductor materials supply chain.
TSMC is aggressively expanding CoWoS capacity, aiming to nearly quadruple output by late 2026 with monthly wafer production potentially reaching 190,000 to 200,000 by 2027, yet bottlenecks in sub-components like wafer-level substrates and MLCCs limit throughput. Despite these efforts, lead times and supply gaps persist, although the supply-demand gap for CoWoS packaging is expected to narrow from 20% to 10% by the end of 2026. Meanwhile, geopolitical tensions and fierce competition among industry giants—Nvidia projected to consume 55% of CoWoS capacity in 2027, with Google increasing its share—create a high-stakes “clash of the titans” scenario for scarce advanced packaging resources.
Intel’s EMIB-T advanced packaging technology is emerging as a formidable alternative to TSMC’s CoWoS, boasting over 95% interconnect yields and attracting major customers such as MediaTek, Ampere Computing, AWS, Tesla, and Google’s TPU v9 project. This diversification in advanced packaging solutions could reshape supply dynamics by alleviating some pressure on TSMC’s constrained capacity, signaling a potential shift in the semiconductor supply chain landscape as AI hardware demands continue to escalate.
Nvidia’s Scaling Headaches
Nonlinear manufacturing challenges have forced Nvidia to scale back ambitious GPU designs and delay next-gen racks, revealing the physical and logistical limits of ultra-dense AI systems.
By mid-2026, Nvidia confronted significant manufacturing hurdles that forced a strategic redesign of its Rubin Ultra GPU from an ambitious four-die configuration with 16 HBM4E stacks and 1 TB memory to a more conservative two-die design. This shift, driven by complex challenges in packaging, yield, power delivery, and cooling, underscores the nonlinear scaling difficulties inherent in advanced semiconductor integration, where larger packages exponentially increase error sources and mechanical instability. Consequently, Nvidia’s AI accelerator strategy must now compensate for reduced per-package performance through increased package counts or larger rack deployments, potentially raising system-level costs and complicating supply chain dynamics for key HBM memory suppliers like SK hynix and Samsung.
Nvidia’s Kyber NVL144 AI rack system, designed to integrate 144 Rubin Ultra GPUs, has been delayed by over 12 months to 2028 due to the extraordinary manufacturing complexity of its 78-layer PCB midplane. This midplane, a laminated assembly of three 26-layer boards nearly a square meter each, replaces roughly 20,000 copper cable connections and must maintain impeccable signal integrity at 448G+ SerDes speeds, making yield and reliability a formidable challenge where a single defect can scrap the entire board and disrupt all 144 GPUs simultaneously. These manufacturing obstacles not only stall Nvidia’s roadmap but also ripple through the broader AI hardware ecosystem, highlighting the practical limits of scaling dense GPU deployments despite robust demand.
Efforts to mitigate Kyber’s delay through interim solutions, such as the NVL72x2 back-to-back rack architecture that paired two 72-GPU Oberon racks, were ultimately abandoned after hyperscalers like AWS, Microsoft Azure, and Google Cloud rejected the design due to its unconventional layout and operational complexity. This cancellation reflects the broader infrastructural constraints in hyperscale data centers, where power delivery, cooling, and cable management systems are not readily adaptable to novel rack orientations, thereby limiting Nvidia’s short-term scalability options and reinforcing the challenges of evolving AI infrastructure within existing ecosystem boundaries.
Despite these setbacks, Nvidia’s near-term data center compute revenue outlook remains robust, with Rubin Ultra systems entering full production and shipments expected in late 2026, projecting revenues approximately 20% above Wall Street consensus in the second half of fiscal 2027. To bridge the gap left by Kyber’s delay, Nvidia plans to increase sales of the Oberon rack as the mainstay platform during the Rubin Ultra era. However, the ongoing challenges with co-packaged optics technology and PCB manufacturing not only constrain Nvidia’s scaling ambitions but also open competitive opportunities for rivals like AMD and Google, while raising broader industry concerns about the maturity of critical enabling technologies that underpin the AI hardware scaling roadmap.
Scarcity Mania Fuels Valuations
Investors are driving up stocks of obscure component suppliers by thousands of percent, betting on future bottlenecks as AI infrastructure constraints redefine who wins—and who gets sidelined—in the supply chain.
By mid-2026, evolving supply chain bottlenecks in critical AI infrastructure materials such as indium phosphide have dramatically reshaped investor strategies and competitive dynamics. Companies like AXT Inc. and Lumentum Holdings have seen astronomical stock surges—8,484% and 1,100% respectively—despite stagnant or declining revenues, underscoring how market valuations are increasingly driven by scarcity narratives rather than current financial performance. This bottleneck-driven valuation phenomenon has created a competitive moat for suppliers controlling rare materials, as the 'story matters a lot' in investor sentiment amid AI compute demand outpacing supply.
As 2026 progressed, the focus of AI infrastructure bottlenecks shifted from GPUs to memory bandwidth and optical networking, prompting investment funds like the WisdomTree Artificial Intelligence and Innovation Fund to outperform by targeting these emerging constraints. Hyperscaler spending patterns have prioritized memory, storage, and networking components, signaling a strategic pivot in capital allocation that distinguishes AI adoption plays from infrastructure bottleneck exposures. This evolution highlights how investors are refining their approaches to capture value in the increasingly complex AI compute ecosystem.
Looking ahead to 2027, analysts from Nomura and Mizuho warn of an 'epic shortage' in AI semiconductor supply chains, with bottlenecks expanding beyond advanced nodes to packaging, substrates, and optical components, driving unavoidable price hikes and squeezing shipment growth. This scarcity intensifies competition for limited advanced packaging capacity at TSMC, where NVIDIA is projected to consume 55% of CoWoS wafers and Google 27%, heightening capacity battles that could marginalize AMD, AWS, and others. Meanwhile, Intel’s EMIB-T packaging technology, boasting yields above 95%, emerges as a credible challenger to TSMC’s dominance, with major customers like Google and MediaTek exploring adoption, potentially reshaping competitive positioning in AI infrastructure.
NVIDIA’s delay of its next-generation Kyber NVL144 rack architecture to 2028—caused by manufacturing challenges in producing the complex PCB midplane—introduces a significant bottleneck that disrupts its AI infrastructure roadmap and opens a competitive window for AMD and Google. Despite this setback, NVIDIA’s Rubin systems continue strong production, shipping to cloud giants including Amazon, Microsoft, and Google, with data center compute revenue forecasted to exceed Wall Street consensus by 20% in the latter half of fiscal 2027. This juxtaposition of near-term resilience against longer-term supply constraints underscores the evolving competitive landscape shaped by hardware manufacturing limits and capacity allocation battles.








