AI’s bottleneck shifts to memory bandwidth

Artificial Intelligence Made Simple ↗

The gist

AI's breakneck pace is now throttled not by raw compute, but by the brutal limits of memory bandwidthffwith supply chain gridlock, triopoly pricing, and new architectures all battling to keep up.

What to know

SRAM Outpaces HBM Limits

AI hardware is now defined by how close and fast memory can serve data, with SRAM-based chiplets shattering energy and bandwidth barriers but running up against capacity walls.

By late 2025, it became clear that memory bandwidth, rather than raw compute power, is the fundamental bottleneck in AI inference workloads. d-Matrix’s innovative digital in-memory compute (DIMC) approach exemplifies this shift by integrating SRAM directly into compute cores, achieving staggering bandwidths up to 150 TB/s—orders of magnitude faster than traditional HBM-based GPUs capped around 2-8 TB/s per chip. This architectural choice not only accelerates data access but also slashes energy costs by a factor of ten, dropping from roughly 3 pJ/bit with HBM to 0.3 pJ/bit with SRAM, underscoring the critical role of memory proximity and bandwidth in efficient AI inference. However, the limited SRAM capacity (256 MB per chiplet) constrains model size, necessitating multi-chiplet designs to scale both memory and compute for large models, highlighting the ongoing tradeoff between bandwidth and capacity in AI hardware design.

The roofline model, extensively analyzed by early 2026, elucidates the stark disparity between compute-bound and memory-bound AI workloads. Prefill operations, with high arithmetic intensity (e.g., 4096 FLOPs per byte), fully utilize GPU compute resources, whereas decode phases operate at roughly 1 FLOP per byte—far below the H100’s ridge point of ~295—causing GPUs to idle over 95% of the time waiting for data. This structural mismatch means that despite GPUs’ massive parallelism and tensor cores optimized for compute-heavy tasks, the decode phase is fundamentally memory bandwidth limited, a bottleneck that no amount of raw compute scaling can overcome. Consequently, memory bandwidth—not peak FLOPS—dictates inference throughput, especially during token generation where the entire model weights and growing KV cache must be streamed repeatedly from HBM.

The KV cache, which stores attention states to avoid recomputation, has emerged as a dominant driver of memory bandwidth demand and cost in AI inference. Its size grows linearly with context length and model parameters, often rivaling or exceeding model weight size, thereby imposing a 'bytes-per-new-token tax' that limits batching and concurrency. This dynamic severely constrains throughput and economic efficiency, especially for large context windows reaching 128k to 256k tokens, as highlighted by d-Matrix CEO Sid Sheth and corroborated by multiple analyses. While software optimizations like FlashAttention and speculative decoding have mitigated some bandwidth pressure by reducing memory accesses and improving cache efficiency, they cannot fully overcome the fundamental physical limits imposed by memory bandwidth and capacity.

The AI hardware ecosystem is responding to these memory bandwidth bottlenecks through both architectural innovation and strategic industrial partnerships. Companies like Nvidia are investing heavily in securing and co-developing next-generation HBM memory with SK hynix, as evidenced by their multiyear collaboration to optimize HBM4 and HBM4E production and design using AI-driven manufacturing automation. Meanwhile, emerging SRAM-centric designs from d-Matrix, Groq, and Cerebras leverage on-chip SRAM to deliver ultra-high bandwidth and low latency, albeit with capacity tradeoffs. These efforts reflect a broader recognition that memory bandwidth and capacity—not just raw compute—are the gating factors for AI inference performance, with memory supply constraints expected to persist for years due to limited manufacturers and costly fabrication processes.

Sources
CNBC - TechnologyQCwireArtificial Intelligence Made SimpleVik's NewsletterChinaTalkArtificial Intelligence Made Simple

Packaging: The New Moore’s Law

TSMC’s CoWoS and 3D stacking have become the linchpin for AI progress, with industry giants pouring billions into packaging tech that now dictates who can build the fastest AI chips.

Advanced packaging technologies such as TSMC’s CoWoS have become indispensable in overcoming the physical and economic limitations of transistor scaling, enabling chipmakers like Nvidia to integrate massive chiplets—exemplified by the Blackwell (B200) with 208 billion transistors, more than doubling its predecessor’s count. This evolution extends Moore’s Law by shifting from traditional 2D scaling to complex 3D stacking and chiplet integration, dramatically boosting memory bandwidth from DDR5’s 70 GB/s to over 1,200 GB/s, which is critical for AI workloads that are bottlenecked by memory access rather than raw processor speed. The industry’s strategic commitment is underscored by an estimated $100 billion investment over the past three years, reflecting packaging’s central role in scaling AI compute beyond conventional fabrication limits.

Despite wafer fabrication capacity expansions, TSMC’s CoWoS advanced packaging remains the pivotal bottleneck in AI hardware supply, with capacity capped at roughly 20,000 wafers per month and lead times stretching up to 78 weeks. Nvidia’s dominant reservation of about 60% of CoWoS output for its GPU architectures exemplifies the fierce competition for this scarce resource, which commands pricing increases two to four times faster than wafers themselves. This scarcity has driven diversification, with alternative packaging providers such as OSATs (ASE, SPIL, Amkor) and Intel’s EMIB 2.5D packaging gaining traction, while geopolitical factors—like Nvidia’s pivot to Amkor for China-bound H200 chips following export license developments—highlight packaging’s strategic supply chain sensitivities.

Advanced packaging innovations are not only enabling the integration of heterogeneous memory types such as HBM and emerging High Bandwidth Flash (HBF) but are also driving new AI compute architectures and ecosystem expansions. SK Hynix and Micron’s hybrid H³ architecture, combining HBM and HBF, targets unprecedented memory capacity and bandwidth within compact footprints, addressing AI inference’s distinct demand for high-speed reads. However, the complexity and yield risks inherent in such heterogeneous stacking have made companies like Nvidia cautious in early adoption, shifting initial development toward cloud giants like Google and Meta, which are pushing custom ASICs to leverage these packaging breakthroughs. This dynamic underscores packaging’s role as a critical enabler and bottleneck for next-generation AI memory scalability and chip design services.

Looking ahead, the packaging landscape is evolving rapidly with TSMC’s development of next-generation CoPoS technology, slated for pilot production by mid-2027 and early adoption by Nvidia’s Feynman platform in 2028, signaling ongoing innovation to meet escalating AI demands. Concurrently, industry-wide investments are surging: ASE’s capex rose to $8.5 billion with 15 new plants planned, while memory suppliers like SK Hynix are raising billions to expand packaging capacity, though major new HBM packaging facilities are not expected before 2028. This expansion is vital as advanced packaging capacity continues to constrain AI hardware scalability, even as alternative providers and new packaging methods proliferate to alleviate supply pressures and enable breakthroughs in on-device AI and heterogeneous inference architectures.

Sources

HBM Supply: An Oligopoly’s Grip

A handful of memory giants and a single advanced packaging vendor control the fate of AI hardware, triggering bidding wars and production shifts that reshape the entire semiconductor landscape.

The supply chain for high-bandwidth memory (HBM) is tightly controlled by a triopoly of SK Hynix, Samsung, and Micron, with SK Hynix alone commanding roughly 53-62% of the market. This oligopolistic control, combined with the intensive wafer space requirements of HBM4—about three times that of standard DRAM—has led to fully booked production capacity through 2026 and beyond, driving unprecedented pricing power and forcing manufacturers to repurpose traditional DRAM lines, thereby constricting broader memory availability. Nvidia’s multi-hundred-billion-dollar strategic partnership with SK Hynix exemplifies how securing long-term HBM supply is now as critical as GPU architecture itself, underscoring the structural nature of this bottleneck in AI hardware availability and pricing dynamics.

Advanced packaging capacity, particularly TSMC’s proprietary CoWoS (chip-on-wafer-on-substrate) process, has emerged as a critical chokepoint in AI accelerator production. Despite aggressive capacity expansions aiming to nearly quadruple output by late 2026, CoWoS remains sold out through 2026 and into 2027, with Nvidia securing upwards of 60% of this scarce resource. This scarcity has triggered a private bidding war among leading AI players, elevating TSMC’s margins to historic highs and compelling companies like SK Hynix to invest billions in new U.S. packaging facilities to achieve vertical integration. Meanwhile, alternative packaging approaches such as Intel’s EMIB-T and outsourcing to OSATs like Amkor are gaining traction, reflecting the strategic importance of packaging capacity in shaping AI hardware availability and competitive dynamics.

Power infrastructure constraints are increasingly recognized as a fundamental bottleneck in scaling AI hardware deployments, with data center electricity demand projected to nearly double by 2030 and grid interconnection delays stretching up to seven years. This mismatch between rapid compute demand growth and slow power capacity expansion—on decade-long timelines—forces hyperscalers like Microsoft and Google to invest heavily in dedicated energy assets, including nuclear plant restarts and multi-billion-dollar acquisitions of renewable power firms. These developments highlight that power availability and cost-efficiency considerations are now central to AI hardware supply chains and site selection strategies, adding a complex layer to the memory and packaging bottlenecks.

The fierce competition for scarce memory and packaging resources has catalyzed a wave of strategic partnerships and long-term procurement frameworks among AI hardware players, fundamentally reshaping market dynamics. Nvidia’s $500 billion multi-year deal with SK Hynix and Samsung’s $200 billion MoU with Broadcom exemplify how companies are locking in supply chains that integrate memory, foundry, and packaging capabilities to mitigate risk amid capacity constraints and long qualification cycles. These alliances not only secure critical components but also foster industrial co-design and AI-driven manufacturing automation, reflecting a shift from transactional procurement to deeply collaborative ecosystems that influence pricing, availability, and geopolitical supply resilience.

Sources

Architectural Rethinks Break Barriers

Radical designs—like wafer-scale engines and hybrid memory stacks—are rewriting the rules, prioritizing bandwidth and memory proximity over raw compute to leapfrog traditional bottlenecks.

D-Matrix’s digital in-memory compute (DIMC) architecture exemplifies a breakthrough in overcoming AI memory bottlenecks by tightly integrating SRAM cells with MAC compute cores, achieving an extraordinary memory bandwidth of up to 150 TB/s—far surpassing traditional HBM-based chips capped around 2 TB/s. This tightly coupled design, organized into 256 cores per chiplet interconnected all-to-all, allows the Corsair platform to scale memory and compute capacity by linking multiple chiplets on an organic substrate, addressing the SRAM capacity limits for large models that require tens of gigabytes of weight storage. CEO Sid Sheth highlights that by avoiding reliance on DRAM, which faces supply constraints, their architecture sidesteps common bottlenecks, enabling up to 10x faster inference and 5x better energy efficiency compared to Nvidia GPUs for latency-sensitive AI workloads.

High Bandwidth Flash (HBF) technology, pioneered by SK Hynix and adopted by Google in TPU architectures, introduces a complementary memory paradigm that leverages 3D-stacked NAND flash to deliver 10 to 16 times the capacity of HBM at roughly one-fifth the cost. While NAND’s limited write endurance confines HBF’s use to mostly static model weights rather than dynamic KV caches or activations, this hybrid memory stacking approach—combining HBM and HBF—addresses the capacity and cost bottlenecks inherent in scaling large AI inference models. Advanced packaging techniques like silicon through vias (TSV) and hybrid bonding are critical enablers here, underscoring the growing importance of heterogeneous memory integration in AI hardware evolution, particularly as cloud giants like Google, Meta, and AWS drive early adoption rather than traditional GPU vendors.

Cerebras Systems and China’s DFSX exemplify radical architectural innovations that prioritize memory bandwidth and proximity over transistor scaling to break through AI compute bottlenecks. Cerebras’s wafer-scale engine (WSE-3) integrates 44 GB of on-chip SRAM and nearly one million AI cores on a single 5nm wafer, achieving memory bandwidth thousands of times greater than GPUs and enabling training of trillion-parameter models. Meanwhile, DFSX’s DF1000 and DF2000 chips employ 3D near-memory compute with wafer-level hybrid bonding to stack memory directly atop compute layers, achieving memory bandwidths up to 15 TB/s per chip and 960 TB/s in multi-chip supernodes—surpassing NVIDIA’s GB200 NVL72 system. By focusing on memory throughput rather than raw FLOPS, DFSX challenges the conventional wisdom of node miniaturization, demonstrating that architectural innovation and advanced packaging can overcome both memory and fabrication bottlenecks even on mature process nodes.

Looking ahead, d-Matrix is advancing its architecture with next-generation 3D memory stacking that vertically integrates multiple DRAM stacks atop compute units, aiming to multiply both memory capacity and bandwidth within a smaller footprint to accelerate token generation for latency-critical AI inference. This hybrid bonding approach, combined with heterogeneous deployment strategies that pair their accelerators with traditional GPUs, reflects a broader industry shift away from GPU-only solutions toward specialized architectures that directly address memory bandwidth bottlenecks. As Bhoja of d-Matrix notes, stacking four DRAM layers on compute enables many more fast tokens in a compact form factor, highlighting how packaging innovations are becoming as crucial as transistor-level improvements in the race to optimize AI compute performance.

Sources

Software Squeezes More from Memory

Extreme KV cache quantization, smarter decoding, and new kernel tricks are wringing out up to 5× gains in AI throughput—often outpacing hardware advances in breaking the bandwidth wall.

By early 2026, it became clear that the decode phase in AI inference is heavily memory-bandwidth bound, with arithmetic intensity as low as 1–4 FLOPs per byte compared to hardware thresholds like the NVIDIA H100’s 295 FLOPs per byte, causing GPUs to idle while waiting for KV cache data. Software and model-level strategies such as aggressive quantization (e.g., INT4), kernel optimizations like FlashAttention, and architectural innovations targeting reductions in KV cache size and bytes loaded per token have emerged as critical tools to raise arithmetic intensity and reduce data movement, thereby improving inference efficiency despite hardware memory limits.

Leading AI labs and companies including Anthropic, Google, and DeepSeek have driven substantial cost reductions and throughput improvements by innovating around KV cache compression, reuse, and management, fundamentally shifting AI economics. Techniques like the OScaR framework push extreme KV cache quantization to near-lossless levels with up to 5.3× memory footprint reductions and 4.1× throughput gains, while software-hardware co-design and kernel-level optimizations—such as those enabling FlashAttention to reach 85% of peak GPU performance—maximize compute utilization by overlapping data loading and computation to mitigate memory bottlenecks.

Recent advances in model architectures and decoding workflows, including mixture of experts (MoE), multi-token speculative decoding, and structured multi-step delegation pipelines, have further alleviated memory bandwidth constraints by reducing the bytes-per-token load and improving inference fidelity over long contexts. NVIDIA’s NVFP4 quantization format and speculative decoding methods like EAGLE-3 enable up to 5× speedups without accuracy loss, illustrating how algorithmic and software innovations continue to outpace hardware improvements, which currently account for only about one-quarter to one-third of efficiency gains.

Despite these strides, the fundamental quadratic growth of the KV cache remains a persistent challenge, especially for emerging use cases like autonomous agents with long, mutable histories that resist caching. This has spurred exploration of unconventional hardware-software co-design approaches, such as dedicated decode machines using older GPUs and SRAM accelerators, alongside advanced kernel-level optimizations that repurpose underutilized compute units and employ software emulation to overlap operations, aiming to unlock remaining optimization headroom and sustain cost-efficient AI inference at scale.

Sources
Artificial Intelligence Made SimpleAI for Software EngineersArtificial Intelligence Made SimpleData GravityHugging Face Daily PapersAI Engineer

Supply Chains Shape AI’s Limits

Multi-year shortages in memory, packaging, and even power infrastructure are forcing hyperscalers into direct supplier deals and record investments, setting hard ceilings on AI’s growth curve.

By 2026, the AI infrastructure landscape is grappling with entrenched bottlenecks in High Bandwidth Memory (HBM) supply and advanced packaging capacity, prompting hyperscalers like Microsoft, Google, and Meta to engage directly with suppliers such as Micron, Samsung, and SK Hynix in Korea to secure scarce resources. This scarcity is exacerbated by the wafer-intensive nature of emerging memory tiers like HBM4 and HBM4e, which demand up to three times more wafer space than traditional DRAM, forcing manufacturers to reallocate production lines and further constrict traditional memory availability. Concurrently, the shift toward chiplet-based AI accelerators from Nvidia, AMD, Google, and Amazon intensifies pressure on advanced packaging technologies, particularly TSMC's CoWoS, which is sold out through 2026 and heavily allocated, underscoring the critical role of packaging as the new frontier in overcoming memory bandwidth bottlenecks.

The industry’s response to these multi-layered constraints involves massive capital investments and strategic long-term partnerships aimed at expanding and diversifying supply chains across memory, packaging, and optical components. Notably, Nvidia’s $2 billion investment in laser suppliers Lumentum and Coherent highlights the growing recognition of optical interconnects as vital to scaling AI compute clusters, while semiconductor giants like TSMC and Amkor have inked decade-long agreements to secure advanced packaging capacity. Meanwhile, emerging technologies such as Intel’s EMIB-T and High Bandwidth Flash (HBF) memory stacks—offering up to 10 times the capacity of HBM per stack—are gaining traction as strategic alternatives to alleviate bottlenecks and enable new deployment models, including on-premises air-cooled AI systems that reduce reliance on costly GPU proliferation.

Despite aggressive expansion efforts, supply chain constraints are projected to persist through 2027 due to the inherently long lead times—often two to three years—for building new fabs and packaging facilities, compounded by cascading shortages in materials like copper, indium phosphide, and power delivery components. This has led to an unprecedented scale of capital deployment, with companies like ASE raising 2026 capex to $8.5 billion and hyperscalers driving global data center projects toward gigawatt-scale compute deployments. Industry leaders, including Nvidia’s Jensen Huang, warn that these bottlenecks, particularly in HBM and power infrastructure, will cap AI revenue growth to roughly doubling annually, emphasizing that overcoming these multi-front constraints requires not only hardware innovation but also systemic changes in budget structures and resource pooling to unlock scalable AI performance.

Looking ahead, the AI infrastructure ecosystem is evolving from a focus on raw compute power toward memory-centric architectures and integrated system designs that balance compute and memory bandwidth. Companies like Rebellions AI are pioneering this shift by combining large on-chip SRAM with HBM3 to maximize energy efficiency, while collaborative efforts across semiconductor manufacturers, software developers, and telecom operators aim to accelerate deployment of 3D chip architectures and advanced packaging technologies. This global, partnership-driven approach not only addresses supply chain resilience and AI sovereignty concerns—as seen in South Korea’s sovereign AI initiatives and France-Taiwan manufacturing collaborations—but also fosters innovation in rack-scale systems optimized for cost and energy efficiency, signaling a maturation of the AI infrastructure landscape beyond mere chip speed to holistic system performance.

Sources

Get the stories behind the trends

Deep-dive reporting and the weekly brief, in your inbox.