Cerebras rockets past Nvidia in AI speed, but margin squeeze and cloud bets rattle investors

How They Make Money

The gist

Cerebras has leapfrogged Nvidia in AI inference speed with its monster wafer-scale chip—but margin pains, cloud gambits, and sky-high expectations are giving investors whiplash.

What to know

The Wafer-Scale Gamble

Cerebras’ lightning-fast hardware outpaces Nvidia, but faces an uphill battle against entrenched CUDA software and costly engineering hurdles.

Cerebras has carved out a formidable technological edge with its wafer-scale AI processor architecture, boasting a speed advantage ranging from 15 to 21 times over Nvidia GPUs for inference workloads. This leap is driven by its third-generation Wafer-Scale Engine (WSE-3), which integrates 900,000 AI-optimized cores and 44GB of on-chip SRAM on a single wafer, drastically reducing chip-to-chip communication bottlenecks that typically slow down multi-GPU clusters. CEO Andrew Feldman encapsulates this innovation as a 'giant wafer scale chip' designed specifically for inference speed, enabling faster token generation and lower latency critical for interactive AI applications.

Despite Cerebras’ raw hardware speed superiority, the company faces a steep uphill battle against Nvidia’s entrenched software ecosystem, particularly the CUDA platform, which dominates AI training and inference workloads. Nvidia’s modular GPU lineup, including Blackwell and Hopper architectures, enjoys broad native support across all major large language model frameworks and enterprise stacks, creating a powerful developer lock-in that Cerebras must overcome through costly custom engineering and specialized compilation. This software moat not only complicates Cerebras’ market penetration but also underpins Nvidia’s impressive 75% gross margins, contrasting sharply with Cerebras’ negative operating margin guidance despite its landmark $20B+ deal with OpenAI.

To counter Nvidia’s software dominance, Cerebras is strategically expanding beyond traditional enterprise sales by integrating its wafer-scale technology into cloud marketplaces such as AWS Marketplace, Microsoft Marketplace, and IBM watsonx Model Gateway. This approach aims to democratize access to its high-speed inference capabilities via self-service APIs and broader developer engagement, potentially loosening Nvidia’s software lock-in. By aligning its WSE-3 powered CS-3 systems with interactive AI use cases that prioritize low-latency token generation, Cerebras bets on a classic infrastructure play: building foundational hardware first and waiting for the AI ecosystem to catch up.

Sources

Cloud Deals Redefine Growth

Multi-billion dollar partnerships with OpenAI and AWS are shifting Cerebras’ business from bespoke hardware to recurring cloud revenue—reshaping its entire commercial strategy.

Cerebras has cemented its commercial leadership in AI infrastructure through landmark multi-year agreements with OpenAI and Amazon Web Services, collectively valued at over $20 billion. The centerpiece is the OpenAI deal, which commits to 750 megawatts of high-speed inference compute capacity over several years, enabling OpenAI to port GPT 5.4 and 5.5 models onto Cerebras’ wafer-scale chip architecture. This deal not only validates Cerebras’ speed advantage in AI inference workloads but also marks a significant revenue milestone as these deployments ramp up.

Strategically expanding beyond direct hardware sales, Cerebras is aggressively integrating into cloud marketplaces to broaden its commercial footprint. Collaborations with AWS leverage a disaggregated inference model where AWS’s Trainium chips handle prefill stages and Cerebras’s CS-3 systems execute high-speed decode, enabling scalable and efficient AI inference. Moreover, Cerebras solutions are now accessible through multiple cloud marketplaces including Microsoft, IBM watsonx, Vercel AI Gateway, OpenRouter, and Hugging Face, targeting startups and AI-native companies and accelerating recurring revenue streams.

This strategic pivot towards cloud-based AI services is reflected in Cerebras’ financial performance, with core revenues soaring 92% year-over-year to $191.3 million in Q1 2026, driven by a remarkable 167% jump in cloud and services revenues to nearly $80 million. This shift towards renting back data center capacity and emphasizing inference consumption models signals a deliberate move to recurring, service-oriented revenue, even as the company navigates margin pressures inherent in this transition.

Sources

Sky-High Hopes, Shaky Margins

Despite blockbuster contracts and dizzying valuation multiples, Cerebras’ reliance on a handful of customers and eroding margins make its future a high-stakes balancing act.

Investor sentiment toward Cerebras post-IPO reflects a complex blend of bullish optimism and cautious skepticism. While analysts from UBS, Morgan Stanley, and Citi emphasize the company's wafer-scale engine and multi-year contracts as strong growth drivers, the stock's high valuation multiples—trading around 64 to 90 times forward revenue compared to Nvidia’s 12 times—and significant price volatility, with annualized volatility exceeding 244%, temper enthusiasm. This elevated valuation demands near-perfect execution, as any delays in OpenAI deployments or margin pressures could sharply compress multiples, underscoring the market’s heightened expectations for AI companies today.

Cerebras faces pronounced revenue concentration risks that continue to unsettle investors despite its impressive $20 billion-plus backlog. Historically reliant on two major UAE clients accounting for 86% of 2025 sales, the company has shifted this dependency toward marquee partners like OpenAI and AWS through landmark deals. While these relationships underpin optimism, analysts and investors remain wary that any disruption or slowdown in these key accounts could destabilize financial performance and market confidence, especially given the gradual slowdown in revenue growth guidance and the critical need to convert backlog into recognized revenue swiftly.

Margin pressures present a persistent challenge for Cerebras, with gross margins declining from approximately 42% in 2024 to around 39% in 2025, and forecasts dipping further to 36%-41% in 2026—well below industry leaders like Nvidia’s 74.9%. This erosion stems largely from the complexity and defect-prone nature of wafer-scale chips and operational factors such as renting back its own chip systems to fulfill backlog. Although CEO Andrew Feldman anticipates margin improvements as new data center capacity comes online, these margin headwinds have triggered stock sell-offs and fuel investor caution, highlighting the tightrope Cerebras must walk to sustain its lofty valuation.

Stock volatility remains a defining feature of Cerebras’ post-IPO journey, with the share price surging 68% on debut only to tumble over 39% from its peak within weeks, serving as a cautionary tale for investors in high-profile tech IPOs. Factors such as the impending expiration of insider lock-up agreements—potentially unlocking stakes worth billions held by CEO Andrew Feldman and CTO Sean Lie—add supply-side pressure that could further unsettle the market. Despite this turbulence, all ten analysts currently covering the stock maintain buy ratings with an average price target near $294, reflecting a cautiously optimistic consensus that hinges on the company’s ability to scale operations and diversify its customer base.

Sources

Scaling Pains Hit Hard

Margin erosion and data center bottlenecks threaten Cerebras’ cloud ambitions, as surging demand collides with infrastructure delays and capital strain.

Cerebras is grappling with significant margin contractions, primarily due to its strategy of renting back previously sold gear to customers in order to meet surging demand, which has slashed margins by 10 to 15 points. This approach, combined with heavy infrastructure spending to support its AI cloud transition, has pushed gross margin guidance down from 47% in Q1 to an expected 36-41% range for the year, reflecting the steep costs of scaling its wafer-scale chip technology amid operational complexity.

Data center capacity constraints represent the largest bottleneck for Cerebras’ expansion, as the pace of AI innovation clashes with the slow-moving realities of real estate and construction. CEO Andrew Feldman highlighted the irony that while AI technology races ahead, the company is limited by physical infrastructure hurdles such as power availability, permits, and supply chain shortages, with nearly half of the planned 16 gigawatts of US data center capacity delayed or canceled, forcing Cerebras to aggressively pursue global partnerships like the 120-megawatt deal with Bell Canada slated for 2027.

The strategic pivot toward an AI cloud model, heavily reliant on inference workloads for marquee customers like OpenAI, is a double-edged sword for Cerebras. While cloud and services revenue surged 178% year-over-year to $82.8 million, this transition introduces heightened operational risks including infrastructure bottlenecks, capital intensity, and customer concentration, with the company’s profitability and balance sheet discipline under pressure as it bets on converting blockbuster deals into sustainable growth amidst a capital-intensive buildout.

Sources

Part of these trends

Get the stories behind the trends

Deep-dive reporting and the weekly brief, in your inbox.