AI's pay-per-token era: SaaS abandons subscriptions as cost pressures mount

a16z

The gist

As AI compute costs skyrocket, SaaS giants are ditching subscriptions for pay-per-token pricing—forcing enterprises to rethink everything from budgets to vendor strategies.

What to know

The End of Flat-Rate AI

Granular, pay-as-you-go billing is replacing all-you-can-eat subscriptions in AI, forcing enterprises to rethink spending, diversify providers, and demand smarter, value-based pricing.

The AI and SaaS industries are decisively moving away from traditional flat-rate subscription and seat-based pricing models toward granular, consumption-based billing that more accurately reflects the highly variable and escalating costs of frontier AI workloads. Anthropic’s recent decision to end flat-rate subscriptions for Claude Enterprise, requiring customers to switch to pay-as-you-go API billing, exemplifies this shift, underscoring the unsustainability of 'all-you-can-eat' pricing in the face of soaring compute demands. This transition mirrors historical industry evolutions, such as the move from perpetual licenses to recurring subscriptions and from CapEx to OpEx cloud spending, signaling a broader recalibration of how enterprises forecast and manage technology expenses.

Enterprises are increasingly diversifying their AI provider portfolios to avoid dependency on single vendors’ pricing whims, leveraging multiple models and open-weight alternatives to gain economic resilience and flexibility. This multi-provider strategy aligns with the consumption-based pricing paradigm, enabling organizations to optimize costs dynamically and prevent overpaying for unused capacity or hitting restrictive usage caps. As Anthropic matures from a startup subsidizing compute costs to a major enterprise software vendor, its pricing evolution toward consumption-based models reflects a market-wide push for fairness and predictability, where customers pay precisely for the value and volume of AI usage they consume.

While consumption-based pricing offers fairness and scalability, it introduces complexity that providers like Vercel have addressed by dedicating multiple teams to develop sophisticated metering and billing systems capable of tracking AI usage at token or CPU cycle granularity. Customers desire predictability amid fluctuating workloads—such as seasonal spikes on Black Friday—forcing providers to balance flexible pay-per-use models with cost transparency and simplicity to reduce friction. Hybrid approaches combining tiered plans with scalable usage metrics and value-based pricing linked directly to delivered AI outcomes are emerging as pragmatic solutions to manage this complexity and align pricing with customer-perceived value.

Despite some lingering reliance on seat-based subscriptions—illustrated by OpenAI’s recent $100 per user per month offering—industry consensus suggests these models are temporary stopgaps unable to sustainably manage the ballooning compute costs of AI workloads. OpenAI’s subscription attempt appears partly aimed at capturing Anthropic’s enterprise customers amid pricing anxieties, but the inevitable trajectory points toward consumption-based billing as the standard, driven by the need for cost efficiency, fairness, and alignment with actual AI usage. This evolution is poised to reshape SaaS economics fundamentally, rendering traditional ARR metrics obsolete in favor of pay-per-token or pay-per-inference models that better capture AI’s dynamic value delivery.

Sources
The InformationMind the ProductThe Information's TITVThe a16z ShowMore or Less PodcastWhat's Hot 🔥 in Enterprise IT/VC

FinOps Tools Reshape AI Budgets

Real-time cost tracking and the rise of distilled, task-specific AI models are slashing enterprise spending, shifting the market toward efficiency and clear ROI over unchecked growth.

By early 2026, innovations in AI FinOps tools have become indispensable for enterprises grappling with unpredictable and variable AI spending. Vendors like Vantage offer platforms that break down costs per API call, workflow, or tenant, enabling providers to align pricing with actual usage and safeguard margins as AI adoption scales. Meanwhile, managed service providers such as Rev.io leverage AI-driven internal dashboards to monitor ROI and token utilization, empowering teams to justify AI expenses through clear value demonstration.

The emergence of distilled AI models like OpenAI's GPT 5.4 Mini and Nano exemplifies a strategic shift toward cost-efficient AI adoption, particularly for handling simpler tasks at a fraction of the cost. These smaller, specialized sub-agents reduce expensive token consumption by delegating routine functions—such as document summarization or categorization—to cheaper models, a tactic that can dramatically lower daily expenses for small businesses from hundreds to mere tens of dollars, thereby enhancing ROI and sustainability.

Despite current subsidies keeping API token prices relatively low, the looming rise in AI infrastructure and R&D costs underscores the necessity for granular cost tracking and spending guardrails. As companies like Gamma demonstrate, selecting AI models that balance performance, speed, and cost—sometimes favoring longer-tail, less powerful models—can achieve profitability within months. This nuanced approach to AI FinOps reflects a broader industry trend toward commoditization of AI models, where rapidly declining costs will further shape future spending management strategies.

AI FinOps strategies must also reconcile the tension between delivering premium product experiences and managing escalating costs. For example, RevenueCat opts for the highest-end models to ensure superior output quality and user satisfaction, accepting higher initial expenses with the expectation that AI prices will decline over time. This balancing act highlights how cost management tools and pricing models are evolving to align AI output quality, speed, and affordability with competitive market positioning.

Sources
Authority Hacker PodcastSub Club by RevenueCatChannelholic

Margin Squeeze Hits AI Giants

Skyrocketing infrastructure costs and investor skepticism are driving tech titans and enterprise adopters to prioritize profitability, disciplined scaling, and third-party AI management over reckless expansion.

Enterprises scaling AI infrastructure grapple with soaring compute costs and margin compression, forcing a delicate balance between aggressive growth and sustainable cost management. Despite the high expenses and uncertain immediate ROI, Big Tech giants like Google have doubled their share prices following Gemini and TPU launches, underscoring the strategic imperative to lead the AI race even as chatbots face commoditization and persistent losses, as highlighted by the ongoing 'race to the bottom' in pricing.

Market realities and investor sentiment sharply constrain AI infrastructure expansion, with Microsoft and Oracle experiencing stock declines of 30% and 60% respectively since late 2025, reflecting financial discipline imposed by wary capital markets. OpenAI’s recent difficult funding round and pivot toward enterprise monetization further illustrate the pressing need to convert AI compute investments into sustainable revenue streams amid high cash burn and margin pressures.

Operationally, scaling AI infrastructure challenges many enterprises, particularly in sectors like lending where cloud storage costs have surged unpredictably and in-house expertise remains scarce. Consequently, many lenders emulate their cloud adoption journey by relying on third-party providers to manage AI scalability, while adopting a 'start small' approach by integrating off-the-shelf AI tools into workflows to mitigate upfront infrastructure risks and costs.

The financial implications of AI infrastructure scaling extend beyond raw costs to include strategic pricing and margin management, as SaaS companies confront a structural shift from historically high margins toward retail-like profitability due to token commoditization and deflationary AI economics. Enterprises face margin compression exacerbated by AI price wars where new entrants subsidize token costs, compelling incumbents to match cuts and triggering cascading price reductions, making sustainable growth contingent on intelligent workload routing, multi-tool deployment for risk mitigation, and a focus on value-driven differentiation rather than pure cost-cutting.

Sources
PhiloinvestorSiliconANGLE theCUBEMind the ProductSiliconANGLE theCUBECode Story: Insights from Startup Tech LeadersThe Few Bets That Matter

Get the stories behind the trends

Deep-dive reporting and the weekly brief, in your inbox.