AI price war heats up as enterprises demand 90% cut

The gist
Enterprises are demanding a 90% cut to AI token prices, igniting an industry-wide price war as tech giants and vendors scramble to make large-scale AI adoption economically viable.
What to know
- Palo Alto Networks’ CEO Nikesh Arora is calling for a phased 90% AI token price drop by 2028, starting with a 20% cut in year one to unlock enterprise AI at scale.
- Meta slashed Muse Spark 1.1 token prices to $4.25 per million, forcing rivals like OpenAI and Anthropic to rethink pricing amid skyrocketing enterprise demand and backlash.
- Innovations like Stanford’s RouteLLM are enabling up to 85% cost savings with near-GPT-4 quality, pushing enterprises to adopt smarter routing and infrastructure to survive the price squeeze.
Enterprise AI Budgets Under Siege
Major companies are slashing AI investments and overhauling procurement strategies as runaway token costs force hard choices between innovation and operational viability.
Nikesh Arora, CEO of Palo Alto Networks, has issued a clarion call for a dramatic 90% reduction in AI token pricing within the next two years, underscoring that without such a steep price drop, scalable and sustainable enterprise AI adoption remains out of reach. Arora specifically outlines a phased approach: a 20% cut in the first year followed by an additional 70% the next, framing this as essential for companies to realistically scale AI workloads and achieve viable ROI.
The urgency of this pricing overhaul is already palpable among major enterprises like Uber, which exhausted its entire 2026 AI budget by April, forcing CTO Praveen Neppalli Naga to reconsider AI investments versus traditional engineering hires. This budget strain exemplifies how soaring token costs are not just theoretical concerns but immediate barriers that compel companies to ration AI usage and rethink vendor partnerships based on price per token.
In response to escalating token expenses, enterprises are fundamentally reshaping their AI procurement and governance strategies by embedding cost control as a core design principle. This evolution includes instituting tiered AI model usage—reserving frontier models for mission-critical tasks while routing less sensitive workloads to older or open-source alternatives—and integrating price-per-token metrics into vendor evaluation scorecards, thereby elevating cost considerations alongside traditional capability benchmarks.
AI Usage Caps Reshape Workflows
Strict daily token limits and intelligent routing are redefining enterprise AI, shifting focus from experimentation to maximizing ROI through disciplined model selection.
Enterprises such as Uber, Meta Platforms, and Salesforce have shifted from unchecked AI experimentation to rigorous cost governance by rationing AI token usage and instituting daily caps, like PMG’s $50-per-user limit, to prevent budget overruns during peak demand periods. This strategic pivot prioritizes business value over sheer volume, reflecting a broader industry trend where firms enforce usage limits and engage in ongoing education to align AI consumption with measurable outcomes, as exemplified by Dollar Shave Club’s triage system to reserve costly models for critical tasks.
To optimize token spend, enterprises are increasingly adopting sophisticated AI model routing strategies that dynamically allocate requests to the most cost-effective models based on task complexity and organizational roles. Companies like MindBridge and others implement routing layers that act as abstraction layers, enabling CIOs to direct tokens toward specialized models—such as Gemini flash for dashboard coding or fine-tuned models like Cobalt for specific codebases—thereby avoiding the expensive default use of top-tier models for simple queries and balancing accuracy with cost efficiency.
Cost governance in enterprise AI is evolving into a software-driven discipline where decision-making balances accuracy and token cost through iterative workflow optimization embedded within enterprise systems. As Monti Saroya highlights, automated systems must discern the minimal intelligence level required to maximize data accuracy while minimizing expenses, a necessity given the impracticality of manual model selection. This approach is complemented by strategic partnerships with specialized hardware providers like SambaNova and Intel to mitigate rising compute costs exacerbated by supply constraints and geopolitical factors.
Effective AI cost management also hinges on cultural and operational adaptations, including educating employees about the high costs of cutting-edge AI tools and encouraging the use of less expensive models when appropriate. Enterprises like those led by Matthias Steinberg advocate allocating dedicated budgets for AI experimentation to foster internal learning and transparency, while grappling with the complexity of multiple AI tooling sources that necessitate standardizing usage patterns to streamline development workflows and optimize overall token consumption.
Meta Ignites AI Price Showdown
Meta’s aggressive price cuts are forcing competitors to rethink their business models as enterprises flock to commercial providers for predictable, scalable AI solutions.
Meta Platforms has dramatically intensified the AI pricing wars with its Muse Spark 1.1 API, slashing token prices to an aggressive $4.25 per million tokens. This bold move directly pressures competitors like OpenAI and Anthropic to reconsider their pricing strategies amid growing enterprise backlash against exorbitant token fees. By aggressively undercutting prices, Meta is reshaping the battle for cost-efficient superintelligence and forcing a market-wide reevaluation of value versus cost in AI service offerings.
Anthropic’s adoption of a dual-pricing AI model strategy exemplifies the fierce competition among leading AI vendors grappling with soaring infrastructure costs and a volatile funding environment in 2026. This nuanced pricing approach reflects the broader market dynamics where providers must balance profitability with the urgent demand for scalable, affordable AI solutions. Meanwhile, enterprises are increasingly scrutinizing token costs as a critical factor in vendor selection, underscoring how pricing models now directly influence market share battles.
The escalating token costs are not only fueling intense price competition but also accelerating a significant shift in enterprise AI adoption patterns. Despite the availability of open source models, which have seen production usage drop from 19% to 11%, 44% of enterprises now rely on commercial providers like Anthropic, highlighting a consolidation around major vendors. This trend underscores how rising costs and the complexity of managing financing and technical stacks are driving enterprises toward established players who can offer more predictable and scalable pricing.
Smart Infrastructure Trumps Model Hype
Enterprises are slashing costs by routing tasks across open-source and specialized models, with infrastructure intelligence and private deployments outperforming brute-force model power.
The escalating complexity of AI workflows, such as multi-agent pipelines and retrieval-augmented generation, has significantly increased computational overhead per task, driving up token consumption and costs. This challenge is compounded by many enterprises initially adopting AI tools without robust ROI measurement frameworks, leading to unchecked token usage and a strategic imperative to optimize cost governance as AI scales.
Intelligent routing of AI requests across multiple models, providers, and geographies has emerged as a transformative innovation for cost optimization. Stanford’s RouteLLM research demonstrated that such routing can achieve up to 85% cost savings while preserving 95% of GPT-4 quality, highlighting a shift where infrastructure intelligence—exemplified by Telnyx’s private global IP backbone and GPU clusters—now trumps raw model sophistication in driving enterprise AI efficiency.
Open-weight and open-source AI models are reshaping market dynamics by enabling enterprises to deploy AI on private infrastructure, thereby reducing vendor lock-in, enhancing auditability, and slashing token costs. For instance, Vercel’s June data showed open-weight models handling 29% of token volume but accounting for less than 4% of spending, while innovations like GLM-5.2’s reinforcement learning and inference optimizations narrow the performance gap with proprietary frontier models, fueling a bifurcated market of high-cost frontier and low-cost high-volume models.
The AI market’s current reliance on heavy subsidies from venture capital and cloud providers obscures the true cost of token consumption, a situation poised to change as AI companies transition to public markets demanding profitability. This impending subsidy withdrawal is accelerating investments in cost-efficient infrastructure and open-weight models, as seen in Anthropic’s $80 billion cloud commitments and strategic partnerships with SpaceX and Broadcom, signaling a broader industry pivot from pure model provision to platform orchestration for sustainable margin and scalable enterprise adoption.








