Persistent Workspaces, HBM Bottlenecks, and Cloud Cost Controls

By DripPublished

The gist

Machine learning is shifting from model access to control points: secure execution, scarce memory supply, and embedded cost governance now determine who captures value.

This week’s developments

Persistent Workspaces Move Into the Agent Security Stack

August 11 and September 8 marked the next step: xAI began beta testing Grok Bot as a per-user cloud computer, and Meta launched Muse Secure VM with the same persistent-workspace model. NVIDIA is formalizing a secure agent workspace around one persistent VM per user, while Bromure is selling persistent Linux VMs for coding agents. The common thread is that agents no longer reset after each task; they retain filesystem state, browser sessions, terminal history, and often credentials, making the runtime boundary itself the security product.

That pushes the story beyond orchestration into infrastructure control. If agents live in durable environments, policy enforcement, registries, tool permissions, and isolation become core platform decisions, not workflow add-ons. Salesforce’s Koa makes that explicit: it is a CRM reasoning model built on NVIDIA Nemotron 3 Super, post-trained on synthetic data from nearly three decades of Salesforce CRM knowledge, and deployed inside Salesforce’s trust boundary through Agentforce and Data Cloud. Salesforce says the model weights stay under its control and inference runs on Salesforce-controlled infrastructure. Baidu’s Dazi points to the same architecture for sovereign deployments. For practitioners, the progression is clear: the workspace, control plane, and vertical runtime layer are now where governed execution gets priced and operationalized.

Where will control and value accrue in persistent agent workspaces?

If you operate in this industry

  • Persistent agent workspaces make runtime control part of the product.
  • If your agents keep state, security, policy, and tool access become core architecture choices—not add-ons you can defer.

Sources

If you sell into this industry

  • Governed persistent VMs are becoming the new enterprise agent baseline.
  • Shift roadmap and GTM toward secure workspaces, permissions, and auditability; buyers will pay for control, not just orchestration.

Sources

If you invest in this industry

  • Value is moving from agent apps to the secure workspace layer.
  • Watch platform owners and infra control planes; point tools without durable runtime control risk margin compression and bundling.

Sources

Meta’s MTIA Push Exposes HBM as the New Bottleneck

Broadcom’s roughly $350 billion in AI chip orders, despite supply risk, shows demand has not softened; it has been pre-allocated. The constraint has now narrowed further than regional capacity, packaging, or power: HBM is the gating input that determines which systems get built and for whom. Meta’s response is the clearest strategic signal. The company is scaling “hundreds of thousands” of MTIA inference chips, expanding clusters and networking, and reducing exposure to scarce Nvidia supply where workloads allow. That makes this week less about whether compute is available in the abstract and more about which operators can actually secure memory-backed systems at scale. Capacity planning is shifting from GPU reservations to memory-backed systems, and the winners will be vendors that can secure HBM, redesign around lower allocations, or replace general-purpose accelerators with custom silicon and tighter stack control.

How do we secure HBM supply before competitors do?

If you operate in this industry

  • HBM, not GPUs, is now the real constraint on scaling AI.
  • Secure memory-backed capacity early, or redesign workloads around custom silicon and lower-HBM systems before rivals lock supply.

Sources

If you sell into this industry

  • HBM access now decides whose AI systems can actually ship.
  • Shift roadmap and GTM toward memory-efficient designs, custom silicon, and supply-backed offers; generic accelerator pitches are weaker.

Sources

If you invest in this industry

  • AI demand is intact; value is shifting to memory-controlled supply.
  • Favor HBM, packaging, and vertically integrated silicon plays; scarcity now determines who can convert demand into revenue.

Sources

AWS and Alibaba Push Cost Controls Into the Stack

OpenAI’s July 30 price cuts kept the pressure on, with GPT-5.6 Luna falling 80% from $1/$6 to $0.20/$1.20 per million input/output tokens, while GPT-5.6 Terra dropped 20% from $2.50/$15 to $2/$12. But this week’s more important shift was that the savings started showing up inside the stack itself: Alibaba cut Qwen3.8 audio input costs, and AWS added Kimi K3 to Bedrock with explicit cost controls, not just another model endpoint. AWS also exposed prompt caching at $0.30/M cached reads and $3.75/M cached writes versus $3.00/M uncached input, plus a cross-Region inference profile it says is about 10% cheaper than geographic routing. That extends the workflow land grab into a more mature phase, where vendors compete on total workload cost rather than discounted agent access. Orchestration platforms are reinforcing the shift with task-, action-, and outcome-based billing, absorbing falling inference costs into more predictable software economics. As text, audio, and video all compress together, standalone model APIs lose pricing power, and the next advantage goes to platforms that control routing, caching, packaging, and workflow ownership.

Where will cost-control value accrue across the AI stack?

If you operate in this industry

  • Cost controls are moving into the platform, not just the model.
  • Own routing, caching, and workflow economics or get squeezed as model margins commoditize and platform bundles win on total workload cost.

Sources

If you sell into this industry

  • Buyers now want cheaper workloads, not just cheaper tokens.
  • Shift roadmap to caching, routing, and usage controls; pricing power is moving to platforms that can prove end-to-end savings.

Sources

If you invest in this industry

  • Value is shifting from model APIs to workload-control platforms.
  • Favor vendors with routing, caching, and orchestration leverage; standalone model exposure looks increasingly commoditized.

Sources

Stay ahead in Machine Learning

Get the weekly Machine Learning brief in your inbox — the developments, what they mean by vantage, and what to do next.