AI agents close the loop on IT ops with live telemetry
The gist
AI agents from AWS, Microsoft, and Splunk are closing the loop on IT operations by tying incident response and software delivery directly to live production telemetry.
What to know
- From August 27 to mid-September 2026, AWS, Microsoft, and Splunk all launched integrated AI agent stacks that automate detection, investigation, and remediation using real-time production data.
- Microsoft cut its monthly human-handled incident tickets from 70,000 to about 27,000, while Splunk rolled out a hybrid-ready agent stack at Splunk.conf26.
- Databricks’ AI SRE now supports 150+ teams, handles over 2,000 daily investigations, and dramatically reduces engineer debugging time—showcasing the shift to proactive, telemetry-driven IT ops.
Hybrid AI Agents Go Live
AWS, Microsoft, and Splunk are converging on real-time, production-scale AI agents that unify code, telemetry, and enterprise context—pushing automation beyond the cloud into tightly governed, hybrid environments.
The late-Q3 wave became unmistakable on August 27, when AWS and Dynatrace publicly presented Kiro, AWS DevOps Agent, and Bluebox as a single integrated workflow linking agentic development and operations to live runtime context. AWS said the stack spans “AWS, multicloud, and on-premises environments,” while insisting that “to close the loop between code generation and production context, Kiro and AWS DevOps Agent rely on real-time production intelligence,” a launch framing that squarely joined agentic AI, telemetry, and hybrid observability.
By mid-September, that pattern had broadened from vendor showcase to enterprise and platform rollout: Microsoft described AI agents embedded in its internal incident management tooling, saying, “The monthly volume of tickets that require human intervention rose as high as 70,000 in January 2025. Eighteen months later, we’d reduced it to less than half—about 27,000—through the application of automation and AI.” Splunk used its September 14, 2026.conf26 event to announce an integrated agent stack spanning on-prem AI, SOC automation, and federated data access—positioning agents to operate with enterprise-controlled context.
That Splunk launch mattered because it showed cloud-era agent automation being adapted for tightly governed, hybrid estates rather than only greenfield environments: “The stage at Splunk.conf26 in Denver, September 14, 2026.” Splunk updates include: “AI in your own environment: Cisco AI POD for Splunk and self-managed AI Assistant are available now, including for air-gapped deployments,” alongside federated search expansion to AWS CloudWatch Lake and Databricks, evidence that late-Q3 releases were converging on integrated, production-scale agent operations across distributed enterprise data footprints.
Telemetry Powers Proactive Ops
Agentic AI now closes the loop by consuming live production signals, enabling multi-agent workflows to autonomously detect, investigate, and remediate incidents—turning IT from reactive firefighting to prevention-first operations.
What turns agentic AI from a chatbot into an operations system is not the model alone but the telemetry-fed loop around it: continuous production signals, service relationships, deployment history, and organizational context assembled into a machine-readable graph that lets agents sense, reason, act, and adapt. In that setup, the job shifts from waiting for flames to spread to spotting smoke early, with multi-agent workflows that can detect, investigate, remediate, and then surface what happened back to teams as part of a prevention-first operating model.
The practical breakthrough is that agents no longer start from a blank prompt; they start from filtered evidence, correlating logs, metrics, traces, and recent changes, and asking the data why until they reach an actionable mitigation step. That matters because “This context assembly can consume 60-80% of an engineer's time,” while “If you reduce noise by 90%, fewer incidents enter your IT Service Management (ITSM) tool, like ServiceNow.” Once telemetry is fused across hybrid environments, the loop can extend into governed remediation: agents query production monitoring systems directly, compare current behavior with historical baselines, generate mitigation plans from prior successful actions, and hand off evidence-backed fixes for review when confidence is high. The result is not a lab demo but an operating pattern at scale—“AI SRE currently supports over 150 teams at Databricks, handling more than 2,000 investigations daily and saving engineers hours of debugging time”—showing how autonomous investigation and code remediation can move IT from reactive response toward proactive prevention with humans still in control.

