Amazon and Datadog Put Journey-Level Observability on QA’s Radar

Amazon and Datadog are nudging QA toward journey-level scoring, where live user and agent flows help teams prioritize the tests that matter most.

Updated

What is this trend?

Journey-level observability is pushing QA to score real user and agent flows by tracing logs, metrics, and outcomes so teams can focus testing on the journeys most likely to fail.

  • Amazon is tying traces, logs, metrics, and evaluators to live agent quality scoring.
  • Datadog is inferring critical user journeys from real traffic and surfacing high-risk paths.
  • QA focus is shifting from broad coverage to login, onboarding, checkout, and other revenue flows.
  • Observability is becoming a decision aid for prioritizing tests, not just diagnosing incidents.
  • The signal is early, but live journey data is now shaping where QA scrutiny lands.

What’s the latest?

Amazon’s “Amazon Unifies AI Agent Observability and Evaluation” is the clearest signal this week: it combines distributed tracing, automated quality scoring, curated datasets, and experiment workflows

How it developed

  1. Release readiness scoring, synthetic test data, and risk-based QA reshape go/no-go decisions

Go deeper

Curated long-form picks on this trend — podcasts, videos, and analysis, by seniority.

Stay ahead in Quality Assurance / Quality Control

Get the weekly Quality Assurance / Quality Control brief in your inbox — the developments, what they mean by seniority, and what to do next.