AWS CloudWatch Omni Targets Agentic AI Observability With Built-In Evaluation Engine
AWS's CloudWatch Omni is generally available, adding evaluation tools and unified traces to explain why AI agents fail.
An agentic system can return a clean response, meet its latency target and throw no errors while still giving a customer the wrong answer, calling the wrong tool or pulling from a stale knowledge base, the report notes. By every traditional metric the system is healthy, even though the business outcome is a failure. AWS cites an IDC forecast of more than 1 billion deployed agents by 2029, a volume SiliconANGLE says no operations team can review manually.
The centerpiece of Omni is an evaluation engine rather than a set of dashboards. According to the report, Omni captures every trace and ships with 17 built-in evaluators that score coherence, helpfulness, faithfulness and routing correctness, among other metrics. Teams can compare prompt versions in a playground, build test datasets from production traffic and catch regressions automatically. Evaluators can also run continuously against live traffic, so quality drift is flagged in the same way a CPU spike would be.
Sony is an early adopter. "At Sony, our enterprise-wide agentic AI platform now supports hundreds of proof-of-concept and production workloads," said Masahiro Oba, senior general manager of the AI Acceleration Division at Sony. "At this scale, observability and evaluation are essential. With Amazon CloudWatch Omni, I can go from a single trace directly to evaluation, AI analysis, comparison or dataset creation." Oba also said assembling evaluation datasets is often a business-side bottleneck, and that one-click dataset creation from live traces removes it.
Omni runs outside the AWS Management Console. Developers get a native extension for Visual Studio Code, Cursor and Kiro, where traces appear as an agent runs locally and no AWS account is required. Operators get a standalone web experience with single sign-on through existing identity providers such as Okta and Microsoft Entra ID. Both share a single data layer, so the trace a developer debugs is the same one an operator investigates. AWS acknowledges that its console was built for infrastructure administrators, not for the site reliability engineers, AI engineers and application owners who now carry operational responsibility. Meeting developers inside the integrated development environment, where AI coding assistants such as Claude Code and Codex can set up instrumentation, shifts quality work to the point where problems are cheapest to fix, according to the report.
Omni's differentiator, in SiliconANGLE's assessment, is that agent traces, application telemetry and infrastructure signals all live in the same CloudWatch data store. An investigation can therefore start with an agent receiving a bad tool result, move to an application programming interface error from a capacity-limited service and end with an exhausted database connection pool. AWS DevOps Agent is enabled by default in investigation sessions, correlating signals and maintaining a full investigation history.
Capital One was a design partner. "Capital One operates one of the largest observability footprints in financial services," said Parvez Naqvi, managing vice president of cloud platform and resilience engineering at Capital One. "As a design partner for Amazon CloudWatch Omni, we helped shape a single AI-powered observability solution that will give our engineers topology-aware intelligence and natural-language querying across all telemetry from a single surface, with full data ownership through OpenTelemetry."
For heavily regulated industries such as banking, the report states that data portability is critical, and that it depends on data ownership. Captured investigation history is described as underrated for the same reason, serving in regulated industries as audit evidence of how an AI incident was handled.
Editor's Summary
AWS has made CloudWatch Omni generally available, combining agent traces, application telemetry and infrastructure signals in one data store with 17 built-in evaluators to score agent output quality. Sony and Capital One are early adopters and design partners, and AWS is betting that evaluation, not uptime, becomes the core measure as agent deployments scale toward an IDC forecast of more than 1 billion by 2029.