SIGNAL
Tracking the global AI frontier — labs · research · agents · policy
Frontier Signal
Practice

How Cornerstone OnDemand cut database diagnosis by 78% with Amazon Bedrock

Cornerstone OnDemand built Orion AI, a multi-agent system on Amazon Bedrock and Strands Agents, to turn database operations from reactive firefighting into proactive automation. A three-person team cut database diagnosis from 45 minutes to 10, a 78% reduction, in six months. See the design decisions other teams can reuse.

How Cornerstone OnDemand cut database diagnosis by 78% with Amazon Bedrock
Primary source aws.amazon.com ↗

Published October 7, 2026 · Category: AI Practice

Overview

Cornerstone OnDemand, Inc. (Cornerstone) is a global leader in workforce readiness solutions, serving 140 million users across 186 countries. The company built a multi-agent AI system that transforms database operations from reactive firefighting into proactive, self-orchestrating workflows. The system, called Orion AI, uses Amazon Bedrock and Strands Agents, an open source agent orchestration framework from AWS, to coordinate specialized agents.

Before Orion AI, Cornerstone’s Enterprise DataOps team would spend up to 45 minutes per database incident manually querying system views and cross-referencing logs before handing off across teams. With Orion AI, database diagnosis dropped from 45 minutes to 10, a 78% reduction. A three-person team delivered the system in six months.

In this post, we walk through the operational problem Cornerstone faced, how they designed Orion AI, the outcomes they measured, and the design decisions other teams can reuse.

The operational challenge

Before Orion AI, Cornerstone’s Enterprise DataOps team operated reactively, with four recurring pain points:

  • Database performance investigations took approximately 45 minutes per incident, spread across multiple tools and system views.
  • Database lifecycle workflows required 10 or more manual steps, from establishing connections to communicating status updates.
  • Reporting between the site reliability engineering (SRE) team and the data team ran on a 15-minute lag.
  • Redundant and overlapping alerts created noise that buried the signals that mattered.

Together, these meant engineers spent their time on manual coordination instead of resolving issues.

The solution: Orion AI

Orion AI is a multi-agent system where a coordinating agent (the meta-orchestrator) delegates work to specialized child agents arranged in a hub-and-spoke topology. Engineers interact with it through a web application, asking questions and approving actions in the same place they already coordinate operational work. The agents cover infrastructure monitoring, database diagnostics, database lifecycle operations, customer analytics, and knowledge-driven support.

Two tenets guided the build: data privacy using AWS controls under the shared responsibility model, and deep integration with existing operational tools. Amazon Bedrock provided managed access to foundation models, and Strands Agents provided the orchestration layer.

Results

Orion AI delivered measurable improvements in both diagnosis speed and accuracy and freed the team from manual coordination. By automating manual bottlenecks and streamlining cross-team workflows, Cornerstone achieved the following outcomes.

Metric Before After Improvement
Database diagnosis time 45 minutes 10 minutes 78% faster
Manual lifecycle steps 10+ steps 1 interaction 70% reduction
SRE-to-data-team reporting lag 15 minutes Instantaneous Real-time
Redundant alerts High baseline volume Filtered 65% reduction (median)

Faster performance troubleshooting

Reducing database diagnosis time was Orion AI’s largest single efficiency gain to date. Previously, engineers manually connected to the affected SQL Server instance, queried system views for blocking chains and wait types, cross-referenced logs to isolate long-running queries, and correlated evidence to form a root cause hypothesis. This averaged approximately 45 minutes per incident.

With Orion AI, three specialized agents share the investigation, each owning a distinct stage. Beyond diagnosis, Orion AI carries issues through remediation: identifying the root cause, recommending the fix, and creating a populated Jira ticket assigned to the right on-call engineer. A four-step manual handoff collapses into a single interaction.

Automated database lifecycle management

Tasks that previously required over 10 manual steps, from establishing database connections to running cross-system queries to communicating status updates, now run as a single natural language interaction. Orion AI identifies the right tools and data sources, runs queries across systems, validates results, and returns a unified response. From the engineer’s perspective, the work is reduced to prompting Orion AI and validating the findings it reports back.

Real-time monitoring and reduced alert fatigue

Orion AI removed the 15-minute SRE-to-data-team reporting lag, replacing periodic manual checks with continuous cross-system visibility.

Orion AI reduced redundant alerts by a median of 65 percent through deduplication, threshold filtering, and cross-signal correlation. For every 10 alerts previously generated, only 3–4 now reach engineers. Alerts route directly to the responsible teams without an intermediate alerting layer.

Key design principles

Three decisions shaped Orion AI, and you can apply them to your own multi-agent projects:

  1. Split agents by domain, not by task complexity: Each agent carries a narrow set of tool integrations for one operational domain. This keeps a model’s context focused and improves the accuracy of tool selection, rather than asking one general-purpose agent to do everything.
  2. Default to keyword routing, fall back to semantic search: Predictable requests are matched by keyword for speed. Ambiguous requests fall back to semantic search for accuracy. This keeps latency low without sacrificing correctness on the harder queries.
  3. Scope conversational memory to sessions and bypass it for live metrics: Persistent memory supports multi-turn conversations, but operational questions always read current system state. Bypassing memory for live metrics helps prevent stale data from contaminating real-time diagnostics.

Architecture overview

Orion AI is deployed as containerized services on Amazon Elastic Container Service (Amazon ECS). Users interact through a web application that routes each request to the appropriate agent, invokes the required tools, and assembles the response. The following diagram shows how these components connect within the AWS Cloud.

Architecture diagram of Orion AI within AWS Cloud. Web Application Users and external integrations (chat, issue tracking, incident management, and metrics dashboards) connect to a TaskExecutor agent running on Amazon ECS inside a virtual private cloud (VPC). The TaskExecutor acts as the meta-orchestrator hub, routing to specialist spoke agents including a Database Diagnostics agent, a Session Blocking Analysis agent, and other agents. A Portal-Tools MCP server inside the VPC connects the agents to the Data Plane, which holds SQL Server databases and the DATAOPS API. Amazon DynamoDB provides short-term memory. Amazon Bedrock, outside the VPC, provides model invocation, Amazon Bedrock AgentCore memory for long-term memory, and Amazon Bedrock Knowledge Bases backed by Amazon Simple Storage Service (Amazon S3). Amazon CloudWatch and AWS X-Ray provide observability.

Figure 1: Orion AI architecture: a meta-orchestrator on Amazon ECS routes requests to specialized agents, which reach data sources through a Portal-Tools MCP server and use Amazon Bedrock for models, long-term memory, and retrieval augmented generation

Architecture deep dive

The preceding design principles show up concretely in the architecture. The rest of this section walks through Orion AI’s core components.

Building the hub-and-spoke pattern with Strands Agents

Orion AI expresses its topology with a handful of Strands Agents primitives. The Agent class is the building block for both specialist agents and the meta-orchestrator. Each agent’s tools are ordinary Python functions marked with the @tool decorator. This derives a tool’s specification from the function’s type hints and docstring, meaning there is no separate schema to maintain. BedrockModel wraps the Amazon Bedrock model each agent runs on, and MCPClient connects agents to external tool servers.

Details

The topology:

  • The hub is a single Strands Agent acting as the meta-orchestrator (the TaskExecutor). It holds no domain tools — only routing tools (find_relevant_agents, call_agent) and control-flow tools (emit_plan_step, emit_confirmation_gate) so it can reason through a request step by step.
  • The spokes are independent Agent instances, loaded lazily through a decorator-based registry. Each carries its own model, domain tools, and localized system prompt.

At runtime, the hub invokes a specialist by name through the registered call_agent tool. Each specialist runs its own tool-calling loop and returns synthesized text back to the hub, which assembles the final response.

Hybrid routing

Orion AI uses keyword-first routing with a semantic search fallback. Approximately 80 percent of queries resolve through an in-memory keyword fast-path in under a millisecond. When keywords are insufficient to determine intent, the system falls back to semantic search powered by Amazon Titan Text Embeddings V2 to route by meaning. For model availability by AWS Region, see Supported models by AWS Region in Amazon Bedrock.

When direct tool invocation isn’t appropriate, a fallback chain activates: Amazon Bedrock Knowledge Bases, the fully managed Retrieval Augmented Generation (RAG) capability, retrieves operational documentation so responses are grounded in approved procedures rather than generated without context.

Domain-specific agents and their boundaries

Orion AI uses 13 domain-specific agents. The three SQL Server agents illustrate how domain splitting maps to the stages of a real investigation:

  • Database diagnostics agent surfaces blocking chains, wait types, and long-running queries, then translates technical findings into business impact with remediation recommendations. It answers, “what is happening and what does it mean.”
  • Session blocking analysis agent runs multi-step investigations with reasoning models to map blocking chains to a root blocker. It answers, “why is this happening and what is the root cause,” delivering prioritized recommendations from immediate session termination to longer-term architectural fixes.
  • Real-time SQL diagnostics agent queries SQL Server instances directly through Cornerstone’s DATAOPS API for live blocking data and high-CPU session analysis. It answers, “what is true right now.”

The remaining agents follow the same principle of narrow ownership: infrastructure monitoring, operational analytics, database lifecycle management, customer analytics, knowledge retrieval from runbooks, notification routing, compound query decomposition, Availability Group listener resolution, and model connection warm-up.

Tool integration and service connectivity

Orion AI reaches data sources through two patterns:

  • Model Context Protocol (MCP) for tools that several agents share (for example, real-time SQL diagnostics). A Strands @tool function calls an MCPClient over streamable HTTP to a Portal-Tools MCP server, which reaches SQL Server and the DATAOPS API. Jira integration follows the same approach through external Atlassian MCP tools.
  • Direct SDK or REST API calls for sources that don’t need the shared MCP interface. This includes metrics and dashboards, on-call schedules, and Amazon Bedrock Knowledge Bases through AWS SDK for Python (Boto3).

This split means each agent uses the lightest mechanism that fits its source, while the MCP server centralizes the tools that several agents share. These connections run over TLS, and credentials such as the MCP token and REST API authorization are supplied per request, not embedded in agent code.

Conversational memory

Amazon Bedrock AgentCore, a platform to build, connect, and optimize agents at scale with any framework or model, provides the cross-session memory layer. Orion AI coordinates context through a central memory manager that reads from three tiers in parallel, each with its own timeout and token allocation.

Orion AI coordinates context across three tiers. Short-term memory keeps same-session context in Amazon DynamoDB, encrypted at rest by default, paired with a rolling conversation summary that Amazon Nova 2 Lite generates asynchronously. Long-term memory provides cross-session recall through AgentCore memory, a capability of Amazon Bedrock AgentCore, namespaced by user ID. On task completion, Orion AI calls create_event to trigger automated extraction, summarization, and consolidation. A third tier, inter-agent memory, is an ephemeral in-memory scratchpad that carries findings across sub-agent steps within a single task.

The memory manager retrieves across the tiers within a 500-millisecond hard timeout against a 4,000-token budget, backed by a least-recently-used cache. When a request concerns current system state, agents bypass memory entirely and read live data, so stale context does not contaminate real-time diagnostics.

Responsible AI and guardrails

Because DataOps safety depends on domain-specific rules, such as which database operations are disruptive, the team built custom guardrail logic rather than relying on generic content filters. Orion AI applies controls across four layers:

  • Prompt-level safety constraints injected into every agent’s system prompt, blocking dangerous recommendations (for example, terminating critical database processes) and enforcing non-disruptive methods.
  • A human-in-the-loop confirmation gate that pauses destructive operations until the user confirms within a five-minute window, defaulting to deny on timeout.
  • Custom guardrail modules for input validation (length limits, prompt- and SQL-injection blocking), output sanitization (redacting secrets and personally identifiable information), rate limiting, role-based access control, and per-request cost tracking.
  • Routing-level protection that blocks answering operational queries from memory, forcing live tool execution for current-state questions.

Observability

Orion AI uses standard AWS observability services for monitoring and debugging. Amazon CloudWatch provides metrics and structured logging for routing confidence, latency, and agent invocation activity. AWS X-Ray traces distributed calls across agent executions for end-to-end request path visibility.

Lessons learned

The design decisions that made Orion AI possible are ones you can reuse. Splitting agents by domain let each agent carry narrow tool integrations without bloating a single model’s context. Using keyword routing by default with semantic search as a fallback kept latency low without sacrificing accuracy on ambiguous requests. Scoping conversational memory to sessions while bypassing it for live metrics kept real-time diagnostics trustworthy.

The team also made pragmatic infrastructure choices. They deployed compute on Amazon ECS because Amazon Bedrock AgentCore runtime, a capability of Amazon Bedrock AgentCore, wasn’t available at the start of the project, and they’re now evaluating it as a future migration option. If you’re building today, weigh the same tradeoff for your own stack: start with what’s stable and available now, and revisit as managed capabilities mature.

Conclusion

Orion AI reduced database diagnosis from 45 minutes to 10, avoided 70 percent of manual lifecycle steps, and cut alert noise by 65 percent. A three-person team delivered it in six months on Amazon Bedrock. The patterns behind those results are ones other operations teams can adopt: domain-scoped agents, hybrid routing, and session-scoped memory with a live-metrics bypass.

To get started with multi-agent orchestration on AWS, review the AWS Solutions Library guidance.


About the authors

Harpreet Chawla

Harpreet Chawla

Harpreet, Cornerstone OnDemand, Inc., is a Strategic Data & AI Infrastructure Executive with over 20 years of experience driving large-scale digital transformation and operational excellence for global enterprises. He specializes in scaling mission-critical data platforms on AWS while pioneering the shift toward Agentic AI and Autonomous Platform Engineering.

Derek Ziehl

Derek Ziehl

Derek is a Senior Technical Account Manager (TAM) at AWS. He has a background designing large-scale network systems and managing cloud migrations. As a TAM he enjoys enabling customers to run resilient, optimized workloads on AWS.

Madhu Pai

Madhu Pai

Madhu is a Principal Specialist Solutions Architect for Generative AI and Machine Learning at AWS, where he helps ISVs and enterprises move AI from experimentation to production at scale. He partners with organizations across the industry to navigate the real barriers to AI adoption: governance, compliance, data readiness, and trust. Drawing on deep field experience, Madhu brings a pragmatic, vendor-neutral perspective to how teams can responsibly scale AI.

Nishchai JM

Nishchai JM

Nishchai is a Senior Solutions Architect at Amazon Web Services, specializing in Data, Analytics, and Generative AI for Independent Software Vendors (ISVs). He helps customers modernize their data platforms, build large-scale distributed applications, and unlock value from AI/ML workloads on the cloud.

Source

Originally published at aws.amazon.com.

Related Articles

F
Frontier Signal Desk

Frontier Signal tracks the global AI frontier — labs, research, agents, creation tools and real-world practice — straight from primary sources. Tip the desk: editorial@news.tunx.ai

Email the desk →
From our network: explore the AI assistant platform behind this site. Visit tunx.ai →
Note: This story is aggregated and summarized from the primary source linked above; the original publisher retains all rights. Details may evolve after publication — always confirm against the source. Nothing here is professional, legal or investment advice.

Related Stories

More from Practice →