HPE Zerto is deploying an on-premises AI troubleshooting system with Amazon Bedrock, linking live recovery data to guarded, actionable guidance.

HPE Zerto has built an agentic troubleshooting system that runs inside a customer’s on-premises environment while using Amazon Bedrock to power its reasoning layer. The system connects a natural-language interface to live disaster-recovery data, product documentation, and operational tools, aiming to help resilience teams investigate problems without manually switching between dashboards and support materials.
The architecture, detailed in an AWS Machine Learning Blog post, reflects a practical constraint for enterprise AI: the model service may be cloud-based, but sensitive operational context and tool execution need to remain close to the customer’s infrastructure. HPE Zerto says its system is deployed as a pod within the Zerto product and is accessed through the existing operations interface.
For AI builders and enterprise technology teams, the important development is less a new chatbot than the boundary HPE Zerto has drawn around it. Agents can reason over current environment state and retrieve relevant knowledge, but their access is mediated through local APIs, retrieval systems, and security controls. That design is intended to make an AI assistant useful during outages and configuration issues without turning a general-purpose model into an unrestricted operator.
According to the AWS Machine Learning Blog, the system uses agents built with Strands Agents. Those agents run as an on-premises pod and can draw on three main information sources: locally stored conversation history, internal tools, and external knowledge retrieval.
The internal tool path is exposed through a local Model Context Protocol server. That server provides structured access to Zerto Manager APIs, allowing the agent to inspect live information from the customer environment rather than relying only on static documentation or a user’s description of an incident.
The external path uses Amazon Bedrock Knowledge Bases to retrieve public documentation, runbooks, and other operational material. This combination is significant because troubleshooting disaster-recovery systems usually requires both current state and procedural context. An alert may show what is failing, while a runbook or product guide explains the safe sequence for investigating or correcting it.
The interface is embedded in the existing Zerto user experience. HPE Zerto says it uses Server-Sent Events to stream investigation progress to the user, exposing intermediate activity instead of waiting for a single final response. The stated use cases include answering configuration questions, helping mitigate health problems, assisting with setup and feature adoption, and summarizing risk, service-level agreement exposure, or recovery readiness.
That does not mean the model independently controls the entire recovery environment. The source describes a system designed to reason, act, and respond through the tools made available to it. The practical safety of the deployment therefore depends on the permissions, validation logic, and action boundaries implemented around those tools—details that the AWS post discusses at an architectural level but does not fully quantify as a production performance report.
HPE Zerto’s target users manage cyber resilience, continuous data protection, and disaster recovery across hybrid or multi-cloud environments. In those settings, operational data can include the health of protected workloads, replication status, alerts, events, site relationships, and recovery readiness. The company’s stated challenge is that this information is fragmented across interfaces and documentation, slowing decisions during outages or cyber incidents.
Running the agentic layer inside the customer environment addresses part of that problem. The assistant can be placed beside the systems it needs to inspect, while session history is stored locally and the product remains accessible through the familiar administrative UI. For enterprises, that may simplify deployment and reduce the need to move operational context into a separate external application.
The model layer still relies on Amazon Bedrock. AWS says HPE Zerto selected the service for controlled access to foundation models, the ability to evaluate different models, and integration with guardrails, observability, and retrieval services. HPE Zerto also uses Amazon Bedrock Guardrails to apply policy, compliance, and security controls to prompts and generated output.
This hybrid arrangement illustrates a deployment pattern increasingly relevant to enterprise AI: keep data access and execution near the system of record, while using a managed model platform for inference and model choice. It can offer flexibility, but it also creates dependencies across the local product, the network path to model services, retrieval quality, and the permissions granted to each agent.
The available evidence is an AWS Machine Learning Blog case study written around HPE Zerto’s architecture. It confirms the described components and deployment approach, but it does not provide independent validation of troubleshooting accuracy, response latency, customer adoption, reduced support-ticket volume, or measurable recovery-time improvements.
Claims about faster, more informed recovery decisions and reduced operational burden should therefore be read as the system’s intended outcomes rather than independently verified results. The source also does not specify which foundation models are used in production, how model routing is handled, or how often agents are permitted to take direct actions versus returning recommendations to an operator.
A second AWS Machine Learning Blog post about Intuit’s Intuit EWOK Agent offers a related reference point, not additional evidence about HPE Zerto. Intuit describes an agent that translates plain-language failover requests into validated workflows, while a deterministic execution system performs the infrastructure changes. Intuit says teams have used that system for eight months and reports that supported recovery workflows can complete in about 20 minutes, but those are Intuit’s own claims about its separate platform.
The two cases share a broader design principle: let an AI model interpret intent and select from bounded capabilities, while deterministic systems enforce policy and execute high-impact operations. That pattern is more transferable than any specific benchmark, but organizations should still test it against their own data quality, failure modes, and change-control requirements.
For builders, HPE Zerto’s approach highlights grounding as an integration problem rather than a prompt-writing exercise. The agent needs reliable access to current product state, a defined tool schema, relevant documentation, and enough session context to avoid repeatedly asking for information already available in the conversation. A local MCP server can provide a structured interface, but it also becomes a critical control point for authentication, authorization, logging, and API compatibility.
For enterprise buyers, the key evaluation questions are operational. Can the system distinguish a stale alert from a current condition? Does it cite or expose the source of a recommendation? Can administrators limit tools by role and environment? What happens when the model is unavailable, the retrieval result is incomplete, or the requested remediation could increase data-loss risk? The AWS post establishes the architecture, but it does not answer all of those procurement and governance questions.
The system’s value will likely be highest where teams face high information density and uneven operator experience. A natural-language layer can reduce the time needed to assemble context and explain symptoms, particularly for less experienced administrators. But in disaster recovery, helpful explanation is not the same as safe execution. Any move from diagnosis to remediation needs explicit permissions, confirmation steps, audit trails, and a reliable fallback to human operators.
The next meaningful signals will be evidence of deployment beyond the architecture description. HPE Zerto could clarify whether the assistant is generally available, which product editions and environments support it, and whether customers can configure model choice or tool permissions.
Builders should also watch for published measurements covering answer accuracy, retrieval quality, false recommendations, time saved during incident investigation, and the rate at which users accept or reject suggested actions. Those metrics would show whether the system improves operational work rather than simply adding a conversational interface.
On the platform side, model flexibility will remain an important test. HPE Zerto says Amazon Bedrock allows evaluation across foundation models based on quality, latency, and cost. Real-world comparisons, deployment guidance, and clearer details on data residency and network failure behavior would help enterprises assess whether that flexibility produces a practical advantage.
HPE Zerto’s system is a useful example of where enterprise agentic AI is becoming concrete: inside an existing operational product, connected to live state, and constrained by the APIs and controls that already govern the environment. The strongest part of the design is the separation between a conversational reasoning layer and the product’s underlying operational interfaces.
The harder question is proof. In disaster recovery, an agent must be judged not only by whether it can explain an alert, but by whether its guidance is current, auditable, and safe under pressure. Until HPE Zerto or its customers publish outcome data, the announcement is best understood as a credible deployment pattern and an architecture to evaluate—not yet evidence that agents have solved recovery operations.