A zone-redundant architecture on paper is no guarantee that failover actually works. Microsoft's new Azure Infrastructure Resiliency Manager links Availability Zones, Chaos Studio and Monitor together and lets you simulate a zone outage before a real one does.
Many IT teams assume their Azure environment is resilient the moment workloads are spread across multiple Availability Zones. That is an assumption, not proof. Without a test that actually takes a zone offline, nobody knows whether failover works as intended: whether a database switches over smoothly, whether an AKS cluster stays up without the right node pool, or whether an application silently depends on something that only existed in that one zone. The architecture diagram says everything is redundant. Reality only proves that once a zone actually goes down, and by then it is too late to fix anything.
This summer, at Build 2026, Microsoft introduced the Azure Infrastructure Resiliency Manager, now in public preview and available to all Azure customers. The tool does not add anything new to the underlying building blocks: Availability Zones, Azure Advisor, Azure Chaos Studio, Azure Monitor and Azure Copilot already existed. What was missing was a coherent process tying those separate pieces together into an ongoing resiliency strategy, instead of a collection of disconnected dashboards that nobody structurally compared.
The core of the problem is that resiliency often looks fine on paper but has never actually been tested. Advisor gives recommendations, Chaos Studio can simulate failures, Monitor shows metrics, but nothing automatically connected the three. Teams rarely saw the gap between the architecture they thought they had built and the configuration actually running in production. The Resiliency Manager makes that gap explicit: it shows, per resource, whether the intended zone redundancy is actually configured, not only at creation but also after later changes.
Microsoft organises the process into three phases. In the Start Resilient phase, the tool helps design new workloads with the right zone-redundant foundation before anything goes live. In the Get Resilient phase, the Resiliency Agent scans existing environments, flags misconfigurations, and explains the risks and trade-offs in plain language rather than an error code or policy notice. The Stay Resilient phase is where the tool proves its value: continuous validation through repeated drills and monitoring, so that a change quietly undermining resiliency, for example an engineer accidentally moving a resource out of its zone-redundant configuration, gets noticed quickly instead of only during the next outage.
The most concrete addition in this preview is availability zone failure drills, powered by Azure Chaos Studio. Instead of hoping failover works during a real outage, the drill simulates a zone failure in a controlled setting: virtual machines in the target zone are shut down, zone-redundant databases are forced to fail over, and AKS node pools in that zone are stopped. Within minutes you see whether your application absorbs the outage or still falls over, without a customer ever noticing and without having to wait for a real outage to get the answer.
The preview already covers a broad range of resource types: virtual machines, databases with zone-redundant configurations, AKS clusters and networking components. Microsoft says it is actively expanding coverage, so expect more resource types to follow in the coming months. For organisations already running zone-redundant workloads, that is reason enough to run the first scan now rather than waiting for the tool to be fully built out.
A missed failover costs more than just downtime. An outage lasting half an hour does not only affect the application involved, it also affects the trust of customers who cannot use a service at that moment, and internal confidence in your own architecture. For organisations with a service-level agreement toward customers, every unplanned minute counts, and those minutes are usually easier to prevent than to explain afterwards.
In practice, start with the Resiliency Agent, accessible via Azure Essentials in the portal. Let it run a first scan across your key production environments and see which resources look zone-redundant on paper but are not in practice. Then plan a first drill on a non-critical or staging environment, not straight on production, to get a feel for what a zone outage actually does to your architecture. Only once that exercise runs without surprises is it sensible to repeat drills periodically in production, for example every quarter or after every major architecture change.
For most organisations, this is a chance to turn resiliency from an assumption into a proven fact, without needing an actual outage to find the weak spots. Not sure whether your Azure environment is genuinely zone-redundant, or want help setting up and running your first drills? Zarioh is happy to help design and test a resilient cloud environment.
Zarioh Digital Solutions
IT specialists from Utrecht, the Netherlands. We help businesses with Microsoft 365, AI agents, hosting and telephony, and share what we learn in practice. Follow us on LinkedIn

Cloud & Infrastructure

Cloud & Infrastructure

Cloud & Infrastructure