Skip to content
All articles

AWS Architecture

What multi-AZ buys, and what it doesn't

Availability zones are a failure-domain definition. Multi-AZ survives a power or fibre loss, not a bad deploy, an expired certificate or a corrupt migration, and only with spare capacity.

· 4 min read

A table of failures against what survives them. Power, cooling or fibre loss and a rack or host failure: multi-AZ yes, multi-region yes. A regional control-plane incident: multi-AZ no, multi-region yes if the standby needs no API calls in the failed region. A bad deployment, an expired certificate and a corrupt schema migration: no in both columns. An exhausted account quota: multi-AZ no, multi-region partly. Below: three zones survive losing one only if two can carry peak load, so run at or below 66% utilisation.

"Region" and "availability zone" sound like marketing. They are really definitions of failure domains, and knowing exactly which failures each one contains is most of designing for them.

A region is a geographic area with its own control plane, such as ap-south-1 in Mumbai or eu-west-1 in Ireland. Regions are almost completely independent: separate APIs, separate service endpoints, separate failure domains, and no data moves between them unless you move it.

An availability zone is one or more data centres inside a region, with independent power, cooling and networking, linked to the other zones by low-latency connections of typically single-digit milliseconds. Zones fail independently, for physical reasons: a power event, a cooling failure, a fibre cut.

What each boundary covers

FailureMulti-AZ survives itMulti-region survives it
Power, cooling or fibre loss in one data centreYesYes
A rack or host failureYesYes
A regional control-plane incidentNoYes, if the standby needs no API calls in the failed region
A bad deploymentNoNo, and active-active makes it worse
An expired certificateNoNo
A corrupt schema migrationNoNo: replication faithfully copies the corruption
An exhausted account quotaNoPartly, because quotas are per region

Three rows say no in both columns, and the quota row is covered only in part. A deployment pipeline crosses every zone and region without noticing them, the same certificate sits in every zone, and a corrupt migration replicates. Changes like these cause a large share of outages. Multi-AZ protects against infrastructure failure, multi-region against regional failure, and neither protects against you. A team asking for a third zone when its last four incidents were deploys is optimising the column that was already covered.

The two boundaries cost different things

Crossing a zone costs one to two milliseconds and a per-gigabyte charge on the traffic, which is small enough that spreading across zones should be the default for everything. Crossing a region costs tens to over a hundred milliseconds; Mumbai to Ireland is roughly 120 ms round trip. That is an architectural constraint rather than a line item. A synchronous cross-region call inside a request path spends more time on the wire than most services spend on their entire response, and no amount of tuning wins it back.

Common Mistake

Calling a service multi-AZ because it runs in three zones. Lose one zone and capacity drops by a third; if the other two saturate, the service failed along with the zone. Surviving the loss of a third means serving peak load on the remaining two thirds, which means running at or below 66% utilisation at peak. A zone-removal drill tells you which kind of deployment you actually have.

Get one diagram a week

A short article built around one engineering diagram, from the same library as these courses.

One diagram-led article a week on AI and systems engineering. We email you once to confirm, and every newsletter has an unsubscribe link. Privacy policy