"Region" and "availability zone" sound like marketing. They are really definitions of failure domains, and knowing exactly which failures each one contains is most of designing for them.
A region is a geographic area with its own control plane, such as ap-south-1 in Mumbai or eu-west-1 in Ireland. Regions are almost completely independent: separate APIs, separate service endpoints, separate failure domains, and no data moves between them unless you move it.
An availability zone is one or more data centres inside a region, with independent power, cooling and networking, linked to the other zones by low-latency connections of typically single-digit milliseconds. Zones fail independently, for physical reasons: a power event, a cooling failure, a fibre cut.
What each boundary covers
| Failure | Multi-AZ survives it | Multi-region survives it |
|---|---|---|
| Power, cooling or fibre loss in one data centre | Yes | Yes |
| A rack or host failure | Yes | Yes |
| A regional control-plane incident | No | Yes, if the standby needs no API calls in the failed region |
| A bad deployment | No | No, and active-active makes it worse |
| An expired certificate | No | No |
| A corrupt schema migration | No | No: replication faithfully copies the corruption |
| An exhausted account quota | No | Partly, because quotas are per region |
Three rows say no in both columns, and the quota row is covered only in part. A deployment pipeline crosses every zone and region without noticing them, the same certificate sits in every zone, and a corrupt migration replicates. Changes like these cause a large share of outages. Multi-AZ protects against infrastructure failure, multi-region against regional failure, and neither protects against you. A team asking for a third zone when its last four incidents were deploys is optimising the column that was already covered.
The two boundaries cost different things
Crossing a zone costs one to two milliseconds and a per-gigabyte charge on the traffic, which is small enough that spreading across zones should be the default for everything. Crossing a region costs tens to over a hundred milliseconds; Mumbai to Ireland is roughly 120 ms round trip. That is an architectural constraint rather than a line item. A synchronous cross-region call inside a request path spends more time on the wire than most services spend on their entire response, and no amount of tuning wins it back.
Common Mistake
Calling a service multi-AZ because it runs in three zones. Lose one zone and capacity drops by a third; if the other two saturate, the service failed along with the zone. Surviving the loss of a third means serving peak load on the remaining two thirds, which means running at or below 66% utilisation at peak. A zone-removal drill tells you which kind of deployment you actually have.


