What 99.9% Uptime Actually Means
“Three nines” sounds strict. 99.9% available — surely that is basically always up? It works out to nearly nine hours of downtime a year, or about 43 minutes in a 30-day month. Whether that is fine or alarming depends entirely on what your service does.
The nines, in real time
| Uptime | Name | Downtime per month | Downtime per year |
|---|---|---|---|
| 99% | two nines | 7h 12m | 3d 15h 36m |
| 99.9% | three nines | 43m 12s | 8h 45m 36s |
| 99.99% | four nines | 4m 19s | 52m 34s |
| 99.999% | five nines | 26s | 5m 15s |
Two patterns jump out. First, the jump from 99% to 99.9% removes almost four days of yearly downtime — a big, usually worthwhile improvement. Second, every nine after that cuts the remaining budget by about 90%, and the absolute savings get smaller while the engineering effort to achieve them grows sharply.
Why each nine is so expensive
Going from 99.9% to 99.99% means your yearly outage allowance drops from nearly nine hours to under an hour. In practice that requires:
- redundancy across independent failure domains, not just a spare server
- automated failover that is regularly tested, not a runbook someone follows at 3am
- a deploy process tight enough that releases rarely cause incidents
- on-call coverage with people who can actually fix things
- monitoring that catches problems in seconds, not when a customer emails
Each of those adds cost and operational complexity. Most products do not need more than three nines. Picking a target higher than your users require mostly buys you stress and a larger infrastructure bill.
SLA, SLO, and what you actually measured
Three related terms that get muddled:
- SLA — the availability you promise in a contract, often with refunds or credits if you miss it.
- SLO — the internal target your team works to. Set it stricter than the SLA so you have a safety margin before penalties kick in.
- Actual uptime — what you measured over the period. The gap between this and your SLO is your early-warning signal.
One more contract detail: many SLAs exclude announced maintenance windows from the downtime calculation. Read the definition of “downtime” and “excluded events” before you compare your measured number to the promise.
Pick the target from the cost of being down
The right availability target is a business decision, not an engineering one. Work out what an hour of downtime actually costs you — lost revenue, staff who cannot work, recovery effort — and compare that to the cost of the next nine. If an outage costs a few hundred dollars an hour, chasing five nines makes no sense. If it costs tens of thousands, three nines is probably not enough.
Do the maths
The uptime / SLA calculator converts any availability percentage into exact downtime per day, week, month and year, and also works backwards from a measured outage to the percentage it produced. Pair it with the cost of downtime calculator to put a dollar figure on each minute — the number that tells you how much reliability is worth buying.