Seven things the document must specify
Two parties can sign a page containing “99.9% availability” and still disagree about whether a given Tuesday counted as downtime. These are the items where that disagreement comes from.
- Measurement window. Monthly, rolling thirty days or quarterly. A monthly window resets after a bad month.
- What counts as downtime. Total unavailability, or degradation? A system that accepts jobs but cannot dispatch them is operationally down.
- Maintenance exclusions. A window, a notice period and a cap. Without a cap, maintenance excludes everything.
- Coverage hours. Business hours in AWST, extended hours or 24/7. A commitment that sleeps overnight is a different product.
- Response versus resolution. First human response within a time is not a fix within a time.
- Dependency carve-outs. Third-party APIs, gateways, telco, the client's own network and identity provider — named, not assumed.
- Who declares an incident. If the supplier starts the clock, it starts late.
The cost curve of nines
The jumps are not linear, because each additional nine requires that a class of failure becomes architecturally impossible rather than merely recoverable. Roughly, 99.9% permits a little under forty-five minutes of unavailability a month; 99.99% permits about four and a half.
Two to three nines is largely operational discipline — monitoring, patching, decent deployment. Three to four means removing single points of failure everywhere: multi-AZ deployment, automated data-tier failover, stateless nodes, and a rollback needing no human judgement at 3am. Four to five means designing out correlated failure across regions with a team large enough to rehearse it.
For most mid-market systems the third nine is worth having. The fourth is an architectural programme with recurring cost, often bought by organisations whose real exposure is a four-hour outage every eighteen months.
Availability targets should follow the cost of an outage, not the number that looked good in a proposal.
Fast recovery beats another nine
If your realistic bad day is a corrupted deployment or a failed restore, redundancy across zones does not address it — a bad release reaches every zone at once. What reduces exposure is detecting quickly and reverting confidently.
That means backups with a restore performed on a schedule and timed; runbooks written for the person on call; a rollback rehearsed in anger; infrastructure defined in code; and alerting that reaches a human in minutes.
Two numbers frame the rest. RPO, how much data you can lose measured in time, and RTO, how long restoration may take. An operator who tolerates a two-hour outage but cannot lose a day of field capture wants excellent backups and no multi-region architecture at all.
Support tiers and indicative costs
Coverage is what you are buying. Indicative monthly ranges for one business-critical application, excluding hosting:
- Business hours, next-business-day response: about A$1,200 to A$2,800, where an overnight delay is inconvenient rather than costly.
- Extended hours, two-hour response in coverage: about A$3,000 to A$6,500, common where WA operations start early or finish late.
- 24/7, thirty-minute response on severity one: about A$8,500 to A$16,000 depending on complexity. A staffing commitment, which is why it dominates the price.
What matters more is that severity levels and exclusions are written down, and that the people who built the system are on the roster — the basis of our maintenance and support arrangements.
What we put in writing
We would rather commit to a modest number we can be measured against than an impressive one with wide exclusions: a defined window and health check, named carve-outs, response targets by severity, a scheduled restore test with the result reported, and an incident review after anything significant.
Availability is rarely limited by the application tier anyway. The failures that hurt originate in unowned connections, as argued in integration debt, or in recovery capability, covered in security by design.
Want an SLA that survives an argument?
We will help define measurement, severity, carve-outs and recovery targets that match what your operation actually loses during an outage.