Six design decisions that become the invoice

Optimisation exercises start with a billing report and end with a list of idle resources. Useful, but it treats symptoms.

  • Transfer paths. Chatty services crossing zones, or egress through a NAT gateway, are billed per gigabyte. Usually an accident of placement.
  • Idle compute. Non-production environments sized like production, running overnight and across weekends in AWST.
  • Untagged resources. If a resource has no owner, nobody questions it and it outlives the project that created it.
  • Oversized databases. Provisioned for a peak that happened once, in the least elastic part of the bill.
  • Log retention. Debug logging left on with indefinite retention becomes a storage line that grows every month.
  • Cross-zone chatter. cross-AZ traffic to a database or cache is often a default, not a decision.

None are finance decisions. They follow from how the system was drawn.

A governance model that actually holds

Governance fails when it depends on someone remembering to run a report.

Tagging and named ownership

Every resource carries an owner, a service and an environment, applied through infrastructure as code. Untagged resources are a policy violation, not a tolerated gap: when something is expensive, the first question needs a name attached.

Per-environment policy with automatic shutdown

Non-production stops outside working hours, with an override for exceptions. The largest reliably recovered saving in most estates we look at, and it takes hours.

A right-sizing cadence

Quarterly, compare utilisation with provisioned capacity using observed percentiles rather than the annual peak, and keep an autoscaling group for genuinely variable load.

Alerts tied to services, not accounts

An account-level alert says the total went up. A service-level alert says which component changed.

Commit only against stable load

Reserved capacity is a bet on steady state: take it for the always-on core and keep the rest flexible, because committing against a growth projection is how you pay twice for capacity you over-estimated.

The metric that changes behaviour is not total spend; it is cost per business transaction.

Cost per server is meaningless outside engineering. Cost per job dispatched or invoice issued is legible to decision-makers, and it exposes the trade-off that matters: if volume doubles and unit cost holds, the platform is scaled correctly; if unit cost rises with volume, the design is fighting you.

The honest account of lift-and-shift

We are sceptical of the claim that migration reduces cost, and say so before starting. Lifting an application largely unchanged onto rented infrastructure usually costs about the same once managed services are included, and sometimes more, because you now pay separately for what came bundled with the old hardware.

What it buys is different and more valuable: environments rebuilt from code in an afternoon, deployment without an outage window, scaling without a procurement cycle, recovery without driving to a data centre. Those justify the staged migration on their own terms.

Where we disagree with standard advice

Two recommendations deserve pushback. The first is aggressive consolidation onto managed services: excellent, priced accordingly, and for steady high-volume workloads the premium is real. We would rather you ran a well-understood database yourself than paid a multiple for convenience nobody can quantify.

The second is treating every cost increase as waste. Check unit cost before cutting — if it is flat while spend rises with volume, the platform is doing its job. Disciplined design still matters, which is the argument in modernisation, integration debt and uptime definitions.

Want the bill explained, not discounted?

We will review the architecture behind the spend and say which lines are design decisions, which are growth, and which are waste.

Cloud Services