Technology · Cloud

Infrastructure somebody must be able to rebuild later

A platform decision sets your monthly bill, your recovery options and who can help at two in the morning. We build on AWS and Azure, in Australian regions, with the estate written down as code.

What cloud decisions actually cost

Compute is rarely the largest line on a mid-market bill; data egress, idle environments, snapshot retention and the engineer time spent maintaining things nobody audits are. We size estates against observed load, tag every resource to an owner and environment, and review spend monthly. Most projects run entirely in one cloud, using Australian regions.

AWS services we build on

4 technologies
ECS Fargate
Lambda
Amazon S3
CloudWatch

ECS on Fargate

The default runtime for containerised applications: no servers to patch and straightforward scaling. It costs more per hour than the equivalent instance, which is the right trade wherever nobody wants to manage operating systems.

Lambda

Used at the edges for scheduled jobs, webhook receivers and file processing. Cold starts and connection limits make it unsuitable as the entire application, though it handles irregular background work economically.

Amazon S3

Object storage for documents, exports and backup copies, with versioning and lifecycle rules declared in code. Accidental exposure comes from hand-made permissions, so access is never configured outside Terraform.

CloudWatch

Metrics, logs and alarms configured alongside the application rather than afterwards. Retention periods are set deliberately and verbose debug output is sampled, because querying unbounded log storage becomes expensive quietly.

Azure equivalents

4 technologies
App Service
Azure Functions
Blob Storage
Application Insights

App Service

The Azure default for web workloads, offering managed hosting and deployment slots for staged releases. It is less flexible than containers for unusual runtimes, which occasionally decides the question on its own.

Azure Functions

Azure's serverless option, applied to the same edge work as its AWS counterpart. Selection between them follows whichever cloud the client already operates rather than any functional difference we consider decisive.

Blob Storage

Hosts documents and exports with lifecycle rules moving ageing artefacts to cheaper tiers. Keys rotate through the platform secret store rather than appearing in configuration files, where they inevitably get copied.

Application Insights

Distributed tracing and failure diagnostics, particularly useful in .NET-heavy estates. Sampling is configured early, otherwise traces either cost more than expected or discard the request someone needed to inspect.

Infrastructure as code and delivery

4 technologies
Terraform
Docker
GitHub Actions
GitLab CI

Terraform

Every environment comes from shared modules, so staging and production differ only by variables. State is remote and locked, plans are reviewed before apply, and no resource is created outside the definitions.

Docker

Images built from minimal pinned bases and scanned before deployment. Reproducibility matters most on inherited platforms, where reliably rebuilding an old application is frequently the first milestone worth celebrating.

GitHub Actions

The pipeline we propose by default: tests, scanning and deployment defined in the repository and visible to whoever maintains the code later. Minutes-based pricing means long jobs are watched rather than assumed free.

GitLab CI

Used where a client already standardises on GitLab, since relocating repositories to suit our preference is an imposition. Self-hosted runners are sometimes required to reach resources inside the client's own network.

Practical notes from production

Cost and reliability lessons that arrived as incidents before becoming standard practice.

01

Egress and NAT charges dominate small bills

Data leaving onshore through a managed gateway is billed per gigabyte, and chatty integrations produce surprising totals. Private endpoints and caching are considered before deployment, not after the first invoice.

02

Untagged resources cannot be attributed

Tagging by owner, environment and cost centre is enforced through policy at creation time. Without it, nobody can answer which department a growing line item belongs to, so nothing ever gets switched off.

03

Restore time must be measured, not assumed

Recreating a large instance from snapshot takes far longer than most expect. Recovery objectives are agreed in writing and rehearsed, because a four-hour restore discovered during an outage is a business problem.

04

Console changes become configuration drift

A security group adjusted directly to fix something today will be quietly reverted by the next apply. Anything genuinely needed is added to the definitions immediately, otherwise the next deploy undoes it without warning.

05

Vendor allowlists need stable addresses

Some Australian platforms accept traffic only from approved addresses, which requires fixed egress. Static addressing is provisioned and documented at build time, since a rotation otherwise breaks the integration silently.

Choose it when, avoid it when

Platform questions are rarely about capability. They are about who absorbs the operational work for the next five years.

Managed services versus self-managed

Choose managed whenever nobody internally will patch queues, brokers or databases. Avoid self-hosting driven by unit-price comparison alone, because the retained ops salary does not appear on any quote.

Serverless versus containers

Choose serverless for irregular or scheduled work and rarely used endpoints. Choose containers for steady applications with database connections and background processes, where warm instances simplify everything downstream.

One cloud versus two

Choose one, almost always. Duplicating a platform across providers doubles the expertise required and halves the familiarity anyone develops, producing an expensive architecture nobody is confident operating at 2am.

Continuous release versus release windows

Continuous release is safe once automated tests, migrations and rollback exist. Before that, or for systems with nightly financial batches, releases belong in agreed AWST windows with someone present afterwards.

Where we draw the line

Operational limits we apply regardless of what a project might otherwise bill.

Kubernetes for five services

We do not run our own cluster for workloads of the size we usually see. The platform needs upgrades, node management and networking expertise of its own, which is hard to justify when a managed container service covers the requirement.

Accounts without the basics

We will not deploy into environments lacking multi-factor authentication, named individual accounts and billing alerts. Administrative access shared through one credential makes every later problem unattributable, which is unacceptable.

Infrastructure nobody wrote down

Manual changes that solve today's problem become tomorrow's outage when the next apply reverts them. Everything returns to code the same week, including emergency fixes, and we include that work in every estimate.

Offshore data by accident

We do not accept designs whose monitoring, logging or analytics quietly send client data outside Australia. Where no compliant option exists for a requested service, we say so and propose something else rather than proceed.

Check whether your cloud matches its bill

We review running estates for spend, resilience and handover risk before recommending any change.

Get a Quote