Defining the debt

Integration debt is the accumulated cost of connections built for an immediate need with no durable contract behind them. A point-to-point link written in an afternoon is often the right call at the time; it becomes debt when nobody owns it, nobody can describe its failure behaviour, and both ends have since changed without telling it.

Most organisations are not suffering from bad integrations. They are suffering from unowned ones, and invisibility is what lets the debt compound.

Four numbers worth measuring

  • Connections. Every path one system uses to read or write another, including scheduled file drops. Most operators are surprised by the count.
  • Schema awareness. How many systems know another system's internal shape — a direct table read, a hardcoded status code. This predicts how expensive change will be.
  • Manual reconciliations. Every place a human compares two reports. Each is a connector the business runs by hand, with a salary attached.
  • Time to detect a failure. Not to fix it — to notice. If the honest answer is month end, that number alone justifies the work.

The interest you pay

Changes propagate

When four systems read one status column, changing its meaning is four coordinated releases with four owners. In practice the change is deferred, then worked around, and the estate gets harder to reason about. Each workaround raises the cost of the next change.

Failures are silent

A connector without alerting fails in the least visible way possible: records stop arriving, nothing errors, and yesterday's numbers look current. Detection latency is what turns a small fault into an operational one, as we argue in uptime definitions.

Ownership is ambiguous

Ask who is accountable when the nightly job does not run and the answer is usually whoever notices first. Ambiguous ownership means no monitoring gets added, because adding monitoring is somebody's job.

An integration nobody owns will be found by finance, at month end, as a number that does not tie out.

How to prioritise retirement

  1. Blast radius. How many people cannot work if it breaks, and for how long. A customer-facing sync outranks an internal analytics feed.
  2. Change frequency. Connections attached to systems you are about to modify carry more risk than stable ones.
  3. Detection gap. Anything slow to detect gets instrumented first, even at low priority. Monitoring is cheap and often reveals the fault immediately.

The ordering is deliberately not “biggest mess first”. Reducing surprise is the goal. In the integration hub case, the first connections replaced were not the oldest but the ones whose failures reached a customer.

What a governed layer looks like

Where the count justifies it, a governed layer is a set of properties every connection must satisfy, not a product purchase.

  • Canonical model. One definition of customer, job and asset, with translation at the edges.
  • Idempotency. Every write carries an idempotency key so a retry cannot duplicate an order or double-count labour.
  • Dead-letter handling. Failures land in a dead-letter queue with payload and reason retained, and someone is alerted.
  • Reconciliation reports. A daily statement that counts and totals agree, visible to engineering and finance alike.
  • Contract tests. Automated checks in CI that a provider still honours the shape a consumer depends on.

Note what is absent: a vendor, a broker, a diagram. Those are implementation choices that should follow from the properties.

The contrarian part

The integration industry has a commercial interest in convincing you that every connection must pass through a hub. We do not agree, and we have talked clients out of hubs. Two systems with one stable, low-volume exchange do not need a canonical model and a governance forum; a monitored, documented direct connector is cheaper to run and easier to debug.

The hub earns its place when the count climbs, when several consumers need the same transformed data, or when you need one place to answer what happened to a record. Below that, spend on monitoring and reconciliation instead. It is the same restraint we argue for in modernisation sequencing, and access control across those connections is covered in security by design.

Count your connections first

We will help you build the baseline and tell you honestly whether a hub is justified in your estate.

Integration Services