A national wholesale distributor moved orders across four systems using fourteen hand-built connectors. We built one broker-led hub — idempotent, reconciled, alerting — without replacing an ERP that could not be modernised.
Every order crossed four systems: an ERP commissioned long before anyone expected an online channel, a CRM holding pricing and credit terms, a third-party logistics warehouse interstate, and a web store trading around the clock. Fourteen connectors joined them — cron scripts, exports and an FTP job untouched for years.
The hub substituted one publish-and-subscribe layer for all fourteen: a canonical model covering orders, customers, stock and shipments; publishers in Python, consumers in Node.js; RabbitMQ between them; and a reconciliation layer reporting stalled records within minutes. The ERP, CRM and 3PL stayed.
The connectors worked individually and failed collectively. Each assumed its own timing, keys and failure handling, so the consequence of one link breaking was unpredictable — and several failed quietly, surfacing during reconciliation rather than trading.
The boundary conditions were clear, and two decided the architecture: order intake could not pause, and no duplicate financial document could reach the ERP. An accounts trail nobody trusts is worse than a delay.
One message hub instead of a mesh. Each system speaks to the hub, never directly to another, and every message is validated against a versioned schema before publication. Consumers are individually replaceable — which is what retiring a connector now means.
Reconciliation — flows and dead letters
Canonical order model
The ERP’s order endpoint was not idempotent: submitting twice created two records. Every message carries an external identifier generated at the edge of the estate, and a deduplication table rejects replays before anything is written, retaining an audit record of the decision.
Multi-step flows needed different handling. Creating an order, reserving stock and requesting shipment can fail independently, so they run as an orchestrated state machine with compensating actions instead of one transaction across three systems. Stock moved to change-data-capture, replacing polling that read whole tables every fifteen minutes, and three years of history was backfilled under the same guarantees during trading hours. Schemas are versioned, with contract tests in continuous integration so a breaking change fails the build rather than production.
Each flow ran old and new in parallel, least consequential first, until reconciliation agreed across a full trading cycle. Cutovers happened in AWST evening windows, one flow switched at a time, the legacy route kept warm for four weeks before deletion.
Parallel running
Failure alerting
First-pass synchronisation settled above 99.8 per cent, all fourteen connectors are decommissioned, and about 26 hours of manual reconciliation returned to the business each week. Failures once found at month end now alert within minutes; two roles moved into customer service.
The setback came during the historical backfill. We had assumed the archive was internally consistent; instead some 1,900 orders existed in the CRM with no ERP counterpart, traceable to a connector that had dropped weekend batches for months. Rather than block the programme we added a quarantine queue with an operator workflow for records that cannot be reconciled automatically — a control nobody requested that is now considered essential.
“The difference is knowing, not speed. If something fails at two in the morning we hear at two in the morning, from a dashboard naming the orders affected. Learning about it a month later was the expensive part.”
We audit integration estates before quoting: what exists, what it costs to keep, and which links to retire first.