Case study · Integration

Fourteen connectors, one governed hub

A national wholesale distributor moved orders across four systems using fourteen hand-built connectors. We built one broker-led hub — idempotent, reconciled, alerting — without replacing an ERP that could not be modernised.

Client: National wholesale distributor (confidential) Sector: Wholesale distribution Duration: 9 months Status: In production since 2023 Team model: Dedicated in-house engineers

Every order crossed four systems: an ERP commissioned long before anyone expected an online channel, a CRM holding pricing and credit terms, a third-party logistics warehouse interstate, and a web store trading around the clock. Fourteen connectors joined them — cron scripts, exports and an FTP job untouched for years.

The hub substituted one publish-and-subscribe layer for all fourteen: a canonical model covering orders, customers, stock and shipments; publishers in Python, consumers in Node.js; RabbitMQ between them; and a reconciliation layer reporting stalled records within minutes. The ERP, CRM and 3PL stayed.

01The estate we inherited

The connectors worked individually and failed collectively. Each assumed its own timing, keys and failure handling, so the consequence of one link breaking was unpredictable — and several failed quietly, surfacing during reconciliation rather than trading.

  • Nightly CSV drops with no receipt or failure notice
  • Retries that duplicated rather than rejected
  • Stock diverging between ERP and the web store, causing oversell
  • No idempotency: a timeout looked identical to a duplicate
  • Two staff reconciling differences found after month end

02Non-negotiables

The boundary conditions were clear, and two decided the architecture: order intake could not pause, and no duplicate financial document could reach the ERP. An accounts trail nobody trusts is worse than a delay.

  • The ERP exposed only a dated SOAP service and flat-file export
  • The 3PL accepted nothing but legacy FTP files
  • Order intake ran continuously, no maintenance window
  • Zero tolerance for duplicated invoices, credits or stock movements
  • Cutover had to be incremental; a single switch-over was refused

03What we built

One message hub instead of a mesh. Each system speaks to the hub, never directly to another, and every message is validated against a versioned schema before publication. Consumers are individually replaceable — which is what retiring a connector now means.

  • A canonical model: orders, customers, stock, shipments
  • Python publishers, Node.js consumers, RabbitMQ broker
  • Idempotency keys and deduplication on every message
  • Dead-letter queues and retries using exponential backoff
  • Change-data-capture readers replacing full-table polling
  • Reconciliation dashboards and support alerting

04Engineering the awkward parts

The ERP’s order endpoint was not idempotent: submitting twice created two records. Every message carries an external identifier generated at the edge of the estate, and a deduplication table rejects replays before anything is written, retaining an audit record of the decision.

Multi-step flows needed different handling. Creating an order, reserving stock and requesting shipment can fail independently, so they run as an orchestrated state machine with compensating actions instead of one transaction across three systems. Stock moved to change-data-capture, replacing polling that read whole tables every fifteen minutes, and three years of history was backfilled under the same guarantees during trading hours. Schemas are versioned, with contract tests in continuous integration so a breaking change fails the build rather than production.

05Sequencing the cutover

Each flow ran old and new in parallel, least consequential first, until reconciliation agreed across a full trading cycle. Cutovers happened in AWST evening windows, one flow switched at a time, the legacy route kept warm for four weeks before deletion.

  • Shadow mode: the hub processed live events and compared silently
  • Legacy routes reversible for four weeks after each switch

06Outcomes, and one honest setback

First-pass synchronisation settled above 99.8 per cent, all fourteen connectors are decommissioned, and about 26 hours of manual reconciliation returned to the business each week. Failures once found at month end now alert within minutes; two roles moved into customer service.

The setback came during the historical backfill. We had assumed the archive was internally consistent; instead some 1,900 orders existed in the CRM with no ERP counterpart, traceable to a connector that had dropped weekend batches for months. Rather than block the programme we added a quarantine queue with an operator workflow for records that cannot be reconciled automatically — a control nobody requested that is now considered essential.

99.8%
First-pass sync success
14
Connectors decommissioned
26 hrs
Reconciliation removed weekly
6 min
Mean time to detect
“The difference is knowing, not speed. If something fails at two in the morning we hear at two in the morning, from a dashboard naming the orders affected. Learning about it a month later was the expensive part.”
Head of Operations — national wholesale distributor

07Technology used

PythonNode.jsRabbitMQAWS ECSAmazon RDSTerraformpytestContract testsCloudWatch

Counting the cost of your connectors?

We audit integration estates before quoting: what exists, what it costs to keep, and which links to retire first.

Request a Quote