The model is rarely the constraint
Capable models are cheap and largely interchangeable, which is convenient and slightly misleading. The difficult part has not moved: deciding what data a system may see, what it is allowed to conclude, and how a person reviews the result before it reaches a customer or a regulator.
When these projects stall they stall for unglamorous reasons — no authoritative record for an asset, rules that live only in the heads of two operators, documents scattered across drives and email. None of that is fixed by changing provider.
What actually blocks mid-market teams
No master data
One site appears under four spellings across three systems, an asset has two identifiers, a customer was duplicated during a migration. Without a canonical identifier every system agrees on, output is confident about entities that do not exist — the same problem underneath integration debt.
Undocumented rules
A model asked to classify or recommend must be judged against something. If the rules determining a correct outcome were never written down, you cannot build an evaluation set, so you have automated a guess and lost the ability to audit it.
Fragmented sources
Job detail in the field app, commercial terms in the ERP, correspondence in email, evidence in PDFs. Assembling one coherent view is the same read-path engineering any serious reporting project needs, and it is most of the effort.
No access-control story for documents
This is what stops projects late, after money is spent. Business documents carry implicit restrictions, and a system that retrieves text without respecting them will show something to someone who should not see it. Permissions must be resolved at retrieval against the existing access model.
If you cannot explain where data came from and who was entitled to see it, you do not have an AI project, you have an exposure.
A readiness checklist
- Canonical identifiers for site, asset, customer and job, each with a documented source of truth.
- Lineage: for any field used in a decision, its producer, last change and transformation.
- Retention and deletion enforced, including for derived artefacts such as extracted text and caches.
- Permissions resolved at retrieval, mirroring the application's role model.
- An evaluation set of genuine historical examples with the correct outcome recorded.
- Human-in-the-loop design: a reviewer, a queue, and source evidence shown beside the suggestion.
Near-term value is extraction, classification and reconciliation
This is where we part company with most of the marketing. The highest-return applications we see are not conversational: they pull structured fields from invoices, dockets and correspondence, classify inbound documents into the right workflow, and reconcile two sources that disagree before handing the difference to a human.
These succeed because they have a verifiable correct answer, a reviewer who already owns the exception queue, and a measurable before-and-after in hours. A general assistant over company knowledge has none of those properties. Document-heavy evidence work like the compliance reporting automation case is a better first target.
Australian privacy expectations and residency
The Privacy Act and the Australian Privacy Principles govern collection, use, storage and disclosure, and apply to many businesses that never saw themselves as handling sensitive data. The Notifiable Data Breaches scheme means an incident involving personal information carries an assessment and possible notification duty, so sending data offshore for inference must be a considered decision.
Residency is the second constraint. Resources, utilities, healthcare and government-adjacent operators often must keep data in-country, which means choosing regions deliberately and knowing where inference runs and logs are retained. That is the substance of security by design, with cost implications in cloud cost governance.
Choosing a first project
Pick something with a bounded document set, a clear correct answer, an existing reviewer and no customer-visible surface. Record the manual baseline first. If the pilot cannot show improvement on the evaluation set and in reviewer queue time, stop: the constraint is upstream in the data.
Work out whether you are ready
Bring us the process you want automated. We will say which checklist items you fail, what fixing them costs, and whether to start this year.