Skip to main content
Back to Insights
Case-Study · 4 min read

Global Beverages: replacing a 70-FTE harmonisation process with a governed lakehouse

A global beverages leader ran master data harmonisation through a 70-FTE manual team across SAP, Salesforce and legacy estates. Unolabs replaced it with a Databricks Medallion lakehouse combining fuzzy matching and LLM-driven classification, backed by verifiable matching logs. The engagement validated £3.3M in operational savings under a shared-risk commercial model.

Insight at a glance

A concise view of impact and engineering focus.

Outcome

£3.3M validated savings

Outcome

70-FTE manual process automated

Outcome

40% faster time-to-analytics

Section 1

What a 70-FTE harmonisation team actually costs

A global beverages leader faced a systemic operational bottleneck: a 70-FTE team dedicated to manually harmonising fragmented master data across SAP, Salesforce and multiple legacy estates. Every record that moved between systems passed through human judgement — matched, classified and reconciled by hand.

At that scale the process was expensive, but headcount was the visible problem rather than the important one. Manual matching is error-prone and inconsistent between operators, and its throughput sets a hard ceiling on time-to-insight for everything downstream. Two operators looking at the same ambiguous pair of records will not always reach the same answer, and nothing in the process records which of them was right.

The firm's ambitions for autonomous supply chain orchestration made the constraint acute. Architectural agility of that kind cannot be built on top of a harmonisation layer running at the pace of manual review, and no amount of investment further downstream will move a ceiling that sits underneath it.

Engineering note

When 70 people sit between source systems and analytics, the headcount cost is the visible problem — the invisible ones are inconsistency between operators and a throughput ceiling on every downstream initiative.

Section 2

Two matching tiers, and why not everything went through the LLM

We replaced the manual harmonisation layer with an engineered AI lakehouse built on a Databricks Medallion architecture, promoting records through bronze, silver and gold layers with quality enforced at each stage. An automated ingestion factory integrated SAP ECC and Salesforce feeds, removing the manual collection step entirely.

The matching engine combines two techniques deliberately. Fuzzy matching handles the high-volume cases where records differ by formatting, spelling or structure. LLM-driven reasoning is reserved for the genuinely hard ones — complex master data classification and matching decisions that previously required human judgement. Sending every record through the expensive path would have been simpler to build and worse in every other respect.

Everything was delivered as governed data-as-a-product, replacing manual operational silos with owned, documented, quality-assured outputs. Commercially, the engagement ran on a shared-risk model tied to validated operational cost reduction, which meant the engineering targets and the commercial targets were the same numbers. That structure removes the usual end-of-programme argument about whether a benefit was real.

  • Databricks Medallion architecture for high-fidelity data processing
  • Automated ingestion factory integrating SAP ECC and Salesforce telemetry
  • LLM-driven reasoning for complex master data classification and matching
  • Governed data-as-a-product design replacing manual operational silos
  • Shared-risk commercial model tied to validated operational cost reduction
Engineering note

The two-tier matching design matters: fuzzy matching handles volume cheaply, and the LLM layer is reserved for classification decisions that genuinely need reasoning — running everything through an LLM would be slower, costlier, and harder to audit.

Section 3

Why the automation had to be more inspectable than the people

The solution delivered an 'Automated Truth Layer' for the global enterprise. Metadata-driven governance is embedded in the lakehouse itself, and every match and classification decision is written to verifiable matching logs as it is taken.

Any harmonised record traces back to the evidence and the logic that produced it, which is what makes the automation auditable rather than merely fast. That auditability was the condition for trusting the output at all: downstream teams and auditors can inspect why two records were merged or how a product was classified, without taking the engine's word for it.

The bar here is higher than the process it replaced. A human operator can be asked afterwards to explain a decision; a matching engine cannot, unless the explanation was written down at the time it was made. Building that logging into the engine rather than around it is why the harmonised estate is accountable as well as accurate.

Engineering note

Automation that replaces human judgement must be more inspectable than the humans were, not less — verifiable matching logs are what turn an AI matching engine into an auditable system of record.

Section 4

The outcome: harmonisation at pipeline speed

The transformation automated the previously manual harmonisation process end-to-end, removing routine manual intervention, and validated £3.3M in operational savings under the shared-risk commercial model.

Time-to-analytics improved by 40%, because harmonised data now flows at pipeline speed rather than review speed. The architecture serves as the trusted foundation for autonomous demand planning and global inventory orchestration across the beverages ecosystem — initiatives previously blocked on data readiness rather than on modelling capability. That is a common diagnosis in our engagements and an unpopular one, because it moves the fix away from the data science team.

Two caveats belong beside those figures. The savings were validated under the shared-risk commercial model agreed with the client rather than by an external audit, and the time-to-analytics improvement is measured against a baseline set by manual review — a baseline that was, by construction, the slowest thing in the estate. Both figures are honest; neither is portable to an estate that starts from a different place.

Engineering note

What we'd flag: end-to-end automation here means routine manual work was removed, not that human oversight disappeared — a system making master-data decisions with LLMs needs a standing exception path and audit review, and the verifiable matching logs exist precisely to support that.

Key Takeaways

What to carry into the next sprint

Back to all insights

Takeaway

Route volume through cheap matching; reserve LLM reasoning for decisions that need it.

Takeaway

Automation replacing human judgement must log its reasoning as it decides, not after.

Takeaway

Tie the commercial model to the same numbers the engineering targets are measured on.

Due diligence

Frequently asked questions

Why not run every match through the LLM?
Cost, latency and auditability. Fuzzy matching resolves the high-volume cases where records differ only by formatting, spelling or structure, and it does so cheaply and predictably. LLM-driven reasoning is reserved for complex master data classification and matching decisions that genuinely need it. Routing everything through a model would be slower, more expensive and considerably harder to inspect.
How do you audit a matching decision an AI system made?
By writing the decision down as it is made. Every match and classification in this beverages lakehouse is recorded to verifiable matching logs, so any harmonised record traces back to the evidence and the logic that produced it. Automation replacing human judgement has to be more inspectable than the people were, because it cannot be questioned afterwards.
Does end-to-end automation mean no human oversight remains?
No. It means routine manual intervention was removed, not that oversight disappeared. A system making master-data decisions with LLMs needs a standing exception path and an audit review rhythm, and the verifiable matching logs exist precisely to support both. Anyone reading 'end-to-end' as 'unsupervised' should confirm the exception handling before comparing this to their own estate.
How was the £3.3M in savings validated?
Under the shared-risk commercial model agreed for the engagement, which tied commercial outcomes to validated operational cost reduction and made the engineering and commercial targets the same numbers. It is not an externally audited figure, and it reflects the cost base of a 70-FTE manual harmonisation process specific to this estate. Ask us how that baseline was constructed.
What does a Medallion architecture add to master data work?
Staged quality. Records are promoted through bronze, silver and gold layers with enforcement at each stage, so a matching decision is taken against data that has already passed defined checks rather than against a raw feed. Combined with an automated ingestion factory across SAP ECC and Salesforce, it removes the manual collection step that preceded every match.
Related engineering assets

Cost out your manual harmonisation before you automate it

We will size the matching work your teams do by hand, separate the cases fuzzy logic resolves from the ones that need reasoning, and tell you what an auditable engine would actually replace.

Book a Harmonisation Review