AI Readiness & Strategy
The AI Problem Nobody Talks About: It's the Data, Not the Model
Enterprise AI fails on ungoverned data, not model choice. How to sequence the data foundation before the AI layer, and three cases where you should not.
Enterprise AI initiatives usually stall on data, not on model choice. Governance-first sequencing is the fix: repair the processes and pipelines producing bad records, rationalise the application estate, then layer AI on top. Deployed the other way round, an agent reproduces ungoverned data as fluent, confident answers rather than as errors, so the failure goes unnoticed.
Every enterprise conversation about AI right now sounds roughly the same. A leadership team gets excited about agentic workflows, GenAI copilots, or predictive automation. A vendor gets brought in. A pilot launches. And within a few months, something quietly stalls.
The numbers bear the pattern out. McKinsey's 2025 global survey found 88% of organisations using AI in at least one function, but only about 7% reporting it fully scaled. MIT's 2025 study of enterprise generative AI put roughly 95% of pilots at no measurable impact on the P&L. Adoption is close to universal. Returns are not.
The instinct is to blame the model. Wrong architecture, wrong provider, wrong use case. In our experience running these engagements, the model is rarely the actual problem. The data underneath it is.
What governance-first actually means
What this means for you: if governance in your organisation is a document rather than a running process, nothing built on top of it inherits any assurance.
Data governance is the set of ownership rules, quality standards, access controls and lineage records that decide whether a given dataset can be trusted for a given decision. Used loosely, the word describes a policy document, a committee, or an annual audit. Used precisely, it describes machinery that runs continuously and produces evidence.
The distinction matters because AI consumes the second kind and is indifferent to the first. A policy stating that customer records must be accurate has no effect on an agent reading those records. A validation gate that rejects a batch, routes the incident to a named owner, and records what was rejected does.
| Decision point | Governance bolted on afterwards | Governance built in from day one | What it costs you if you skip it |
|---|---|---|---|
| When it starts | After the pilot is live | Before any automation is deployed | Rework of every workflow already built on the bad foundation |
| What it produces | A compliance artefact | Validation gates, routed incidents, lineage records | No evidence to show a regulator or a board |
| Who owns "correct" | The technical team, by default | Business-side data stewards, by name | Definitions nobody can defend when they are challenged |
| How long it holds | Until the next clean-up | Continuously, as new sources arrive | Drift returns within months of the last remediation |
| Effect on the AI layer | Confident output over untrusted input | Output traceable to governed input | Answers you cannot act on and cannot audit |
Governance-first is a sequencing claim rather than a technology claim. It says the foundation is built and running before the automation that depends on it is deployed, because a layer that is not load-bearing cannot be made load-bearing retrospectively without redoing the work above it.
The pattern we kept seeing
What this means for you: an agent will not surface your data problems as errors, it will surface them as answers, which is considerably harder to notice.
Across client after client, the same failure mode showed up. Teams would deploy an AI agent or an automation layer on top of data that was inconsistent, ungoverned, or simply untrustworthy. The agent would perform exactly as designed, which meant it faithfully reproduced whatever mess was already sitting in the underlying systems.
This isn't a technology failure. It's a sequencing failure. Governance was treated as a compliance checkbox to tick after the AI was already live, instead of a foundation to build before it. The controls themselves are covered in the governance work that has to come first. This is the order those controls have to arrive in.
Bad inputs, confidently wrong outputs.
Why we built DQ Sentinel
What this means for you: the order in which you spend the budget matters more than which model you eventually put on top of it.
DQ Sentinel is our data quality and governance accelerator: automated validation, anomaly detection against learned baselines, and quality incidents routed to a named owner in the business. It is our answer to that pattern, built to embed governance and compliance principles from day one of an engagement, not bolted on at the end once something has already gone wrong. It is deployed inside a client's environment as part of an engagement rather than a product bought off a shelf.
In practice, that means broken processes and the data plumbing underneath them get identified and fixed before any automation or agentic workflow gets deployed on top. Application portfolios get rationalised first, so it is clear which systems in a client's landscape actually deliver value and which should be retired. Only once that foundation is clean does AI and ML capability get layered on top of it.
The sequence matters more than the technology choice.
A modern, well-governed data foundation with a modest AI layer will outperform an advanced model sitting on top of chaotic data.
Bolted on afterwards, the failure looks the same in every estate. Three systems hold three versions of the truth, nobody can produce an agreed customer count, and the technical team owns the pipes without owning what the numbers mean. The pilot does not fail loudly. It stalls because nobody trusts it enough to act on it.
Governance is a process, not a project
What this means for you: if your data quality work has a completion date, it has a re-infection date too, and it usually falls a few months later.
One mistake we see often: treating data quality as something you fix once and move past. It isn't. Data drifts, systems change, new sources get added, and without the right processes running continuously, the same mess creeps back in within a few months of a clean-up.
DQ Sentinel is built around that reality. It puts ongoing processes in place, not a one-time audit, so data accuracy is maintained as new information flows in rather than re-established every time something breaks. The mechanics of that sit in continuous monitoring of freshness, volume, schema and distribution, and in enforcing the rules in the pipeline itself rather than in a review meeting.
The other piece that gets missed is ownership. Technical teams can build the governance infrastructure, but they can't be the ones deciding what "accurate" or "correct" means for a given business dataset. That has to sit with the business itself, through named domain owners rather than a platform team by default. This is where data stewards come in: business-side owners accountable for the quality and integrity of the specific data domains they know best.
Without stewards in place, governance stays a technical exercise disconnected from the business context that actually gives the data meaning.
What this looks like for a client
What this means for you: the work that decides whether your AI programme survives is rarely the work anyone wants to present at a steering committee.
For banking and financial services clients, this often means DQ Sentinel is doing quiet, unglamorous work: reconciling data lineage, tightening access controls, and making sure regulatory reporting requirements are baked into the data layer itself rather than patched on afterwards. For manufacturing clients sitting on years of fragmented IoT and operational data, it means establishing enough structure and trust in that data before any predictive or agentic capability gets built on top.
Before an AI layer goes near a bank's data, three things have to hold in the data layer itself. Every figure has to trace back to the record it came from, which is what lineage reconciliation means in practice. Access has to be enforced in the data rather than in the application above it. And the reporting obligation has to be built into the layer rather than rebuilt by hand each cycle. For manufacturing the sequence is identical and only the starting mess differs: one schema, one clock and one unit of measure over years of sensor history, before any forecasting is built on it.
Neither of those is the exciting part of an AI transformation story. Both of them are the part that determines whether the exciting part actually works six months later.
When governance-first is the wrong call
What this means for you: sequencing is a default rather than a law, and there are cases where insisting on the full foundation first costs more than it protects.
There are three situations where we would not start here.
The first is a genuinely contained pilot on a single, already-trusted dataset, where the goal is to learn whether a capability is viable at all. Governing an estate to answer a question that a two-week experiment on one clean table can answer is misallocated effort. Keep it contained, keep it away from any system of record, and do not let it quietly become production.
The second is a regulatory or commercial deadline that will not move. If a reporting obligation lands in eight weeks, the correct sequence is to meet the obligation and schedule the foundation work immediately behind it. Naming that as a deliberate deferral, with a date, is very different from never doing it.
The third is an estate already mid-migration. Rationalising an application portfolio that is being restructured anyway means governing systems that are about to be retired. Sequence the governance work to land with the target state rather than the one being dismantled.
None of these is a licence to skip the foundation permanently. Each is an argument about when, not whether.
Three questions to ask before funding an AI initiative
What this means for you: you can test the argument in this article against your own estate this week, without engaging anyone.
Take the AI initiative that stalled, or the one you are about to fund, and work backwards through three questions.
Who is accountable if a number is wrong?
Ask it of every dataset the initiative reads. If the honest answer is the platform team, you have found a missing steward rather than a missing model.
What happens when a dataset arrives malformed?
If a consumer notices eventually, the quality machinery is reactive. A [data quality SLO](/glossary#data-quality-slo) is what makes it proactive.
Are these systems ones you intend to keep?
Governing a source you plan to retire within the year is effort you will spend twice. Rationalise before you govern.
Those three answers are usually enough to tell you whether the estate is ready for an AI layer at all. If it is, the next question is architectural, and it is set out in the four layers autonomous execution depends on.
If an AI initiative inside your organisation has stalled, or never produced the results the pitch deck promised, the honest first question isn't "do we need a better model." It's "can we actually trust the data this model is working from."
Most of the time, the answer explains everything else.
Frequently Asked Questions
Why do enterprise AI pilots stall so often?
In our engagements the common cause of a stalled enterprise AI pilot is the data underneath, not the model. An agent deployed over inconsistent or ungoverned records reproduces those problems as fluent, confident answers rather than as errors, so the failure is hard to spot. McKinsey's 2025 global survey found 88% of organisations using AI somewhere but only about 7% reporting it fully scaled. That distance is largely a foundation problem.
Should data governance come before or after an AI deployment?
Data governance should come before an AI deployment in almost all cases. Governance treated as a compliance step after go-live produces an artefact rather than working machinery, and every workflow already built on the untrusted foundation has to be reworked once the gaps surface. Building it in first means validation gates, routed incidents and lineage records exist before anything automated depends on them.
What is a data steward and why does governance fail without one?
A data steward is a business-side owner accountable for the quality and integrity of a specific data domain. Technical teams can build governance infrastructure, but they cannot decide what accurate means for a dataset they do not own. Without a named steward, nobody can settle a disagreement between two systems, so the definition stays contested and every number built on it inherits that.
Why does a one-time data clean-up not hold?
A one-time data clean-up does not hold because data drifts. Systems change, new sources are added, and definitions move with the business. Without processes running continuously, the same problems return within a few months of a remediation. Governance has to maintain accuracy as new information arrives rather than re-establish it each time something visibly breaks.
Is there ever a case for deploying AI before the governance work?
There are three narrow cases: a contained experiment on a single already-trusted dataset, a fixed regulatory deadline that will not move, or an estate already mid-migration. Each is an argument about sequence rather than necessity. Each should carry a stated condition for when the foundation work begins, whether that is a date, a target state, or the moment the experiment stops being contained.
Unolabs is a Data and AI first engineering consultancy, headquartered in the United Kingdom with engineering operations in Pune and active engagements across the UK, Australia, and Hong Kong. We help enterprises build the architectural foundation for autonomous AI execution — governed data platforms, semantic intelligence, and agentic systems that enterprises can stand behind.
If you are weighing whether to fix your data foundation before funding another AI pilot, book a discovery call and we will work through it with you.
Continue reading
- Data Governance Trends 2026AI-Ready Data Foundations: The Governance Work That Comes FirstAI-ready data foundations decide whether models scale: quality at source, resolved entities, lineage to prediction, and inventories a regulator can audit.11 min read
- Enterprise AI ArchitectureThe Autonomous Enterprise Architecture: Four Layers, and Why AI Programmes Stall Without ThemMost enterprise AI programmes stall on architecture, not models. The four layers of autonomous enterprise architecture, and why sequence decides results.20 min read
- Risk & ComplianceData Observability & Quality ManagementData observability explained: five monitoring dimensions, three pipeline quality gates, a worked data contract in YAML, and a four-phase rollout plan.14 min read
Find out what your data layer would fail on first
Bring the initiative that stalled. We will trace it back through the data that feeds it and show you which definitions are contested, which lineage is unreconciled, and which datasets nobody owns.
Assess Your Data Trust