Strategy
← All insights

Fix the Data for Everyone, Not Just for AI

The data you are worried about feeding to AI already feeds your board pack. Why "fix the data first" gets the scope and the sequencing wrong.

Assessment

COO, CTO, CIO, CFO

Part of our guide: AI transformation for COOs

Somewhere in your business this week a slide will say that the data has to be fixed before AI can start, and the room will nod, because it sounds like prudence. We think it misses something, and the something matters more than the slide.

The data everyone is worried about feeding to AI is the same data that feeds people today. It sits behind the board pack, the operational decisions and the month-end close, and it has been doing so for years. If it is not good enough for a model, it was never good enough for the people who have been working around it every Monday morning. The case for fixing it was there long before anyone mentioned AI.

Where the orthodoxy is right

We will give the orthodoxy its due, because part of it is sound. AI amplifies a data problem in two ways: it goes faster, and it does not pause where a person would have noticed an odd field and quietly worked around it. The bar rightly rises as the person steps further back from the decision, and the peer-reviewed evidence bears that out. In Organization Science, consultants working inside the tasks AI handles well finished faster and at higher quality, while outside that boundary AI users were 19 per cent less likely to produce a correct answer. A system that cannot tell when it has left the boundary needs better inputs than a person does.

That is an argument for fixing the data for every purpose, with the AI use cases setting the bar. It is not an argument for putting AI at the back of the queue until the whole estate is clean.

The disagreement is about scope and sequencing

The orthodoxy, taken literally, says all of the data has to be fixed before any AI starts. In practice the data gets fixed in stages, use case by use case, on the roadmap it always needed. The businesses that wait for the whole estate before starting anything are waiting for a day that does not arrive.

The number behind that is stark, because Dun and Bradstreet's July 2026 survey of 10,000 businesses found only 6 per cent saying their data is fully ready to support AI at scale. In the same survey 34 per cent said they were already scaling AI into production. The May wave of the same tracker listed the barriers in order, with limited data access at 50 per cent, privacy and compliance at 44 per cent and data quality at 40 per cent. Read those facts together and the gap between "fully ready" and "already running" is where most of the value is being created, by businesses that chose a narrow, well-governed set and started.

We would add one caution about the numbers you may have seen in support of the orthodoxy. The often-quoted claim that data preparation is 60 to 80 per cent of an AI project's cost has no primary source that we could trace. Every citation resolves to marketing content quoting a different range. We have set out which AI statistics survive checking, and that one does not.

Start the way your reports already start

We rarely ask people to work from raw data sources, and board reports and decision makers are instead pointed at a curated layer that somebody owns, that has been cleaned and reconciled and that the business already trusts. Data teams call the pattern bronze, silver and gold, meaning the raw feeds, the cleaned and conformed sets and the curated sets built for a purpose. Most mid-market businesses have some version of it whether or not they use the names.

AI can start the same way, whether the set a use case needs is a table of policies or a corpus of contracts, and the sequence is one your finance function would recognise.

Pick a use case where a wrong answer is recoverable and a person still owns the decision. That constraint does most of the governance work on its own, because it keeps the boundary the evidence describes inside a human's reach.

Build or reuse the curated set that serves it, on the platform you already have. A new platform is a reason to wait, and waiting is the failure mode we are trying to avoid.

Give that set an owner, access control and lineage. Those three are the whole of data governance at the scale of one use case, and they are far easier to establish for one curated set than for an estate.

Let each live use case show you which part of the wider estate to fix next. The roadmap writes itself from evidence rather than from a survey of everything that is wrong, and the curated set becomes the baseline you measure the results against.

Budget the data work as a line, not a gate

There is a cost reason to run it this way as well, because a survey of 396 organisations published in July 2026 found the top source of unexpected AI spend was data platform usage overages at 47 per cent. That put it ahead of model token costs at 43 per cent. Data work is where the money goes in any sequence, and treating it as a prerequisite gate hides that cost in a programme nobody has started. Treating it as a line on each use case keeps it visible, owned and measured. We covered the wider picture in the real cost of AI for a COO.

The cost of waiting

Organisations waiting for everything to be perfect are making perfect the enemy of good. They hand the advantage to competitors who started with a narrower, well-governed set, and they miss the learning that only comes from running something real, the learning that tells you what to fix next.

Fix the data because your people deserve it, and don't make AI wait for it.

Working out which curated set to start with, and what it will take to stand it up, is part of what a Breathe discovery sprint does. If your AI programme is currently parked behind a data clean-up with no end date, that is a conversation worth having.

Sources

Work with us

Breathe: Find your starting point

A focused sprint that maps your capabilities, scores AI opportunities and builds a three-year total cost of ownership view, so you know exactly where AI fits, what to do first and what it will cost.