JOURN3Y LogoJOURN3Y
Data Foundations
September 7, 2026

Why Your AI Project Is Really a Data Project

Back to Blog

There is a moment in most enterprise AI programmes where the conversation stops being about AI.

It usually arrives when two people look at the same output and disagree about whether it is right. Not because the model hallucinated — because one of them counts a customer differently from the other, and both have counted that way for years.

At that point the programme is no longer an AI project. It is a data project wearing an AI project's clothes, and it will not progress until that is acknowledged.

The three problems underneath

Almost every stall we see traces back to one of three things, and none is exotic.

The same word means different things. Finance counts a customer by billing entity, sales counts by relationship, and support counts by login. All three are correct within their own function. An agent asked "how many customers do we have" will pick one and sound certain.

The information is in a system nobody indexed. The answer exists, in a shared drive or a legacy database or somebody's inbox, but the AI layer cannot reach it. The output is not wrong so much as blind — and blindness is invisible in a demo built on curated data.

Nobody owns the definition. Not a technical problem at all. There is no forum where "what is an active account" gets settled, so it is settled independently in a dozen places.

Why this is worse with AI than with dashboards

These problems are not new. Organisations have lived with inconsistent definitions for as long as they have had reporting.

What changes is exposure. A dashboard is built by an analyst who knows the caveats, used by people who have learned them, and updated slowly. The inconsistency is contained by the small number of people involved.

An AI layer removes that containment on purpose. That is the point of it — anyone can ask anything, without an intermediary. Which also means the definitional mess is now available to everyone, at speed, phrased with total confidence and no caveats at all.

AI does not create the problem. It industrialises it.

What actually needs to be true

Less than people fear, and more specific than a data strategy.

The handful of terms that matter have one agreed meaning. Not every field in the warehouse. The ten or fifteen that appear in the questions people actually ask — customer, account, revenue, active, churn.

The systems where the work lives are reachable. Including the unglamorous ones. The exception is usually the system that matters most.

Permissions are modelled at the source. Not as a layer on top. What an agent can see for a given person should be exactly what that person could already see.

Somebody owns each definition. A name, not a committee. When the definition needs to change, there is a person who decides.

What this does not mean

It does not mean a two-year data programme before you are allowed to build anything. That advice is common, and it is how organisations spend a fortune arriving at a governed warehouse nobody has tested against a real question.

The sequencing that works is narrower. Pick the first job. Establish what that job needs — usually a handful of definitions and two or three systems. Fix those. Build it. The next job will share most of the foundation, which is why the second agent is so much faster than the first.

You do not need the whole house in order. You need the rooms you are about to use.

A test worth running

Ask four people in different functions the same question about your business. Something basic: how many customers do we have, or what did we sell last month.

If you get four numbers, you have not found a problem with your people. You have found the thing that will make your AI programme feel unreliable, before you have spent anything discovering it the expensive way.

Tags:
#DataFoundations#DataStrategy#EnterpriseAI#DataGovernance#AIStrategy