Skip to content
All insights

Data Readiness

"Our data isn't good enough for AI" is usually half true

The most common reason organisations delay is a fear about data quality. It's a reasonable fear, and it's wrong about as often as it's right.

NXTVIS Engineering7 min read

Almost every first conversation includes some version of this sentence, usually delivered apologetically: our data is a mess, we're probably not ready. It is worth taking seriously, because sometimes it is exactly right. But it is stated far more often than it is true, and the belief costs organisations years.

What actually blocks a model

There is a short list of genuine blockers, and 'messy' is not on it.

  • **The outcome was never recorded.** You cannot train a model to predict machine failure if nobody wrote down when machines failed. This is the real blocker, and it is a collection problem with a timeline, not a quality problem.
  • **The volume genuinely isn't there.** A few dozen examples of a rare defect will not support a reliable detector. Sometimes the honest answer is to wait and accumulate.
  • **There's no ground truth.** If nobody can say what the right answer was, supervised learning has nothing to learn from.
  • **It's legally inaccessible.** Consent, jurisdiction and contractual restrictions are real constraints and worth establishing early.

What people mean by 'messy', and why it rarely matters

Inconsistent formats, missing fields, duplicate records, free-text where there should be categories, three spellings of the same supplier name. All of this is normal, all of it is expected, and handling it is a routine part of the work rather than a reason not to start. Models are considerably more tolerant of noise than spreadsheets are, and every organisation's data looks like this, including the ones with impressive AI deployments.

Nobody has clean data. They have data, and a team that stopped apologising for it.

The test that settles it in an afternoon

Rather than debating readiness, pull a sample of a few hundred records of whatever you would want to predict, then check three things: is the outcome recorded, are the inputs present at the time the prediction would need to be made, and can a knowledgeable person confirm the outcome is correct. If yes to all three, you are ready, whatever the formatting looks like.

That last condition catches the subtle failure: data that includes information only available after the fact. A field populated by the person resolving the case is not available at the moment you need to predict the case, and a model trained on it will look brilliant in testing and fail completely in production.

And if you genuinely aren't ready

Then the correct project is instrumentation, and it is cheap. Start recording the outcome you eventually want to predict. In twelve to eighteen months you will have a dataset nobody can buy and a model that is straightforward to build. Organisations that do this deliberately end up years ahead of the ones that waited for the technology to get easier.

Establishing which of these two situations you are actually in takes us a few days and costs nothing. It is one of the more common outcomes of the audit, and one of the more valuable, because it converts a vague fear into a dated plan.

Want this looked at in your own systems?

The audit is free, the report is yours to keep, and there is no obligation to build anything with us.