Somewhere in most organizations there is a slide that says the data isn’t ready. It appears whenever someone proposes an AI project, it is never wrong exactly, and it stops the conversation cold.
The trouble is that “ready” has no meaning at the organizational level. Data is ready for something. A model that drafts a first-pass response to a customer question needs almost nothing of what a model that prices a deal needs. Asking whether your data is ready, unqualified, produces either a two-year remediation program nobody funds or a shrug. Both outcomes end the project.
A more useful question: for this specific use, what would have to be true about the data for the output to be trustworthy? That question has an answer, and the answer is usually much smaller than the remediation program.
Four conditions, assessed per use
Four conditions decide it, and they decide it for one defined use at a time.
Ownership. Who decides what this data means? Not who administers the database — who adjudicates when two systems disagree about a customer’s name, or what counts as an active matter. If nobody can answer that, the AI output will inherit the ambiguity and present it with more confidence than the source deserved. Unowned data does not become owned because a model consumed it.
Quality, defined against the use. Quality is not an abstract score. A retrieval system over policy documents can tolerate inconsistent formatting far more easily than it can tolerate three versions of the same policy with no reliable way to tell which one is current — the wrong-but-confident answer comes from the version conflict, not the formatting. A use built on structured, numerical history usually has the opposite problem: it needs the fields consistent and parseable, but a well-dated series does not suffer the same “which one is current” ambiguity a document does. Write down what “good enough” means for the specific use before measuring anything, or you will measure the wrong thing thoroughly.
Access. Can the system reach this data, under what identity, and with what permissions? This is where AI projects most often meet the real constraint. The data exists, the quality is fine, and it sits behind an access model that either cannot be extended to a service account or — much worse — can be extended so easily that the AI system ends up seeing more than any individual user should. Retrieval systems are very good at surfacing the exact document somebody forgot to secure.
Integration. How does the data get from where it lives to where the model needs it, how often, and what happens when that breaks? An impressive pilot built on a one-time export is not a system. It is a demonstration with a half-life.
Unstructured data counts, and usually dominates
Discussion of data readiness drifts toward databases, because databases are what data governance has historically meant. But most of what makes an AI system useful inside a professional-services firm is unstructured: documents, email, knowledge collections, prior work product, file shares that have accumulated for twenty years.
That material has the same four conditions, in less tractable form. Ownership is murkier. Quality is largely a question of whether anyone can tell current from superseded. Access is genuinely hard, because document permissions are usually set item by item and drift out of sync with who actually needs to see what. And integration means deciding what gets included in a retrieval collection and what gets kept out — a decision with real consequences that is frequently made by whoever happened to set up the pilot.
Preparing document collections properly — deciding scope, respecting the existing permission model, keeping the collection current as the source changes — is most of the work in a retrieval system. It is also the part that is easy to mistake for a week of tidying, rather than the substance of the work.
Better data does not guarantee better output
Worth stating plainly, because the opposite is widely implied: fixing your data does not make a language model reliable. It removes one category of failure — the model confidently repeating something wrong, stale, or ambiguous in the source. It does not remove the model’s capacity to be confidently wrong on its own, even about data that is well governed. Human review where the stakes justify it, evaluation against cases you actually care about, and a clear account of what the system is not for remain necessary regardless of how good the inputs are — the same point sits at the center of the NIST AI Risk Management Framework, which treats data quality as one input to trustworthy AI rather than the whole of it.
Data readiness is a prerequisite, not a guarantee. Anyone selling it as a guarantee is selling something else.
Where this leaves you
If a use is stalled behind a data problem, the productive move is to name the use, name which of the four conditions actually fails for it, and size the work of fixing only that. Sometimes the honest answer is that this particular use is not worth what it would cost to make it trustworthy — which is a perfectly good outcome, arrived at in weeks rather than after a failed implementation.
What you want to avoid is the slide. It is not that it is wrong. It is that it is unfalsifiable, and unfalsifiable objections do not get resolved. They get inherited by whoever proposes the next project.