Every AI system is limited by the data it learns from or reads. A forecasting model trained on patchy sales history will produce patchy forecasts. A document assistant pointed at a shared drive full of outdated policies will confidently quote outdated policies. Checking AI data readiness before a project starts is the cheapest way to avoid months of disappointment. This article gives you seven questions to ask, with simple checks you can run, and what to do when the answers are not encouraging.
Different AI needs different data
"Data" means different things depending on the kind of AI you plan to use:
- Predictive machine learning (forecasting demand, predicting churn, scoring leads) learns from historical, structured records, usually in databases, with known outcomes.
- Document and language AI (assistants that answer from your documents, chatbots) depends on a well-organised, current, accessible collection of text.
- Computer vision (defect detection, counting items) needs many example images, labelled consistently.
Keep your intended use in mind as you go through the questions; not all apply equally.
Question 1: Does the data exist, and can you find it?
List the data the use case requires and where each piece lives: which system, which tables or folders, who owns it. Gaps often appear immediately. Perhaps the reason a customer cancelled was never recorded, or delivery dates exist only in a courier's portal. If a crucial input is not captured anywhere, the first step is to start capturing it.
Question 2: Is there enough of it?
There is no universal number, but some rules of thumb help. Forecasting needs enough history to see seasonal patterns repeat, ideally more than one full cycle. Classification needs enough examples of each category, especially the rare ones you care about most. If you want to predict a type of event that happened only a handful of times, no model will learn it reliably.
-- How much history and how many examples of each outcome?
SELECT MIN(created_at) AS first_record,
MAX(created_at) AS last_record,
COUNT(*) AS total
FROM support_tickets;
SELECT resolution_category, COUNT(*) AS examples
FROM support_tickets
GROUP BY resolution_category
ORDER BY examples;
Question 3: Is it accurate and complete?
Measure, do not assume. Check the share of missing values in fields that matter, impossible values, and duplicates:
SELECT
COUNT(*) AS rows_total,
AVG(customer_id IS NULL) AS pct_missing_customer,
AVG(delivered_at IS NULL AND status = 'delivered') AS pct_delivered_no_date,
AVG(delivered_at < shipped_at) AS pct_delivered_before_shipped
FROM shipments;
(In MySQL, a true comparison counts as 1, so AVG gives a fraction; in PostgreSQL, cast with ::int.) Also look for silent changes over time: a field whose meaning changed when a system was upgraded, or a period when data entry was skipped. Models trained across such breaks learn confused patterns.
Question 4: Are outcomes labelled consistently?
Supervised machine learning needs labels: the known answer for each historical example, such as "this ticket was a billing issue" or "this invoice was paid late". Labels are often the weakest part of business data. Different staff use categories differently, "other" becomes a catch-all, and categories are renamed mid-year. Pull a random sample of a hundred records and have two people label them independently. If they disagree often, the model will be learning from noise.
Question 5: Does it reflect the situation the AI will face?
Historical data reflects historical conditions. If you changed pricing, entered a new region or redesigned a product line, older data may mislead. Watch for bias in the data too: if past decisions were unfair to certain groups, a model trained on them will learn the same pattern. For document collections, check how much content is outdated, duplicated or contradictory, and whether there is a clear "current version" of each policy.
Question 6: Can systems actually get to it?
Data that exists only in a legacy system with no export, in scanned paper files, or in individuals' email inboxes is not ready, however good it is. Check:
- Can it be extracted automatically and repeatedly, not just once by hand?
- Can records from different systems be linked by a reliable shared identifier?
- How fresh will it be when the AI uses it: real time, daily or monthly?
- Will the live data the model sees in production be in the same format as the training data?
Question 7: Are you allowed to use it this way?
Data collected for one purpose cannot always be used for another. Consider what customers were told when their data was collected, contracts with partners and suppliers, confidentiality obligations, and the data protection laws that apply to you, which differ by country. If an external AI service will process the data, check its terms on retention and training use. Remove personal data the use case does not need.
A simple AI data readiness scorecard
| Area | Ready | Needs work | Blocker |
|---|---|---|---|
| Existence | All key inputs captured | Some inputs partial | Key input never recorded |
| Volume | Ample history and examples | Thin for rare cases | Too few examples to learn from |
| Quality | Few gaps or errors | Fixable gaps | Mostly missing or unreliable |
| Labels | Consistent and agreed | Some inconsistency | No outcomes recorded |
| Access | Automated extraction possible | Manual or partial | Locked in inaccessible systems |
| Rights | Use is clearly permitted | Needs review or consent changes | Use not permitted |
A single blocker usually means pausing the AI project and fixing that first.
If your data is not ready
That is a common and useful finding, not a failure. Practical next steps:
- Start capturing missing data now, with validation at entry.
- Clean and deduplicate the most important records.
- Agree category definitions and relabel a sample.
- Consolidate data into a reporting database or warehouse with reliable links between sources.
- For document AI, retire outdated content and assign owners to each collection.
- Meanwhile, consider tasks that need less historical data, such as document extraction with general-purpose models.
Our database management team handles data consolidation and cleanup, and our AI and machine learning development team can run a readiness assessment for a specific use case.
Key takeaways
- AI data readiness depends on the type of AI: structured history, document collections or labelled images.
- Check existence, volume, quality, labels, representativeness, access and usage rights.
- Run simple SQL checks and a labelling agreement test before committing budget.
- If the data is not ready, fixing it is the most valuable first project.