The tolerated pain is treating data as available when it is not fit for purpose
A business can have large volumes of data and still have weak inputs for AI. Customer records may contain duplicates. Product names may vary between systems. Historical decisions may have been recorded inconsistently. Important fields may be missing or based on definitions that changed over time.
None of this stops an AI project from running. That is what makes the problem dangerous. A model can produce an answer from poor inputs with the same confidence and speed it produces an answer from good ones.
Why businesses tolerate the problem
AI projects are visible and exciting. Data cleanup is slower, less glamorous and often owned by nobody. Teams therefore start with the tools they can buy rather than the information conditions the use case requires.
NIST's AI Risk Management Framework treats validity and reliability as core trustworthiness characteristics. NIST's work on qualifying data for AI also stresses understanding data source, collection method, coverage, variance, integrity and sensitive information. The implication is practical: data quality is not an abstract data-team concern; it is part of whether the AI system is fit for its intended use.
How the pain becomes money
- Decision error: poor inputs can distort prioritisation, classification, forecasting or recommendations.
- Rework: employees spend time checking, correcting and overriding outputs that cannot be trusted.
- Wasted AI spend: licences, model usage and implementation effort are consumed before the underlying information problem is solved.
- Remediation cost: bad outputs that reach customers or operations may require correction, communication and review.
- Low adoption: staff stop using AI when repeated errors destroy confidence, leaving the organisation paying for capability it does not trust.
Measure data fitness against the use case
There is no universal data-quality score that makes a dataset “AI ready.” Fitness depends on the decision being supported. A missing field may be irrelevant for one use case and critical for another.
Before deployment, define which data fields matter, where they originate, who owns them, how often they change and what error level the business can tolerate. Test the AI against known cases and track where poor data rather than model behaviour causes failure.
AI rework cost
outputs requiring correction × average correction time × loaded labour cost
Data-remediation effort
records requiring correction × average remediation time × loaded labour cost
These are operational measures, not universal benchmarks. They show whether AI is reducing work or merely relocating it into verification and cleanup.
Faster is not the same as better
Automation changes the economics of error. A person may make a bad decision slowly and locally. An AI-enabled workflow can apply the same flawed input across hundreds of records quickly. That can improve productivity when controls are sound; it can also scale mistakes when they are not.
For high-impact use cases, validation and human review should reflect the consequence of being wrong, not the novelty of the technology.
When this becomes commercially urgent
Priority is high when AI outputs influence pricing, customer eligibility, employee decisions, financial forecasts, operational scheduling or other material outcomes; when source systems disagree; when staff maintain unofficial spreadsheets to correct official data; or when AI pilots repeatedly fail because “the data isn't ready.”
What good looks like
A strong AI use case has defined source systems, documented data owners, measurable quality checks, traceable provenance, test cases, monitoring and a feedback loop for correcting both data and model behaviour. The organisation knows which data the system depends on and what happens when that data is missing or wrong.
Practical next actions
- Choose one AI use case and identify every source field it depends on.
- Document the authoritative source and owner for each critical field.
- Measure missing, duplicate, inconsistent and stale records.
- Create a representative test set with known expected outcomes.
- Separate model errors from source-data errors during evaluation.
- Define human review thresholds based on business consequence.
- Fix recurring data-quality causes upstream instead of correcting outputs forever.
Bottom line
AI can amplify the value of good information, but it also amplifies the consequences of weak information. Before scaling an AI workflow, prove that the data supporting the decision is fit for the job. Otherwise the business risks automating uncertainty.
Sources and further reading
- NIST AI RMF: Generative AI Profile.
- NIST AI Resource Center: AI risks and trustworthiness characteristics.
- NIST: Qualifying data for AI use.