The question I hear most often when a business is getting ready to implement an AI workflow is some version of: "Our data isn't perfect — do we need to fix it all before we start?"
The answer is almost always no. But it comes with an important qualification.
You don't need perfect data. You need the right data to be clean. And those are very different things.
The practical definition of "good enough"
AI workflows act on specific data fields. A renewal reminder workflow needs the customer's name, their email address, and their renewal date. An invoice generation workflow needs the client name, the work completed, the rate agreed, and where to send it. An outreach sequence needs the contact's name, their company, and their email.
Those are the fields that matter for each specific use case. If those fields are accurate and current, the workflow will function correctly. If the industry sector field is incomplete, or the secondary phone number is missing, or the customer's Twitter handle was never added — for those workflows, that doesn't matter at all.
This is where I see a lot of businesses get stuck. They start thinking about data quality as a global problem — everything needs to be fixed before anything can start. That framing is both demoralising and inaccurate. The question isn't "is our data perfect?" The question is "for this specific workflow, are the fields it will read correct?" — the same lens I use in why AI implementations fail on data, not technology.
What I learned in financial services
At the large wealth management firm where I spent several years, we weren't trying to make every field in the client record perfect before we began any automation or profiling work. We couldn't have — the scale was too large and the legacy data too varied for that to be realistic.
Instead, we took a use-case-driven approach. For each initiative — client communications, reporting, regulatory submissions — we identified the fields that mattered for that specific purpose and focused quality effort there. Contact details for outreach. Account ownership for reporting. Advisor assignments for routing. We cleaned those fields, built processes to maintain them, and moved forward. The rest of the record was addressed over time.
That same pragmatic approach applies at any scale — and aligns with data governance obligations under the EU AI Act for high-risk use cases.
A readiness checklist for SMBs
Before starting any AI workflow, I work through four questions with clients:
1. What data will this workflow actually read or act on? Be specific. List the fields. A renewal reminder reads: customer name, email address, renewal date, contract value. List them out and be honest about whether each one is in a single, consistent place.
2. How complete and accurate are those fields right now? Not all fields — just these. Pick 20 records at random and check. Are the fields populated? Are they current? Do they match what you'd expect from other sources you have?
3. Is there one authoritative place for each field, or is it duplicated? If the customer's email address is in three different systems with three slightly different versions, that's the problem to fix — see single source of truth for how to establish one.
4. What's the process to keep those fields current going forward? Data quality isn't a one-time fix. A workflow built on clean data today, with no process for maintaining that data, will degrade over time.
When to fix first, and when to start anyway
There are two scenarios where I'd recommend fixing before starting.
The first is when the core fields are significantly incomplete or wrong — say, more than 20% of records have a problem in the field the workflow will act on. At that point, the automation is going to produce more errors than it saves time. Fix the data first.
The second is when the workflow doesn't include a human review step. If an AI workflow is going to send communications directly to customers without a human approving each one, the data it's acting on needs to be clean. If a human is reviewing and approving before anything goes out, there's more room to start with imperfect data and improve as you go.
Both of those are judgment calls. They depend on the specific workflow, the specific data, and the risk tolerance of the business. Part of what I do in the setup phase is make that assessment clearly — often after mapping how your systems connect.
The cost of waiting
I've seen businesses delay AI implementation for a year, sometimes two, waiting until their data was "ready." In most cases, the data would never have reached a standard they were satisfied with if the improvement work was done in isolation — with no concrete use case driving the prioritisation.
The most effective data quality work I've seen happens in the context of a specific goal. You're cleaning the fields that matter for the renewal reminder workflow. That forces clear decisions about what "clean enough" means, who owns each field, and what the maintenance process looks like. The work is bounded and purposeful rather than open-ended and indefinite.
If you'd like to work through whether your data is ready for a specific workflow, book a free call. We'll look at the actual data together and give you a straight answer — not a vague "your data needs work" but a specific assessment of what's usable now and what would need to change.