Data Readiness for AI: The Five-Question Audit

An older man data engineer and a friendly robot assistant examining five glowing data cubes lined up in a row, a friendly robot assistant marking one

“Our data isn’t ready for AI.” I hear it from IT directors at almost every mid-size organization, usually said with a sigh, as though it closes the conversation. Sometimes it’s used as a reason to delay AI indefinitely. Sometimes it’s the opening line of a proposal for a large data platform project.

Both reactions miss what data readiness for AI actually means. It doesn’t mean all your data is clean, centralized, and governed. No mid-size organization’s data is, and very few large ones are either. It means the data your first few AI use cases depend on is findable, owned, properly permissioned, reachable, and good enough for the job. That’s a much smaller problem, and you can assess it in a week with five questions.

Why data readiness is the dimension that decides pilots

Of the six readiness dimensions I assess, data is the one that most often decides whether an AI pilot succeeds. Security gaps can stop a pilot from starting. Skills gaps slow it down. But data problems let a pilot start, produce plausible-looking results, and then quietly lose users’ trust when the answers turn out to be wrong or stale. By then, the pilot’s reputation is set.

The good news is that data readiness is also the dimension where a focused effort pays off fastest, because you only need to fix the data your chosen use cases touch.

The five-question audit

Four of these questions come straight from the data dimension of the AI Readiness assessment. The fifth is the one that ties them to a specific use case. For each, I’ve included what good looks like, how to check it in about an hour, and where to go deeper.

1. Do you know where the truth lives?

What good looks like: for each question your first AI use case will answer, one system is designated as authoritative, and everyone agrees which one.

How to check: pick the use case, list the five questions users will ask it, and ask three people from different departments which system holds the right answer for each. If you get different answers, that’s your finding.

Go deeper: where should your business data live before AI? includes a one-page source-of-truth map.

2. Does someone own the quality?

What good looks like: each data domain your use case depends on has a named business owner and a steward, with a few written quality rules.

How to check: for each domain, ask “who would fix a wrong record, and who decides what correct means?” If the answer is “IT” or a shrug, ownership is missing.

Go deeper: who should own data quality? includes a one-page ownership charter.

3. Are your documents properly permissioned?

What good looks like: your file shares and collaboration sites have owners, reviewed permissions, and no broad “everyone” access on sensitive content.

How to check: run the sharing reports your platform already provides, and search as an ordinary user for a few sensitive words, such as “salary” or “pricing.” What you can find, an AI assistant can find and repeat.

Go deeper: fix SharePoint oversharing before you turn on Copilot has a 30-day cleanup plan.

4. Can AI reach the data safely?

What good looks like: the systems your use case needs have APIs or standard connectors, accessed through dedicated identities with the minimum permissions required.

How to check: for each system, find out whether an API exists, what license it needs, and how any existing integrations authenticate. Shared admin credentials are the red flag.

Go deeper: APIs before agents includes an AI access worksheet.

5. Is the data good enough for this use case?

What good looks like: the specific fields your use case relies on are complete, current, and accurate enough that a wrong answer would be rare and noticeable.

How to check: pull a sample of 50 records the use case would use. Count how many are missing a field it needs, how many are out of date, and how many contain obvious errors. You don’t need perfect numbers; you need to know whether the problem is small or large.

Go deeper: this question is where data readiness meets use-case selection. If the sample is poor, either fix those fields first or choose a different first use case, as covered in picking your first AI use case.

The asset: a one-week data readiness audit worksheet

Set up a simple table with one row per question and these columns:

  • Question: the five above.
  • Evidence gathered: what you checked and what you found, in a sentence or two.
  • Status: green (ready for the use case), amber (workable with a known fix), or red (blocks the use case).
  • Owner: who will act on it.
  • Next action and date.

Spread the checks across a week: questions 1 and 2 need conversations with business owners, questions 3 and 4 are desk work for IT, and question 5 needs someone who knows the data to review the sample.

At the end, you have a clear picture: which of the five are ready, which need work, and whether the work blocks your first use case or can run alongside it.

The order to fix things in

When several questions come back amber or red, fix them in this order:

  1. Permissions first. Oversharing is the fastest-moving risk, because an assistant turns it into exposure on day one. It’s also mostly admin work.
  2. Ownership second. Named owners make every other fix stick, and they’re needed to decide what “correct” means.
  3. Source of truth third. With owners in place, settling which system is authoritative becomes a decision rather than a debate.
  4. Access fourth. Build connections once you know which systems matter and who approves access.
  5. Quality fifth. With owners and a source of truth, quality fixes can be targeted at the fields that matter.

Most mid-size teams can move their first use case’s data from mostly red to mostly green in one quarter of part-time work.

What data readiness is not

It’s not a data platform project. A warehouse or lakehouse can help later, and it’s worth building when specific needs justify it. It’s not a prerequisite for a first AI use case.

It’s not cleaning everything. Historical data you’ll never use in an AI context can stay as it is.

It’s not a one-time exercise. Run the audit again for each new use case. It gets faster each time, because the ownership and permissions work carries over.

How data fits with the other five dimensions

Data is one of six dimensions in the AI Readiness framework, alongside security, infrastructure, skills, use cases, and governance. It overlaps most with security, because permissions and classification sit in both, and with use cases, because the right first use case is partly determined by which data is ready. For the complete framework, see the 6-dimension AI readiness framework, explained, and for all 24 questions in one place, the AI readiness checklist.

Where does your team actually stand?

The free AI Readiness Score covers data alongside the other five dimensions. It’s 10 questions and gives you a score in a few minutes.

Get your free AI Readiness Score →

Want to see what the full assessment covers first? Flip through a complete 38-page sample report.

Related guides

Frequently asked questions

What does data readiness for AI mean?

Not that all your data is clean and centralized. It means the data your first few AI use cases depend on is findable, owned, properly permissioned, reachable through safe connections, and good enough for the job. That's a much smaller problem than fixing everything.

What are the five questions in a data readiness audit?

Do you know where the truth lives, does someone own the quality, are your documents properly permissioned, can AI reach the data safely, and is the data good enough for this specific use case? Each is scored green, amber, or red with one line of evidence.

In what order should data problems be fixed?

Permissions first, because oversharing becomes exposure on day one. Then ownership, which makes every other fix stick. Then the source of truth, then access, then quality, targeted at the fields your use case actually uses.

Do we need a data platform before starting with AI?

No. A warehouse or lakehouse can help later, when specific needs justify it. For a first use case, reaching the right systems through APIs or connectors, with clear owners and permissions, is usually enough.

Scroll to Top