Ask an IT director where the company’s customer list lives and you’ll usually get a pause, then an honest answer: “It depends who you ask.” Sales has one in the CRM. Finance has another in the ERP, keyed by billing account instead of customer. Customer success keeps a spreadsheet with the columns nobody else tracks, and it’s the one leadership actually trusts.
That’s normal. Most mid-size organizations grew their systems one department at a time, and nobody ever had a reason to reconcile them. The question I get next is whether they need to centralize business data for AI before they can do anything useful with it. My answer is almost always “not all of it, and not yet.” What you need first is to know where the truth lives for the handful of questions AI will actually be asked.
Why AI makes a scattered data estate more visible
An AI assistant answers from whatever it can reach. Point it at a file share with three versions of the pricing sheet and it will pick one, quote it confidently, and never mention the other two. A human analyst knows which spreadsheet is the real one because they asked someone once. The assistant doesn’t have that context unless you give it a single authoritative place to look.
That’s the whole problem in one sentence: AI doesn’t create your data quality issues, it removes the humans who were quietly routing around them.
Here’s how the assessment asks the question, and the 0 to 4 ladder I score it against:
D1. Where does most of your core business data live today?
- Spreadsheets, file shares, and email attachments; no single source of truth
- Several business apps (ERP, CRM, HR) that don’t talk to each other
- Core apps with some point-to-point integrations; reporting is still manual
- A central warehouse or lakehouse covers some key domains
- A governed central data platform covers most domains, with documented access
Most mid-size teams I talk with land at 1 or 2. That’s workable. Level 4 is what large enterprises spend years building toward, and you don’t need it to start. What you need is to stop being at level 0 for the specific data your first AI use case depends on.
The decision you’re actually making
“Where should the data live” sounds like an architecture question. In practice it’s three smaller ones, answered per domain rather than for the whole company:
- Which system is authoritative? When two systems disagree about a customer’s address or a product’s price, which one wins? If the answer is “we’d have to check,” that domain isn’t ready for AI to answer questions about it.
- Should AI read it in place, or from a copy? For most mid-size organizations, reading from the system of record through an API or a vendor connector beats building a warehouse first. Copies go stale; the system of record doesn’t.
- Who can see it? Centralizing data concentrates access. A warehouse that anyone with a BI license can query is an oversharing problem with better performance.
None of those require buying anything. They require decisions, and decisions are what’s usually missing at levels 0 through 2.
Connect first, centralize later
Here’s the rule of thumb I give teams: connect before you centralize. If your first AI use case needs data from two or three systems, connect the assistant or automation to those systems directly and designate which one is authoritative for each field. Build a central warehouse or lakehouse only when you hit one of these triggers:
- You need to join data across four or more systems for the same question, repeatedly.
- Reporting queries are hurting the performance of a production system.
- You need history the source system doesn’t keep, such as point-in-time snapshots or trend data.
- Your auditors or regulators need a controlled, documented copy.
If none of those are true yet, a warehouse project will eat months and a chunk of budget before the first AI pilot sees any value from it. The pattern I keep seeing: a team spends most of a year building a platform their first use case didn’t need, and the pilot loses its sponsor before it ever launches.
When you do hit the triggers, the platform choice matters less than the governance around it. Microsoft Fabric, Snowflake, Databricks, and BigQuery will all hold your data. None of them will tell you who owns the customer table. That question is covered in who should own data quality, and it decides whether a central platform ends up at level 3 or becomes one more place the truth might be.
The asset: a one-page source-of-truth map
Before any architecture conversation, fill this in. It takes an afternoon with the right three or four people in a room, and it’s the most useful artifact I see teams produce at this stage.
Start with the five questions your first AI use cases will need to answer. Not “what data do we have,” but “what will people ask the assistant?” Then fill in a row for each one:
- The question: for example, “What’s the current status of this customer’s open orders?”
- Authoritative system: the one system that wins when sources disagree.
- Other places this data lives: every spreadsheet, export, and shadow copy you know about.
- Business owner: a named person outside IT who can say “this is correct.”
- How AI would reach it: API, vendor connector, database view, or “export by hand.” Be honest.
- Freshness: real-time, daily, weekly, or “whenever someone remembers.”
- Sensitivity: public, internal, confidential, or regulated.
Five rows. When you’re done, look for three patterns:
- Rows with no authoritative system. These aren’t AI-ready, full stop. Fix them before a pilot touches them, or pick a different pilot.
- Rows where “how AI would reach it” is “export by hand.” That’s your integration backlog, and it’s the subject of APIs before agents.
- Rows marked confidential or regulated with no clear owner. These are your risk items, and they should go straight to whoever runs security review.
Whatever’s left is where your first AI use case should live. It’s usually less than you hoped and more than enough to prove value.
Mistakes I see at this stage
Treating the file share as a data source. Unstructured content isn’t useless for AI; document search and summarization are good early wins. But a file share full of exports is not a source of truth. It’s a record of every time someone needed data and couldn’t get it from the real system. If your AI plan depends on those files, the permissions work in fixing SharePoint oversharing before Copilot comes first.
Starting with the platform. Buying a data platform before you’ve mapped sources and owners gets you an expensive new place for the truth to be ambiguous. Map first.
Trying to do every domain at once. You don’t need a company-wide data strategy to run one AI pilot. You need one domain done properly. Customer data or product data is usually the right first choice, because so many use cases depend on it.
Leaving business owners out. IT can tell you where data is stored. Only the business can tell you which version is right. If the source-of-truth session is all IT people, reschedule it.
What moving up one level looks like
If you’re at 0, the goal for next quarter is level 1: get the data that matters out of spreadsheets and into the business apps that should own it, and stop emailing exports. If you’re at 1, level 2 means connecting the two or three systems your first use case needs, even with simple point-to-point integrations. If you’re at 2, level 3 is justified when you hit the triggers above, and not before.
Each step up is a quarter’s worth of work for a mid-size team, not a multi-year program. The source-of-truth map tells you which step to take first. For the wider set of data work that should come before AI, see data readiness for AI: the five-question audit.
Where does your team actually stand?
Where your data lives is one of 24 questions in the AI Readiness assessment, which covers six dimensions: data, security, infrastructure, skills, use cases, and governance. The free version is 10 questions and gives you a score in a few minutes.
Get your free AI Readiness Score →
Want to see what the full assessment covers first? Flip through a complete 38-page sample report.
Related guides
- Who Should Own Data Quality? Why “Nobody” Breaks Your First AI Project
- APIs Before Agents: Giving AI Access to Your Systems Safely
- Fix SharePoint Oversharing Before You Turn On Copilot
- Data Readiness for AI: The Five-Question Audit
- The 6-Dimension AI Readiness Framework, Explained
Frequently asked questions
Do we need a data warehouse before using AI?
Usually not. For a first AI use case, connecting to two or three systems of record through APIs or vendor connectors is often enough. Build a central warehouse or lakehouse when you repeatedly need to join four or more systems, need history the source systems don't keep, or have auditors who require a controlled copy.
What is a source-of-truth map?
A one-page table listing the questions your first AI use cases will answer and, for each one, the authoritative system, other copies of the data, a business owner, how AI would reach it, how fresh it is, and how sensitive it is. It shows which data is ready for AI and which needs work first.
Why does scattered data cause problems for AI assistants?
An assistant answers from whatever it can reach. If three versions of a pricing sheet exist, it may quote any of them confidently. People used to route around bad copies from experience; AI removes that human judgment unless you designate one authoritative source.
Where do most mid-size organizations score on this question?
In my assessments, most mid-size teams land at level 1 or 2: several business apps that don't talk to each other, or core apps with some point-to-point integrations. That's workable. The goal is to stop being at level 0 for the specific data your first use case depends on.




