AI Didn't Invent Your Data Problem

Tall industrial silos with colored bands under a cloudy blue sky
Photo by Waldemar Brandt on Unsplash

I’ve spent some time this month reading through the AI data readiness research that came out this year, and I saw the same number twice. Accenture’s May report says only 7% of companies have the data capabilities to scale advanced AI. A March study from Harvard Business Review Analytic Services and Cloudera says only 7% of respondents describe their organization’s data as completely ready for AI. Same number, two different studies, and they aren’t measuring the same thing.

Cloudera’s respondents named siloed data and trouble integrating sources as their top obstacle, at 56%.

In 2010 I asked whether going around IT to the cloud was building silos that nobody would ever integrate. In 2012 I called it the data disconnect : a team buys Basecamp, or keeps its own web analytics platform while IT runs a different one, and suddenly combining two reports means a pile of spreadsheet engineering. The work gets done, but it gets done in a vacuum, and whatever that data knows stays in the vacuum with it. The next day I wrote about the knowledge stuck inside those collaboration tools , all the things teams said to each other that never made it back into a system the company owned.

Companies have wanted their voice-of-customer, sales, marketing, and business development data to line up for about twenty years. Shadow IT kept buying products that never talked to each other, and the data ended up on laptops and in vendor clouds. AI inherits all of that.

Accenture says the old basics (accuracy, completeness, governance) “remain essential but are no longer sufficient.” Advanced AI also wants context: what your terms mean, how the work actually gets done, the reasoning behind a decision. It’s mostly the collaboration-tool knowledge from that 2012 post. The part of the disconnect we shrugged off back then is the part that causes the most pain for companies trying to implement AI today.

Two different 7s

Accenture’s 7% is its analysts’ grade. They surveyed executives at 2,000 companies across 15 countries and nine industries, sorted them by capability, and 7% qualified as what Accenture calls “data reinventors.” The Cloudera figure is self-reported. The full report surveyed 231 people involved in their organizations’ AI data decisions, and 7% said their data was completely ready. One is an outside assessment and the other is a self-assessment, and they happen to land on the same number. Both are also sponsored by companies that sell data work.

It’s easy to hear that as “you’re not ready, so wait,” and to answer it with a company-wide data program that has to finish before anyone touches an agent. Accenture’s own report argues the other direction: those 7% put their resources on “strategic bets” in core parts of the business, are nearly twice as likely as peers to concentrate there, and build readiness “through a continuous, iterative loop rather than a one-time transformation.” The companies at the top seem to have arrived there by narrowing.

So the move isn’t to “wait for readiness”… it’s to pick the work that already matters and make that data trustworthy first. To do this, pick one workflow that matters, something that moves money or customers for you. Write down only the data it needs and every system that data lives in. Then ask three questions of each system:

  • Can we get the data out, in a form another system can use?
  • Who owns the definition of each term this workflow depends on?
  • Where does the reasoning live (the exceptions, the approvals, the judgment calls), and is any of it written down?

Do that for the workflow that moves money or customers. Readiness compounds from one trustworthy slice that teaches the next; not a company-wide cleanup that has to finish before anyone starts. AI didn’t invent the disconnect… it just made waiting for “completely ready” the more expensive mistake.

Share

Get weekly insights on technology leadership

One idea per issue. No spam.