Malcolm Hawker on fitness for purpose and semantic layer limits
Listen now on YouTube | Spotify | Apple Podcasts | Amazon Music
▶ Watch the full episode and read the transcript: Malcolm Hawker on AI-ready data and identity resolution

When I built data warehouses at IBM, I used to argue with our chief architect about master data management (MDM). My view was that if the manufacturing data was wrong, we should fix it in the source systems where it was created. His view was that we would build the mapping tables, reconcile the identifiers downstream, and leave the source systems alone. He won those arguments, and for years afterward I understood MDM mechanically without ever being convinced it was solving the right problem.
Then I spent the second half of my career marketing data, analytics, and AI platforms at SAS, TIBCO, Alteryx, and Alation, and I still write that copy for clients. Today it says "AI-ready." Much as it pains me, the phrase is on nearly every vendor site, it draws a great deal of budget, and I have yet to see two companies define it the same way. So when Malcolm Hawker came on the Data Faces Podcast, I asked him what it means, and whether the companies claiming their data is AI-ready are fooling themselves.
Malcolm's not really a fan of the term "AI-ready." Of course, data quality still matters, but readiness was never a property your data has or lacks, and treating it as a yes-or-no measurement is why the phrase is ambiguous and nearly meaningless. The industry, he adds, has strong incentives to keep producing phrases exactly like it.
About Malcolm Hawker
Malcolm Hawker is the Chief Data Officer at Profisee, a master data management vendor, where most of his job is external thought leadership. He spent three years as a Gartner analyst covering MDM and governance, and before that held IT leadership and product roles including Distinguished Architect at Dun & Bradstreet. He hosts the CDO Matters podcast and wrote The Data Hero Playbook.1
In this episode, Malcolm and I discuss:
- Why "AI-ready data" is an overloaded term that collapses a spectrum into a binary
- Why the cost of being wrong decides whether an AI use case can go into production
- The difference between defining what a customer means and knowing which customer record is real
- Why telling your CEO "garbage in, garbage out" is a career-limiting move
- The semantic pedantic feedback loop, and how data literacy went from nonexistent to a top-three problem in a single year
Watch the full conversation here:
Ready for what, exactly
Malcolm's definition of AI-ready data is a practical one. Data is AI-ready when it supports the use case in front of it. It maps onto the oldest working definition of data quality, which is fitness for purpose, and it means the same customer table can be perfectly ready for one job and nowhere near ready for the one running beside it.
Large language models (LLMs) were trained on the open internet, which he describes as a complete cesspool when it comes to data quality, and they work well enough that the entire market reorganized itself around them. At the other end sit healthcare, legal, finance, and anything carrying compliance or audit obligations, where accuracy and consistency are not negotiable. Both of those are AI use cases, and no single quality standard covers both.
Malcolm reframes readiness as a question of consequences rather than a question about data. Send a digital coupon to the wrong person, and it costs you close to nothing, which means your data is probably ready today. Model a specific customer's behavior, or decide what offer that person should see next, and you need a far higher standard. In Malcolm's read, that mismatch is also why so many proofs of concept die: organizations take a probabilistic system and drop it into a business process that has always run on deterministic rules, where the output was predictable, auditable, and safely inside a governance policy.
AI models do not behave that way, and the reason they misbehave is rarely recoverable after the fact. Malcolm calls this the attribution problem, and no amount of retrieval patterns, grounding, or knowledge graph scaffolding fully removes it. You can reduce how often the system surprises you, but you can't get to a place where you know why it did what it did every time.
"AI-ready data is data that supports a given use case. I mean, literally, that's it."
— Malcolm Hawker, Chief Data Officer, Profisee
What the semantic layer cannot fix
Semantic layers exist for a good reason, and Malcolm is careful not to argue against them. An LLM cannot look at your raw tables and say anything useful about your customers, because it does not know what a customer is in your business or how your customer table joins to your product table. What it wants is text. Josh Howard put the same point more bluntly on this show: your AI is dumb without your data. Layering definitions over the tables, which is roughly what the market means by a semantic layer, is how the tabular data becomes legible to it.
Knowledge engineers have carried a name for this split for decades. The T-box, short for terminology box, holds the definitions: what counts as a customer, how net revenue is calculated, and what a location is.2 Nearly every conversation in the market right now, whether it is labeled ontology, knowledge graph, or semantic layer, is a conversation about the T-box. Malcolm argues that the other half of the structure has gone mum. The A-box, the assertional box, holds the actual values, and it answers a different kind of question: is Malcolm Hawker a customer, and if five Malcolm Hawker records exist in the system, which one is him?
Identity resolution, master data management, and old-fashioned data quality all live in the A-box, and no quantity of definitions reaches them. A perfectly specified ontology will tell an AI agent exactly what a customer is while remaining completely silent about which of your customer records is the real one. On this show, Matt Hayes made a related argument that AI-ready data clears a much higher bar than analytics-ready data, and Sam Pierson described the industry work to standardize semantic definitions so meaning travels with the data. Malcolm is adding the layer underneath both, where meaning gets you part of the way, and identity gets you the rest.
"If you've got 15 different versions of David Sweenor, all the context in the world isn't going to solve that problem."
— Malcolm Hawker, Chief Data Officer, Profisee
Identity problems are also harder to catch than the failures data teams are trained to look for. People can see a broken dashboard because they notice a number is off. Duplicate customer records produce answers that look entirely reasonable while resting on the wrong person, so the system returns them with total confidence, and nobody has a reason to check.
Never say "garbage in, garbage out" to your CEO
Garbage in, garbage out is probably the most repeated line in data management and analytics. The phrase flattens a spectrum into a binary, the same way "AI-ready" does, and it writes off data that might be entirely sufficient for the job at hand. That alone would make it sloppy, and Brendan Grady made the companion argument on this show about why bad data didn't matter until now, which is that analytics absorbed errors AI will not. What makes the phrase dangerous is what happens when a data leader says it out loud in front of the business.
"First of all, the whole concept of garbage in, garbage out gives me hives."
— Malcolm Hawker, Chief Data Officer, Profisee
Malcolm puts it in a boardroom. You are the CDO, and your CEO is holding a report with two David Sweenors on it and wants to know why there are two of a customer you both know is one person. No data leader answers that question with "sorry, boss, garbage in, garbage out" and keeps their standing in the room. I mentioned to Malcolm that this is career-limiting, because the phrase disowns the entire investment. In that one sentence, the warehouse, the pipelines, and every governance and quality platform on the books are declared powerless against whatever arrived upstream. It hands away the argument that data work changes outcomes, which is the entire argument a CDO is employed to make.
The semantic pedantic feedback loop
Malcolm first posted the most useful idea in this conversation on LinkedIn as a joke, then could not stop finding evidence for it. He calls it the semantic pedantic feedback loop, and it explains why our field keeps inventing new names for things that already had names. I knew to ask about it because Scott Taylor, who came on this show to argue for putting truth before meaning, told me I had to.
It runs in a circle. Analysts, consultants, and thought leaders need to be seen saying something new, and in the analyst business there is a subscription renewal attached to that need, because nobody renews to hear that best practice has not changed much. So new vocabulary gets manufactured. Vendors pick it up and build campaigns around it. Buyers start hearing it from every vendor and at every conference, decide it must be important, and call their analyst to ask what it is. The analyst hears their own coinage echoed back by the whole market and confirms that it is real. Around it goes.
Malcolm's evidence is a receipt from inside Gartner itself. By his account, Gartner's annual CDO survey asks respondents to name their top roadblocks, and in 2017 no respondent named data literacy, because it was not among the options. Gartner added it the following year, and it immediately ranked third. He is pointed about why that particular option travels so well: it locates the obstacle in everyone else's knowledge. The impediment sits with the people being served rather than with the tools, the dashboards, or the data team itself. As he puts it, the thing had a name before it got a new one, and the name was training.
Let me say where I am standing while I write this. I host a podcast, I have written books, and a decade of my career went into product marketing for data platforms. That puts me inside the loop being described, along with the phrase that opened this article. "AI-ready" is the loop's most recent output. Recognizing that buys you one useful habit: every time a new term arrives, ask what it was called last time and whether anything underneath it changed.
What the good ones do differently
Malcolm has an unusually large sample to draw on when he describes what separates the data leaders who succeed. Gartner analysts take client calls they call inquiries, and he did roughly 1,500 of them in three years with CDOs, CIOs, and VPs of data and analytics. In struggling organizations, he heard the same attribution pattern every time. They blamed data literacy, or culture, or the fact that nobody showed up to the governance committee meeting. The blocker was always somewhere else.
Leaders getting lasting results in governance, data quality, and MDM sounded different. They expected to be wrong the first time and to learn from it. They measured themselves on whether the people consuming their data products were successful, and they asked for feedback often enough to find out when they were not. That argument is what The Data Hero Playbook is built on, and the book's premise is that the limiting factor in most data organizations is a set of beliefs rather than a missing capability.1
So the next time someone asks whether your data is AI-ready, the useful response is that the question is incomplete. Ask what the use case is, because readiness is meaningless without one. Ask what being wrong costs, because that number decides whether the use case can go into production. Then ask whether you can tell which record is the real one, because if you cannot, no amount of context is going to rescue the answer.
Listen to the full conversation with Malcolm Hawker on his Data Faces Podcast episode page.
Based on insights from Malcolm Hawker, Chief Data Officer at Profisee, featured on the Data Faces Podcast.
Podcast highlights
- [0:06] Introduction and welcome to the Data Faces Podcast
- [1:18] From IT operations to Distinguished Architect at Dun & Bradstreet
- [2:20] Three years at Gartner, and what an externally facing CDO actually does
- [4:22] The record label he never started
- [5:09] A graduate thesis on Napster, and the case for buying a turntable again
- [7:02] What does "AI-ready data" even mean?
- [7:46] Fitness for purpose, and why the term is overloaded
- [8:05] ChatGPT was trained on the internet and works anyway
- [11:56] The cost of being wrong, and why a coupon is not a claims decision
- [13:10] Semantic layers, the T-box, and the fifteen David Sweenors
- [14:51] Does "garbage in, garbage out" mean anything for unstructured data?
- [15:44] Why the phrase gives Malcolm hives
- [15:56] The boardroom scene, and why the answer is career-limiting
- [17:02] Chunking text destroys the context that made it useful
- [19:26] An underserved market, and back to the semantic pedantic cycle
- [19:57] The feedback loop: pundits, vendors, buyers, analysts
- [23:36] How data literacy became a top-three CDO roadblock in one year
- [24:09] "There was a word for it, and it was called training"
- [25:17] Does master data management apply to unstructured data?
- [25:59] Master data as the connective tissue across business processes
- [28:03] Profiling, entity linkage, and the fox watching the henhouse
- [29:32] Contracts and forms are tractable; email is the hard problem
- [29:54] The Data Hero Playbook, and whether CDOs play defense
- [31:21] Growth mindset versus fixed mindset, drawn from 1,500 Gartner inquiries
- [35:18] What Malcolm is reading3
- [37:29] Where to find Malcolm
Frequently asked questions
What does "AI-ready data" mean?
Data is AI-ready when it supports the specific use case you intend to run on it. There is no universal threshold, because fitness depends entirely on the job. Malcolm Hawker, Chief Data Officer at Profisee, argues the term has become overloaded precisely because the market treats it as a binary state your data either has or lacks. The same customer data can be sufficient for a marketing recommendation and unacceptable for a clinical or compliance decision, so always ask "ready for what."
How is AI-ready data different from analytics-ready data?
Analytics-ready data supports human analysts reading dashboards and reports, where a person applies judgment and notices when a number looks wrong. AI-ready data feeds systems that act on the data without that check, so errors propagate at machine speed with nobody in the loop to catch them. The accuracy, consistency, and identity requirements are correspondingly higher, and the bar rises further as the cost of being wrong rises.
Does a semantic layer or knowledge graph solve AI data quality?
Only partly. Semantic layers, ontologies, and knowledge graphs define what your terms mean, which knowledge engineers call the T-box. They do not establish which specific records are true, which is the A-box. If your systems hold fifteen versions of the same customer, a perfect set of definitions will still not tell an AI agent which one is the real person. Identity resolution and master data management address that half of the problem.
Why do so many AI proofs of concept fail?
A recurring reason is that organizations deploy probabilistic systems into business processes that historically ran on deterministic rules. Those processes were built on outputs that were predictable, auditable, and safely inside a governance policy. Language models do not offer that guarantee, and what practitioners call the attribution problem means you often cannot determine why the system produced a given answer, which makes the failure difficult to diagnose or defend.
How do you apply data quality to unstructured data?
Nobody has solved this well yet. Traditional quality rules need small, structured units, so the standard move is to break text into chunks, but chunking strips away the context that made the text useful to a language model. Practical progress starts with profiling, identifying which documents reference your master data objects such as customers, suppliers, or products, and then determining whether the claims in those documents are true. That last step remains the hardest.
What should a data leader do instead of asking "is our data AI-ready?"
Replace it with three questions. What is the use case, since readiness has no meaning without one. What does being wrong cost, since that determines whether the use case can go into production at all. And can you identify which record is the real one, since context and definitions cannot compensate for unresolved identity. Answering those tells you far more than any readiness score.
About David Sweenor
David Sweenor is the founder and host of the Data Faces podcast, where he talks with the people who are making data, analytics, AI, and marketing work in the real world. He is also the founder of TinyTechGuides and a recognized top 10 AI thought leader and international speaker who specializes in practical business applications of artificial intelligence and advanced analytics.
With over 25 years of hands-on experience implementing AI and analytics solutions, David has supported organizations including Alation, Alteryx, TIBCO, SAS, IBM, Dell, and Quest. His work spans marketing leadership, analytics implementation, and specialized expertise in AI, machine learning, data science, IoT, and business intelligence. David holds several patents and consistently delivers insights that bridge technical capabilities with business value.
Books
- Artificial Intelligence: An Executive Guide to Make AI Work for Your Business
- Generative AI Business Applications: An Executive Guide with Real-Life Examples and Case Studies
- The Generative AI Practitioner's Guide: How to Apply LLM Patterns for Enterprise Applications
- The CIO's Guide to Adopting Generative AI: Five Keys to Success
- Modern B2B Marketing: A Practitioner's Guide to Marketing Excellence
- The PMM's Prompt Playbook: Mastering Generative AI for B2B Marketing Success
Follow David on Twitter @DavidSweenor and connect with him on LinkedIn.
Footnotes
Footnotes
-
Hawker, Malcolm. The Data Hero Playbook: Developing Your Data Leadership Superpowers. Hoboken, NJ: Wiley, 2026. https://www.wiley.com/en-us/The+Data+Hero+Playbook%3A+Developing+Your+Data+Leadership+Superpowers-p-9781394310647. ↩ ↩2
-
Baader, Franz, and Werner Nutt. "Basic Description Logics." In The Description Logic Handbook: Theory, Implementation and Applications. Cambridge University Press, 2003. https://www.inf.unibz.it/~franconi/dl/course/dlhb/dlhb-02.pdf. ↩
-
Hilger, Joseph, Lulit Tesfaye, and Zachary Wahl. Bridging Knowledge, Data, and AI: Harnessing the Semantic Layer Framework to Drive Intelligence. Springer, 2026. https://link.springer.com/book/10.1007/978-3-032-17178-8. ↩