Data Faces Podcast — On Location · CDOIQ Symposium 2026
Kevin Petrie, VP of Research at BARC US, on why 70% of companies say their unstructured data isn’t AI-ready, the metadata that fixes it, and the coming agent cleanup.
Listen: YouTube · Spotify · Apple Podcasts · Amazon Music
About Kevin Petrie

Kevin Petrie is vice president of research at BARC US, where he studies data management, analytics, and enterprise AI adoption. An original Data Faces guest, he has spent his career helping organizations get their unstructured data, metadata, and governance ready for generative and agentic AI.
In this interview
- Why 70% of companies say their unstructured data isn’t ready for AI
- The types of metadata that classify and validate documents for AI use
- How to think about risk when you rent the model and own only the context
- What vibe slop is, and why agent consolidation is the next cleanup project
→ Browse all on-location interviews: Data Faces Podcast — On Location
Full transcript
Kevin Petrie 0:00 Which is still true. All right. We’re ready to roll.
David Sweenor 0:01 We’ll talk about whatever. It’ll be fun. We’ve got past guests, episode two.
Kevin Petrie 0:06 Do I have any food in my teeth?
David Sweenor 0:08 Do I? No. We’re good. We want to know if we had any food in our teeth.
Kevin Petrie 0:14 How is your employee doing? Is he performing okay?
David Sweenor 0:18 Okay. All right. Are we good? We’re recording? All right. Hello, and welcome to the Data Faces podcast. Podcast on location. We’re coming to you live from the CDO IQ event in Cambridge, Massachusetts, right next to MIT. I am sitting next to Kevin Petrie, a return guest on the Data Basis Podcast. He is VP of Research at BarkUS. Kevin, welcome to the show again.
Kevin Petrie 0:43 Thank you, David. I’m glad to be here.
David Sweenor 0:44 And for those who don’t know, you were on episode two. It all began. It all began. You were part of the OG crew. That’s really cool. I want to thank you. Was Sean first? No. He was number two. You were number two. Fantastic. And, you know, I started thinking about the name Data Faces and get behind the people in their professional careers. I like that. So… What did you want to be when you grew up? Or what was your first job? Do you have a preference?
Kevin Petrie 1:19 So I think this is actually consistent. Until now, I hadn’t really thought of it. But growing up, I always wanted to be an author of history in all likelihood. And where that’s morphed now is I would like to write books about the history of innovation. This is in my next life with my abundant free time. Sure.
David Sweenor 1:46 While you’re mountain biking or paddling through Maine.
Kevin Petrie 1:49 Or trying to pay for three college educations for my boys. I want to write about how Ben Franklin and Ford It wasn’t a great guy, but he did a good job with his cars. And other, you know, Turing, other guys had some really interesting innovations that have very clear implications and parallels with today’s innovation.
David Sweenor 2:13 Okay, well, when you’re ready to do that, consider Tiny Tech Guys as your publisher. I will. We’ll get you out there. Thank you. It’s got to be a tiny guy. It can’t be a novella.
Kevin Petrie 2:21 Well, actually, that makes me feel better. It feels a little more achievable if the page count is lower.
David Sweenor 2:27 There is true. There is true. We’ll talk about that. So, you know, you’ve been talking, you know, working for Bark. You’ve been talking a lot about unstructured data. Yes. What’s going on with the unstructured data world? And, you know, like what’s the state? Are people thinking about it? Are we ready for it?
Kevin Petrie 2:44 It’s a mess. It’s a mess. Yeah. I had the slide that I’m going to use in my presentation tomorrow has a picture of my children’s large Lego bin. And it’s just a free-for-all of all sorts of Lego collections over a couple of generations.
David Sweenor 2:59 I have one of those. There you go.
Kevin Petrie 3:01 But there is real value in there. And if you can organize it, you can build some very cool creations. And I’ve got a couple of the ones that my boys have created over the years. to show the before and after. We did a survey of organizations of all sizes, primarily in North America and Europe, and we found that 70% of companies say that less than half of their unstructured data is currently discoverable and usable for AI. Undiscoverable and unusable for AI.
David Sweenor 3:30 70%.
Kevin Petrie 3:32 So we’re not ready. 70% say more than half needs work.
David Sweenor 3:35 So we’re not ready.
Kevin Petrie 3:36 We’re not ready. We need to get ready because there’s an obsession, a rightful obsession with context to keep agents on the straight and narrow, to have them deliver value and have nuanced interactions with people and applications. If you want to do that, you need context, but you need context from unstructured data because that’s the real lifeblood of how your business works. Tables can only tell you so much. Those are trusted facts. Documents, PDFs, emails, images, that’s the proprietary context of a business, and it behooves AI adopters to get their arms around that.
David Sweenor 4:15 All of that dark data we can’t find in organizations. You know, it’s interesting. So I’ve asked a number of guests that have been recording over the past couple days about unstructured data, and I’ve had this… I thought in the market, when people say data quality, they’re thinking of tabular data. Yes, they are. I feel like the market as a whole, unstructured data is sort of, people aren’t thinking about it, or they’re thinking about it, but the vendors aren’t quite there. There’s not a lot of vendors focused on it. Why is this, Kevin? And do you have the same feeling? That’s a good question.
Kevin Petrie 4:50 Well, structured data is a lot easier to work with.
David Sweenor 4:53 It’s easy, right? You can do a distribution and whatever.
Kevin Petrie 4:55 Rows, columns, you have a cell. It’s either right or it’s wrong. It’s a binary deterministic decision, whereas unstructured data involves probabilistic judgment calls about the accuracy, the suitability, the sentiment, and so forth of a given document. And that gets tricky fast. So a lot of vendors, when they talk about context, if they’re not specific, yes, they’re going to be talking about structured data because that’s what they’re most comfortable with. But there are companies, I just had a great session that I moderated with Flexor, exclusively focused on unstructured data. Unstructured.io is another one. And they’re doing a lot to help companies make order out of all these unstructured objects and feed them into context engineering programs, among other things, to support agentic AI.
David Sweenor 5:45 Okay. And then when you’re thinking about unstructured data, let’s just go with text for now because it’s easiest and people get PDFs and Word docs and things like that. What are some of the dimensions people need to be able to think about? Is this like, so if I got a PDF file or a Word doc, how do I ascertain this is a good, valid doc I should use for my context?
Kevin Petrie 6:11 Yeah, good question. A lot of this boils down to metadata. There are different dimensions of metadata. There’s technical metadata that will talk about the structure, schema, and so forth. That’s not going to help too much. You want to make sure it’s not too long or it’s in the right format. There is business metadata that’s going to talk about… terminology, agreed concepts, perhaps master data, which does not get the attention it deserves. And that’s a good measure to start to assess the suitability of a given document. There’s also operational metadata. The best piece there is usage ratings from people who apply tags, trustability scores, and things like that. This document is very well suited for use case one and two, and by the way, it’s current. Then of course there’s governance metadata, we get into things like access rights, role-based access controls, and so forth. And I think all those types of metadata can result in the accurate classification and validation of unstructured data objects for AI innovation.
David Sweenor 7:21 So, that’s interesting. Did one of those, like, talk about the content of the document, meaning it could be AI slop, it could be human slop. Yeah. It could be just gibberish in there. Is one of those, is that factored into that?
Kevin Petrie 7:39 Yeah, so it’s a good question. So that’s where you get into humans.
David Sweenor 7:44 But no one could possibly, nobody reads, number one. I sell books, I tell you that for a fact. Very little people read. Maybe it’s just the tiny tech guys.
Kevin Petrie 7:53 I’m sure they read your stuff. It’s tiny.
David Sweenor 7:56 I know, but you know what I’m saying, though? It’s like, you mentioned humans are applying tags, and there’s probably, I don’t know, you’ll know the stat, but like… 10x more unstructured data than structured data. We can’t get that right. So what hope is to have a human going in or many humans going in and tagging things and reading things? Sure. Is it realistic?
Kevin Petrie 8:17 Well, so the good news is that a small subset of unstructured data probably gives you what you need. and avoids the expensive processing costs that you want to minimize. So if you have a handful of vetted documents that are approved by business experts who wrote and have to read them, legal experts, other subject matter experts, and then you have a set of, let’s say, call service notes, where people whose job it is to accurately reflect customer conversation those are the kind of things where you’re gonna go to the subject matter expert a human whose paycheck depends on the quality of that content and you get them to vouch for its contents at that point you should feel free to put it into the context layer and have an LLM go to town parsing okay here’s another question I had for you so
David Sweenor 9:16 In the world of predictive models, the company owned everything in it. They would get their data, they’d build the model, whatever they wanted, and get some output. They controlled the entire pipeline. So now, they don’t own this middle part. They’re generally not building their own LOM, I would say. They’re going to do some sort of frontier model. Is there risk that comes along with that? And then how does the context interact with that? Could it be at odds? What wins in the end? I’ve always thought about this.
Kevin Petrie 9:51 Yeah, well it’s a good question because most organizations are not going to find it cost-effective and they lack the expertise to meaningfully fine-tune a LLLM. So going in there and starting to tweak parameters to reflect their own business is hard.
David Sweenor 10:07 I feel like that was like last year. Oh, we got to go fine-tune everything.
Kevin Petrie 10:10 And now everybody’s like, oh, forget about that. Forget about that.
David Sweenor 10:13 That’s too much. That was last year’s conference. We’re going to fine-tune the hyperprint and all this stuff. Gone. Not even part of the conversation anymore. It’s true.
Kevin Petrie 10:22 It’s true. they are getting pretty attuned to understand and respond to the content you feed them and to adopt that voice and to start to adopt some of those cultural assumptions and so forth. There’s no question that if you ask Mistral a question versus a Chinese model versus Anthropic a question about how families should behave when we have the loss of a loved one. they’ll probably give different answers because there’s a lot of cultural assumption baked into those. But for most business arrangements, you can have a fair amount of confidence about how to handle situations based on the specific content you’re feeding it through the RAG process and other things.
David Sweenor 11:09 And there’s this whole set of guardrails. Part of my side thing is I’m a master gardener, right? And so I have to answer questions from the hotline. And I got a question in and I was using an AI model to help research. It wasn’t giving me, it just helped scan a million websites way faster than I can. Pictures and this and that. And it was the newest, it was the Fable model. Oh, we can’t answer any questions on biology. It was something like to that effect.
Kevin Petrie 11:40 It was the guardrail.
David Sweenor 11:41 I put it in and it says, I have this question about plants. It just refused to answer it.
Kevin Petrie 11:45 Talk about biological warfare.
David Sweenor 11:47 Oh, it just refused to answer any question about the gardening. So I said, well, Opus will answer it for you.
Kevin Petrie 11:52 Pretty broad-based limitation.
David Sweenor 11:54 It’s pretty weird. Yeah. It’s pretty weird. But the point is, anyways, there’s a lot of guardrails that people need to put into this. And are the guardrails, I guess, there’s obviously guardrails built into the model, but are companies thinking more about their unstructured data on the inputs as well as the outputs? I suppose you’ve got to monitor both.
Kevin Petrie 12:16 Right, so we’ve talked about a lot of the steps to classify, validate, and then prepare, filter, and deliver the unstructured data that’s going to feed the model at inference time. And they’re going to feed machine learning models as well as GNI models. But in terms of the outputs, you do need to have some validation and some controls as well. And that’s where I think sentiment indicators can help. You might have a sentiment indicator that shows, hey, you know, you’re talking to an angry American. You need to dial it down a little bit. Or if you’re talking to an angry person in another country, maybe that’s more cultural. They sound angry, but they’re not. So you do want to start to have sentiment indicators. You want to have kill switches. You want to have certain vocabulary that will trigger a human review and human intervention. So there are a lot of ways in which you do need to monitor those outputs in order to prevent bad decisions, bad recommendations, and then what follows is damaging actions. It’s interesting. And that gets hard.
David Sweenor 13:22 I think I just saw something on LinkedIn. So the fastest way to get to a customer service agent, a real human, is to start swearing to the automated system. Oh, really? It bumps right to a human. I was like, I haven’t tried it.
Kevin Petrie 13:33 Oh, see, that’s funny. Yeah. That brings me back to a very late, frustrating evening with a flight attendant years ago. And when I’d been booted off my plane, I used a word I wasn’t proud of. And she immediately ended the conversation and would not let me get back in line. Yeah.
David Sweenor 13:49 The human judgment call is a little different on that. There you go. All right. Well, another question. I think you might have coined the term vibe slop.
Kevin Petrie 13:58 Is that yours? Vibe slop. I think I got it from the journal.
David Sweenor 14:02 Tell me what is this and what do CDOs need to look for? What is vibe slop?
Kevin Petrie 14:07 Yeah, good question. I think some of this comes from the red cloth or the… I’m going to have to edit this part out. The group that came out with the ability to very quickly create agents.
David Sweenor 14:22 Okay.
Kevin Petrie 14:22 Yeah, some of those initial creators are very worried about vibe slop. And I’m worried about vibe slop because it’s now very possible for you and me you are a skilled data scientist, I am not. Or anyone else who’s helping do all sorts of things, they can now create their own agents. And so knowledge workers of all types are creating agents, one-off ad hoc agents, that help them do their job, or their little team do their job. but it undermines the overall infrastructure of the organization. And so I think that nine to 12 months from now, the buzzword is not going to be agentic AI. The buzzword is going to be agent modernization, agent consolidation, which means cleaning up all the shit. that we put into place in the modern enterprise.
David Sweenor 15:11 We’re going back to the proliferation of BI dashboards. Everybody’s got their own personalized dashboards that works exactly to them. They don’t agree. We’re learning once again. We have to refactor them because you got one that works for you, I got one that works for me, and they don’t agree.
Kevin Petrie 15:27 Ungoverned and self-service. And so we do advisory services at Bark, and we have privately advised several vendors not to talk about the number of agents that they put into production, because 70,000 is not necessarily a good number. Five, 10, 15 really well-governed agents, that’s good. That implies a level of control. It’s not the number. It’s not an arms race. It’s the level of control as you improve your systems.
David Sweenor 15:57 All right, well, fantastic. Well, Kevin Petrie, Vice President of Research at BarkUS, thank you for returning to the Data Faces podcast. It’s been a great conversation. It was a pleasure. Thank you, sir. Cheers.

