TinyTechGuides

AI Model Lake Convergence by 2029 | Stewart Bond, IDC

Data Faces Podcast — On Location · CDOIQ Symposium 2026

Stewart Bond, VP of Data Intelligence Software at IDC, on whether the data lake is obsolete, IDC’s model lake convergence, and governance that runs as code.

YouTube player

Listen: YouTube  ·  Spotify  ·  Apple Podcasts  ·  Amazon Music

About Stewart Bond

Stewart Bond on the Data Faces Podcast at 20th Annual CDOIQ Symposium

Stewart Bond is vice president of Data Intelligence Software at IDC, where he leads research on data integration, data intelligence, and data quality. A returning Data Faces guest, he tracks how enterprise data architectures are shifting for agentic AI, from centralized data lakes to what IDC calls model lake convergence.

In this interview

  • Why centralized data lakes struggle in an agentic world where agents cannot wait for batch processes
  • What model lake convergence means and why IDC expects it to peak by 2029
  • Why data quality, security, and privacy have to move closer to the source before agents and models consume the data
  • How governance changes when policy becomes code that runs on every data access

→ Browse all on-location interviews: Data Faces Podcast — On Location

Full transcript

David Sweenor 0:04 Hello and welcome to the Data Faces Podcast. On location, we are coming to you live from the CDOIQ event in Cambridge, Massachusetts, right next to MIT. Today, I am joined by Stewart Bond, VP of Research at IDC. Stuart, welcome to Data Faces Podcast.

Stewart Bond 0:22 Thanks for having me again. Again.

David Sweenor 0:24 Yes, you were on episode 34. Huge numbers. If you want to hear what Stewart thinks about data intelligence. Watch that one. We’re going to take a little different tact on this one. Okay. All right. So, Stewart, data faces. I always got to ask an icebreaker question. Okay. Before your LinkedIn profile existed, what was your first job?

Stewart Bond 0:50 My first job ever?

David Sweenor 0:52 Yeah.

Stewart Bond 0:53 I was a dishwasher at a mother’s pizzeria restaurant. Okay. And what’d you learn? You didn’t like washing dishes? Actually, honestly, it’s not that I like washing dishes, but that’s my job now. Okay. At home after dinner, I’m the one who washes the dishes.

David Sweenor 1:09 All right. All right.

Stewart Bond 1:12 I could say that maybe the job before that was working, doing some odd jobs around the company that my dad worked for, but that wasn’t really a real job.

David Sweenor 1:21 First one with the paste up.

Stewart Bond 1:23 Yeah.

David Sweenor 1:24 All right. That’s a good one. Well, thanks for sharing that. So, Stewart. You had a session today. Huge attendance there. What was it about?

Stewart Bond 1:34 It was about what the data architecture and what data governance needs to look like as we move towards more agentic enterprises.

David Sweenor 1:45 Okay, and the macro changes, you know, kind of what are they? We’ve been building on this foundation for… I don’t know, a long time now. Maybe Hadoop was sort of the first transition from DB2 and all that stuff. What are the macro changes?

Stewart Bond 2:00 There’s a really profound shift happening right now as we’re moving from centralized to more federated architectures. And one of the first slides in my deck today, in my presentation today, asked or suggested that perhaps the data lake is no longer relevant. It’s already irrelevant. Which is really interesting when you think about it, because that’s where organizations have been investing over the last decade, is creating these data lakes. But data lakes are also, now, some of the data lake vendors might argue with you, or argue with me on this, But essentially the idea of a data lake is it’s a better data warehouse. It’s all about centralizing the data, or at least putting all the data into one thing. And then that one thing, you also have to control the data in that one thing. but arguably data is everywhere in your organizations. It’s pervasive. It’s pervasive. You can’t jam it all in one good thing. No, and agents, as we move to agentic architectures and agentic solutions, they need to access the data in real time where it is it can’t wait for it can’t wait for a batch process to move data from an application into the lake before it can do anything with it it has to go and get that data where it is okay that’s interesting so we were centralized decentralized centralized we’re just going back to more of a federated model yeah because the ai doesn’t care to your point needs to reach into whatever crm or

David Sweenor 3:36 enterprise system is out there, and you can’t get it to a central location if you want it to fast enough. Is that sort of the notion, right?

Stewart Bond 3:43 Absolutely, yeah. Latency is the enemy of agentic AI.

David Sweenor 3:46 Okay, and then as part of this, you know, I think you were talking a little bit about models, and you have a new notion here on like how AI models or machine learning models or analytic models, how they fit into this. If you have a new concept here, I don’t know if a lot of people have heard of. Yeah, yeah.

Stewart Bond 4:03 Tell us a little bit about that. I was on a team at IDC last year that looked at what was going on in AI, what was kind of the future of AI and data. And we came up with this concept called Model Lake Convergence. Okay. By 2029, we’re going to see model-like convergence peak. What does that mean? That means that the model becomes the analytical layer. It’s no longer your data link that’s providing the insights. It’s the model that’s providing the insights, which is really kind of interesting when you think about it.

David Sweenor 4:41 Wait, but I mean, what are you saying? But the data’s still there?

Stewart Bond 4:44 The data’s still there, the data’s still available, whether it’s in your operational systems or in a lake, somewhere in your organization, the data’s still there. But the model’s been trained on what’s happening in your organization. And the first place to get the insights is from the model. You still need to have the data for all the compliance regulatory bits that you need. You need to have the lineage, need to be able to explain what’s happening in the model with the data that’s there. But there will probably be more metadata

David Sweenor 5:17 in the future that handles all that than the data itself because that’s where you’re going to get the explanation of what’s happening and so there’s a lot of questions we could ask about that but you know what does this you know we’re in a data quality conference here yeah what does that say about data quality or model quality or you know how do organizations need to think about that piece of it yeah yeah

Stewart Bond 5:43 So data quality, we know it’s always been a problem. Always will be a problem. And it always will be a problem. Sure. But I think what AI has done is it has really amplified the problem. And there’s this concept that’s been talked about for a while called shift left. The idea there is that you can’t deal with the data quality issues after the data’s ended up in an analytical repository. You have to deal with the data quality issues back at the source. You have to deal with the data quality issues there, because once it comes out of wherever it was, whether it’s an agent using that data directly, whether it’s a model being trained on that data, whether it’s being thrown into an event bus that’s then brokering that and distributing that around the organization. You can’t fix the quality after the fact. The quality has to be, not just quality, the security, privacy, all of those components, all those qualities of service that you need around that data needs to happen closer to the source of where that data came from.

David Sweenor 6:50 And that’s sort of the shift left, and this is really, you know, from what I understand of it is, you know, hey, the developer, the data engineer needs to have more ownership. That’s not really a new idea, though, is it? No, no. And so one of my questions is, they have their vantage, their view of the world, right? And it’s probably a lot smaller than maybe you as… the CEO or even just the owner of the entire data ecosystem. And so I’m just curious, how realistic is that shift left? Because they might get it perfectly tuned for their biopic use case. There’s other considerations that they may not ever even be aware of.

Stewart Bond 7:32 Absolutely. They’re not always aware of what’s downstream. They’re not always aware of where all the consumers are, what they’re doing. There’s a lot happening in the market in that area in terms of some people are calling it genetic data intelligence, where a lot of that information is now being captured. And it’s being captured more real-time, more autonomously. And all of that information is now… more available to the data engineer. So if the data engineer is building a pipeline for something they’re doing and they make a change somewhere, they can now very quickly see all the downstream impacts. They can ask an agent or, I don’t know if it’s an agent, but they can ask AI, what are all the downstream systems that are being impacted by this change? And they can get that answer very quickly without having to go and navigate through a lineage diagram. An agent can tell them where those changes are. So they’re becoming more aware of what’s downstream and what the changes they make, how that impacts what’s going to happen.

David Sweenor 8:37 Yeah, and then maybe another question I have, you mentioned metadata earlier, and we know there’s more metadata than probably actual data. We need to start thinking about metadata quality as well. Is this going to be the new thing? It should be. Or maybe it is a thing already. I don’t know.

Stewart Bond 8:53 It should be. It has been talked about. There are some vendors that are looking at it and taking that into consideration. But for the most part, the metadata should be generated from the data. With some exceptions. But there’s the context pieces.

David Sweenor 9:11 The context piece. It’s so wide.

Stewart Bond 9:13 The context is wider, and the context is more, I’d say, human-informed, potentially. Right. Especially when you get into semantics and all that sort of stuff. There’s a lot of other pieces that come into play that, yeah, could offer some potential bias, some potential… concerns about the quality, the validity of what’s being put into that metadata.

David Sweenor 9:38 Okay, okay, it makes sense to me. So when you were on the Data Faces podcast, I believe it was in March, you said, I have it up here, data governance is an organizational discipline rather than a technology problem, was one of the talking points. And so now you’re sort of new… Another perspective you’re adding on is the government is moving away from control points in communities and towards intent orchestration. Is that correct?

Stewart Bond 10:02 Yeah, intent orchestration. What does that mean? So think about, in my presentation, I had a before and after picture. And so today, we’ve got, when you want to have access to new data, you have to get all these approvals to get that access. Sure. Policy is in documents. Policy is enforced potentially through workflows and through process that’s on top of the data. And the process would potentially go and do something in the technology to give you access to control where you can have access to your knowledge. In the future world of agents, Agents can’t wait. Decisions are being made. Thousands of decisions are being made in milliseconds. They’re not being made by committees.

David Sweenor 11:00 Committees where decisions are going to die, right?

Stewart Bond 11:02 Exactly, yeah. The need to act upon what’s happening is much faster. an agentic world than what it was in the world that we’ve been building. And so that’s why I said latency is the enemy of agentic, right? Because he can’t wait for those things, those things have to happen. So the new underlying, so one of the things is around the governance portion, policy needs to become code. Sure. So that it’s code that’s implemented, it’s not a new document that’s written. Right. And that execution needs to happen every time data’s accessed, every time there’s a transaction, that policy needs to be applied. So that’s why it needs to be code and not necessarily documentation. So it’s not getting away completely from, governance is still an organizational discipline. You still need to have You still need to have direction. You still need to have understanding of what needs to happen and where governance needs to be applied. But in terms of how that governance is implemented, it’s going to be much lower level as code than higher level as process. It’s overlaid on top. The data architecture in the agentic area, the substrate of that, is going to be things like HTAP. We use the idea of HTAP. So it’s bringing analytical and operational workloads together. Okay. Because you can’t wait for the transactions to get to an analytical database before you can do anything on it. You need to be able to access both of those things at the same time. We’re going to see change data capture being used more often to get the events that are happening out of the source systems into some place that can be leveraged We’re going to see event streaming. We’re going to see event stream processing being used a lot more because that’s reaction to events, reaction to what’s happening in near real time. And that’s what agenda is.

David Sweenor 13:07 Yeah, a lot for our viewers and listeners to think about it. Stewart Bond, VP of Research at IDC, thank you for being a repeat guest on the Data Faces podcast. You’re welcome, Dave. Thank you. Cheers. Cheers.