TinyTechGuides

Data Quality for 140 Million Members | Intuit Credit Karma

Data Faces Podcast — On Location · CDOIQ Symposium 2026

Veenit Shah and Puneet Singh of Intuit Credit Karma on data quality for 140 million members, five pillars across 40,000 columns, and an AI agent that finds root cause in minutes.

YouTube player

Listen: YouTube  ·  Spotify  ·  Apple Podcasts  ·  Amazon Music

About Veenit Shah and Puneet Singh

Veenit Shah and Puneet Singh on the Data Faces Podcast at 20th Annual CDOIQ Symposium

Veenit Shah is a Senior Manager of Data Engineering and Puneet Singh is a Senior Data Engineer, both at Intuit Credit Karma, where their team keeps credit scores current for roughly 140 million members. They built the company’s data quality program from startup-era firefighting into five monitored pillars and an AI remediation agent that compresses root cause investigations from 30 to 40 minutes into just a few.

In this interview

  • How Credit Karma moved from reactive firefighting to five data quality pillars measured across 40,000 columns
  • Why every alert gets logged, measured, and analyzed weekly, with brutal prioritization of the noisiest issues
  • How an AI remediation agent runs and refines queries until it lands on root cause and blast radius
  • Where guardrails, anonymized data, and selective agent access fit as the team climbs the maturity curve

→ Browse all on-location interviews: Data Faces Podcast — On Location

Full transcript

Veenit Shah 0:00 Awesome.

David Sweenor 0:03 Hello, and welcome to the Data Faces podcast on location. We’re coming to you live from the CDOIQ event in Cambridge, Massachusetts. Today, standing next to me is Vineet Shah, Senior Manager of Data Engineering at Intuit Credit Karma. And we have Puneet Singh, Senior Data Engineer at Credit Karma. So tell us a little bit about you and your organization.

Veenit Shah 0:27 Sure, yeah, thank you for having us here. I’m Vineet, I work at Credit Karma as a senior data engineering manager. Been here about, coming up on my five years now, and Credit Karma, as you all know, is an organization that helps members enhance their finance journey, right? So this is every aspect of your financial lives, credit scores is one of the major aspects of it, and I lead a team at Credit Karma that allows us to show our members their latest like credit scores for about 140 million members wow that’s that is that is the that is the credit karma population so yeah and uh it’s it’s been a fun ride so far

Puneet Singh 1:12 Hey, I’m Puneet. So I work in Vineet’s team. I’m senior data engineer. Been with Care Command four and a half year. Being in the data engineering space for almost 12 years, started my career as back in Hadoop developer, then to the cloud, and now working with mostly Google Cloud, microservices, ingesting the data, data quality.

Veenit Shah 1:36 That’s my journey so far.

David Sweenor 1:37 So you guys have a wealth of experience. Tell us a little bit, you had a presentation. So tell us a little bit about that presentation and kind of what was it all about for the people who were not at this show?

Veenit Shah 1:49 Yeah, so we came here to share our journey of data quality adoption within Credit Karma. So we were in a pretty reactive state a couple of years ago to now where we are trying to get to more proactive alerting. especially with AI and alerts and enhancements, like how do we get to that next frontier of agentic world? So this all sounds good, but the journey had scars along the way, and we’ve had a lot of learning, so we came here to present and share our learnings.

David Sweenor 2:20 All right, well Puneet, can you tell us a little bit,

Puneet Singh 2:23 prior this transformation what was the state of the world and it sounds like it was a little bit of a mess uh tell us a little bit about it and kind of the the state of the world yeah exactly so uh kate karma like when i joined the company still have that starter feeling so which means that uh just move as fast as you can uh right try to extinguish as much fire as you can there was no set data quality checks, there was no patterns or pillars that we can rely on, and it was a mess, like long hours, nightly calls, so we take a step back and make sure, okay, let’s go to the basic, let’s go to and define the data quality pillars and build the whole framework around it so that we can build the whole posture that we could be on a stable state.

David Sweenor 3:09 Okay, and Vineet, so how did you, you know, so every organization has challenges with data, especially data quality. How did you get sort of the mandate or the project

Veenit Shah 3:25 to go fix all this you know because it’s easy to stay in that firefighting reaction reactionary mode because there’s always another problem to fix so how did you how did this get started yeah i think this goes like twofold so we’ve been when i said scars along the way we’ve had a few incidents that made us realize that our data quality posture is not where we want it to be So I think that was one part of it. And then the other part of it was just blessed with amazing leaders at the company that actually do allow us to trust our instincts and know exactly what is it that we should double down on. So I think both of these combined together got us to a place where we realized, okay, we have a blank board in front of us. How do we start drawing? Okay. How… How long has this journey been?

David Sweenor 4:18 Is it complete or is it always a work in progress? It’s never complete. There’s always more work.

Veenit Shah 4:25 But I think we started, I want to say, at least three years ago. And right now, what we know of where we are at in the journey, I feel like the next generation, GenAI and agentic applications is changing the game. So we actually don’t know what the end state of this journey is, but we do know what the next step is. And we just keep on taking the next step and we’ll figure out if we end up at this conference back again next year. For your next session. Okay.

David Sweenor 4:57 And Puneet, tell us a little bit, how do you go about thinking about data quality? How do you define it and how do you measure it to see you’re moving the needle?

Puneet Singh 5:08 Excellent question. So when we started this journey, the first thing that we did is let’s define the pillars as some of the guests in the first session mentioned that there are 19 different ways of doing data quality. And so you have 19 different measures of data quality? No, we have in the data management, world they had, we defined five of them.

David Sweenor 5:29 19’s a lot, like a handful. Exactly. A handful would be good.

Puneet Singh 5:34 So we have like five different pillars, like timeliness, completeness, accuracy, governance, and data…

Veenit Shah 5:45 Accuracy.

Puneet Singh 5:46 Accuracy, yeah. So these five are most relevant to our data set. So we define those and then we get to work for the most, the issues that is causing us most of the noise around it. So we solve those, but then as Vineet mentioned, that we have to showcase a little bit and that’s where like great leaders trusting our instinct come into picture. We start with small POC, so like there’s one vendor that we use for some of the use cases. When I pitched that idea, it was like, okay, show us some use case that can actually benefit it, measure it, and then come for the whole thing. We did a little bit, we saw the benefit, we were able to measure it, and we just went from there. And currently what we do is, any data quality alert that we have, we have around more than 100 tables, 40,000 columns, and we do data quality on each of them. Each and every alert is logged, measured, and we do analysis on weekly basis. to make sure where the fires are, where we need to get better, and where we need brutal prioritization to remove them or fine tune those.

David Sweenor 6:47 Okay. And so, Vineet talked about all these issues and these alerts. Whose job is it to fix data quality? Is it your group? Is it the business? Can I say agent?

Veenit Shah 7:02 Well, that could be part of the solution. I know, I know.

David Sweenor 7:05 Who owns it, really, in the end?

Veenit Shah 7:07 I think as a team, we show extreme high ownership across the board. So regardless of what level you’re at, whether you’re just starting as a junior engineer or whether you’re a senior manager as well,

David Sweenor 7:20 Right.

Veenit Shah 7:21 It is everybody’s job. It is a fundamental aspect of the impact of what happens when a data quality alert fires is what matters the most. So a simple alert could look like, hey, we have not been able to ingest like X percent of data. But what that actually means is X percent of members out of the 140 million are going to be impacted. And that is a huge impact, right? So I think fundamentally it is everybody’s job here within the team at least to understand and know what some of these data quality nuances are. Okay, okay.

David Sweenor 7:54 And maybe a question is we all understand the importance of data quality. You mentioned all these alerts you have. why doesn’t it come in perfectly every time? We have pretty robust systems. What are the major causes of it to be a mess?

Puneet Singh 8:13 Yeah, excellent question. As I mentioned, we have so much diversity of data and different partners, different FIs on our platform. it’s very hard to monitor those, and we started with the reactive solution that went through all manual interventions, manual checks, and then we widened it and did an omni-direction and all those things, and those are not perfect. So sometimes you get false positive, we dig into that, we make sure that we fine tune it, but even after doing all those things, there could be noise around it. So how we handle this here is using AI a lot. We build an AI remediation agent where every time something happens, the agent goes there, looks into the data quality error, runs the query, as well as refine those query and keep running the queries and analysis till it comes to the conclusion with the blast radius, saving our time from 30 to 40 minutes to just in few minutes. So that’s one thing which is helping us, but to your main question, yes, those are not 100% accurate, sometimes we have to fine tune it, and we are in the nice space now, which we are not a couple of years back. It was a lot more false positive.

Veenit Shah 9:24 I think to add to that, fundamentally at a high level, the issues could arise externally, that is kind of not in our control, versus internally. Yeah, something comes in a mess, and you’ve got to deal with it. It’s the cards that we are dealt with. Right, right. How do we play the hand? And then internally is when we are trying to develop new features and onboard newer partners and newer data sets, that’s when some of the issues arise. So historically, we have seen a lot of external facing issues is something that we have kind of taken a more like proactive approach. Internally as well, I think there’s definitely ways we can enhance our engineering practices better, but that’s like the fundamental way of like, why is it not a perfect world like all the time?

David Sweenor 10:05 Yeah, yeah, I hear you. And so, you’ve mentioned…

Veenit Shah 10:09 agents yes and how is this fundamentally changing the nature of your role right and where do you see it headed i guess yeah i so we actually we shared a few things yesterday in our in our slides there was a slide where we had like you know here’s like the level one here’s the level two here’s the level three and we know we are at level two right now is your level two like of maturity Yeah, it’s like the agent-like maturity in terms of the investigation. We knew that we are at level two, but we actually don’t know what L4 or L5 looks like right now. It is that fast. The nature of the game, it’s the game that’s changing and we’ve got to keep playing with it. So we don’t know what that looks like, but we do know what the next step is, as I mentioned prior, right? So the benefits that we have seen is fundamentally what used to take us do about like 15 to 30 minutes to thoroughly identify the root cause and run the queries ourselves. The agents are actually doing that for us. And we’ve guarded these agents with the proper security and other guardrails. But at the same time, we’ve explicitly given some of our runbooks and instructions to these agents to be like, this is exactly what we want you to do. And it’s all in a protected environment. So we don’t have to worry about the imprecacies of it.

Puneet Singh 11:28 Just to touch base on, we need, so that’s, the AI agents have helped us in two ways. One thing that we need to touch on is on AI investigation side. The other thing that we have done it is, everything in our data pipeline, we have different scale and context set where agents can come and write the PRs and code for us. So within five to ten minutes, I demoed it yesterday, that we can get production-ready thousands of lines of code for each aspect of pipeline, where it’s like pipeline generation, data quality, observability. So that’s the other way, other side of the agents, how it’s helping us.

David Sweenor 12:01 Okay, and maybe just to follow on to that, so agents, there’s a lot of concern. about them. Sometimes they go off the rails. I don’t know if you saw the headlines today, most recent ones.

Veenit Shah 12:15 They have a mind of their own.

David Sweenor 12:17 So how do you protect against something like that? It must be a big concern of yours.

Puneet Singh 12:21 Yes, that’s exactly. So we need touch base on it. That goes to the point that we are in the level two. Things that we want to do on the level three and level four is still under security review. So we want to make sure that’s safe, and anywhere in a company, whichever is a PI data or something goes to a place where no one have access to, including agents. But something which is anonymized, something which can be used for analytics, and that’s what we’re exposing it to agents currently. But on the way, we’ll learn about how and where we can give selective accesses to the agents, what tool can access, what their roles will be, and then we can slowly ramp it up, but it’s still like our journey, as Vineet mentioned.

David Sweenor 13:01 Okay. So, Vineet, maybe, you know, so you’ve been on this journey for about three years and you’re at level two. We don’t know quite where the future will, the fog will lift. But just tell us a little bit about some of the benefits that you’ve just seen from all of this work and this innovation that you and your team have done.

Veenit Shah 13:22 Yeah, I think the benefits are like multifold. We currently have seen the time savings, right? It gives us insane amount of time saved, specifically when a lot of the work that used to take us 30 minutes, an hour to do for investigation, that is compressed within minutes right now. So I feel like that is one of the significant benefits of it. But then on the other hand, it also allows us to think on the next frontier. What are some of the gaps that we did not know that existed? I think that is the major value add that we have seen with some of these agentic world. And I think it just compounds, right? That’s right. Every alert, every data quality issue that we have seen internally, it just starts to compound on its way over.

David Sweenor 14:10 All right, well amazing. Well, Vineet Shah, Vineet Singh, in the data engineering world at Intuit Credit Karma, thank you for joining the Data Faces podcast. Thank you. Thank you for having us. Cheers. Thank you so much. Thank you. Cheers.