Data Faces Podcast — On Location · CDOIQ Symposium 2026
Tom Redman, the Data Doc, on the 3 percent study, why regular people do the data work, and the incentives that manufacture bad data.
Listen: YouTube · Spotify · Apple Podcasts · Amazon Music
About Tom Redman

Tom Redman, known as the Data Doc, is president of Data Quality Solutions and one of the founding fathers of corporate data quality. He led the Data Quality Lab at Bell Labs in the late 1980s, has written for Harvard Business Review and MIT Sloan Management Review for three decades, and authored the famous study finding only 3 percent of companies met basic data quality standards. His latest book is People and Data.
In this interview
- Why quality improves quickly once everyone sees their roles as data creators and data customers
- How misaligned incentives, like badge scans measured by volume, manufacture bad data
- Which layers of the data quality problem are solvable now, and which will take decades
- What the 3 percent study measured, and why people who demand 100 percent accuracy are getting closer to 50
→ Browse all on-location interviews: Data Faces Podcast — On Location
Full transcript
David Sweenor 0:00 Look at this camera. Got it. Okay. Okay.
Tom Redman 0:05 Are we ready to go? Are you ready? And presumably you’re going to cut and paste.
David Sweenor 0:11 I hardly edit at all. Oh, okay. Don’t mess up. I’m just kidding. We’ll edit if we need to.
Tom Redman 0:18 Okay. Well, I’ve got a business manager who’ll want to see it. Okay, yeah. Her job is to make me look good. Fair enough.
David Sweenor 0:24 All right, here we go. Are you ready? I am. Hello, and welcome to the Data Faces podcast live on location. We are coming to you from the CDOIQ event in Cambridge, Massachusetts, right next to MIT. Right now, I am joined with Tom Redman. He is the data doc. Tom, welcome to the Data Faces podcast.
Tom Redman 0:48 Yeah, thanks for having me, David. I’ve been looking forward to this.
David Sweenor 0:50 Yeah, me too. So tell us a little about yourself and what you do.
Tom Redman 0:54 Okay, well, look, my background is I’m a statistician off the Bell Labs for me. While I was there, I was lucky enough to fall into some AT&T business problems that involved data quality. Okay. We set up a lab on data quality, you know, far before anybody knew what was going on, and it was really awesome. We worked on two planes, first of all, solving a important company problems. And then a smaller plane, which is like, well, how do we develop the foundations for this stuff that are going to serve the company over the long term? And so I never would have left, but AT&T had some problems. So I hung my shingle over 30 years ago. There you go. And today, more or less, I wander around and bump into things, mostly around data and data quality, and more and more on getting the organizational stuff right. So, you know, look, I’ve been all over the world over the course of time. People have taught me how to write. I’ve written, I don’t know, hundreds of papers and seven books and have two patents. Mostly what I specialize in is asking dumb questions. I love that. And listening to the answers and then, you know, trying to find some…
David Sweenor 2:21 sleeve some some pathway forward for an organization that aligns with its values sure and helps it get better data yeah yeah it makes a lot of sense to me and you know the name of the show is data faces so i like to get behind the people in their professional career so um before your linkedin profile existed What did you want to be when you grew up?
Tom Redman 2:43 Well, I don’t know. I probably wanted to play Major League Baseball. All right, yeah. What position? Yeah, everybody wanted to be shortstop, so I probably wanted to be shortstop. Look, I went to Bell Labs. I wanted to be a good statistician at Bell Labs.
David Sweenor 2:58 No, no, before that. So baseball player. You wanted to be a baseball player.
Tom Redman 3:01 Okay, a baseball player, sure. I wanted to be a baseball player. All right. Any team? Oh, no, I wanted to be a Yankee.
David Sweenor 3:06 Okay.
Tom Redman 3:06 I mean, you know.
David Sweenor 3:07 Be careful saying that out loud in this time.
Tom Redman 3:09 I’m well aware, but I grew up in the Midwest, but my first major league game was at Yankee Stadium. My grandfather lived on Long Island, and we’d go visit every now and then, and he took me to a game. Okay. You know, and you’re hooked, right, particularly when you’re nine at
David Sweenor 3:27 at the majesty of that stadium. Super cool. Thanks for sharing that. You have a session coming up tomorrow, I believe, and you talk a lot about how people and data are the biggest barriers. So tell us a little bit about what are those barriers and why do they exist today?
Tom Redman 3:49 Well, so first of all, my latest book is called People and Data. And I want to emphasize that people is first. It is the first word in my most important stuff. But what I’ve found over the course of my career is a bunch of little epiphanies. Things that, you know, as I wandered around and I bumped into things, I saw them in a fresh light. Okay. And I was lucky. Many of these things, I was the first person to see them. But one of the things I’m going to talk about tomorrow is this. Is that… Let’s call people with data in their title data people, and everyone else regular people. And I’m a little tongue-in-cheek with that. People who are smart enough not to have data in their title, but in any organization, there’s way more regular people than there are data people.
David Sweenor 4:47 And furthermore, they do all the data work.
Tom Redman 4:50 right so right i mean you know it’s their who defining spreadsheets they’re they’re finding errors and correcting it in databases and so forth they’re working with their customers or not right to to to make things better and and uh They all have two very important roles. They’re all data creators. They’re all data customers. We found that when you get people aligned on those roles, quality improves quickly. It’s almost like you get the organization stuff right, and then quality is easier.
David Sweenor 5:31 Yeah.
Tom Redman 5:33 The vast majority of work on data and work here on data is going to involve what people with data in their title do.
David Sweenor 5:43 Right.
Tom Redman 5:43 Okay. And so the opportunity, first of all, where the action is, the action is with people without data in their title. Right. Regular people. Regular people, right? And so that’s what we have to organize our data programs around, is getting as many of them to see their roles in data. Right, to become good at those things. Getting the regular people to become data people? No, getting the regular people to take on data in their existing jobs. Okay, so like for example here, I go around and I talk to people with boots. Okay, and I talk to them and then what they want to do is they want to scan my badge. Sure. Okay, marketing people are here and they’re scanning my badge for sales people. And they say, well, we get measured on the number of badges we scan. But the last thing the salespeople need is an unfiltered list of just names. Yeah, they don’t want hot lead on every one because it’s meaningless. Right. There’s very, very few hot leads. Just think about the data flow and just think about how screwed up the organization is, is that one group is incented to make meaningless work for another group. That is how they’re measured.
David Sweenor 7:08 I’ve been on that booth many times.
Tom Redman 7:09 You have, right? And so this is an awful lot about what data quality is about. Right? Who is the customer? What do they really need? Right? And I spend a lot of time talking to salespeople that I don’t understand sales very well. And the last thing any one of them ever says is, we need long lists of unqualified leads.
David Sweenor 7:30 Right?
Tom Redman 7:31 Okay?
David Sweenor 7:32 And by the way, this is obvious once you say it.
Tom Redman 7:37 It is obvious what you need to do to fix it. You need to change the incentive. I’m sure there’s lots and lots of different ways to go about it. But organizations seem to be blind to this.
David Sweenor 7:53 And so this epiphany
Tom Redman 7:55 that that it’s regular people that are doing all the work right the marketing people are creating the leads but they’re not trained on what’s what what it means to do a good job doing that that’s sort of like you know an example and a big simple step one okay all right all right so maybe i’m going to switch gears a little bit but
David Sweenor 8:20 The Datadoc. Tell me how that name came to be.
Tom Redman 8:24 Well, the Datadoc, right.
David Sweenor 8:26 You’re known for this, right?
Tom Redman 8:27 Well, I am known for this, right. And by the way, it’s been a very effective brand. But one of my clients, you know, just I was helping them with something, and they thought whatever I did was clever, and they said something like, you know, you’re like the Datadoc. Right, right. And I took it home and I chewed on it a while. And then probably six months later, I had the courage to put it on my business card and see how people reacted to it. And everybody had a positive reaction. Oh, that’s great. And so I kept… And then I began… Well, think about what does the Datadoc mean? What does that brand mean? What does it mean in people’s minds, and how am I going to build on that? And it’s been enormously effective. I’d like to say I had credit. I had the brains to think of it on my own, but it’s simply not true. Yeah, that’s great.
David Sweenor 9:17 And so, you know, 30 years you have been talking about data quality, and data has always been a mess. Will it always be a mess in the future? Are we ever going to solve the data quality issue?
Tom Redman 9:32 Okay, so that’s a really interesting question. First of all, I think that there are layers of data quality issues. Okay. Okay. And the example that we just talked about of passing on good leads, right?
David Sweenor 9:48 That problem is solvable now.
Tom Redman 9:50 Right. That problem is solvable now. And organizations are loaded with problems that are solvable now. Now the next layer of problem might be around metadata and things in a data catalog or data definition. And in many respects, they seem a lot like the marketing and sales. Those ought to be solvable, although we have fewer examples. So by the way, when I say the first category is solvable, I know it because I’ve helped a bunch of organizations do so. But then there’s categories like, well, you know, can we establish a shared vocabulary such that everybody in the company means the same thing when they say customer, right? And I suspect that those are going to bedevil us for a long, long time.
David Sweenor 10:51 There’s always different incentives. I always tell people this story, so I did yield characterization, right? Yield is pretty easy on a semiconductor wafer. The number of good die over the total number tested. There were three systems. They all had different yields for the same wafer. The only way I was able to fix that problem was to turn off two of the systems. And it wasn’t that they were wrong. They just had different contexts and different uses. One was like, what are we going to ship? We could ship partial goods, half goods, quarter goods, whatever. But the business couldn’t agree. So we had to just shut them all off because of that reason.
Tom Redman 11:27 I mean, it really betrays something. There is a richness and a nuance in human language, right? And human language grows and changes. Like thousands of words are added to the dictionary, right? Every day, or excuse me, every year, right? So I don’t know. I don’t know if we’re going to be able to get in front of that. Now, I do think, by the way, if you establish a common vocabulary on, let’s say, 100 key terms, It’ll do your business enormous good. And that should be reachable. Companies, you’ve got to have a business reason to do it. And you’ve got to be able to align people to do it. And actually getting the definitions is not that hard. It’s getting people to start using them that’s way harder. But I just think there’s going to be layers like that. And when we solve that, who knows, the next generation AI may require something different. I’m talking about knowledge graphs now. I don’t think you do knowledge graphs without common vocabulary and so forth. I hope in the next 10 years we really get in front on that first category. Yeah, absolutely. And we get in front of that. I know there’s a big deal between structured and unstructured data. I don’t think that there’s anything philosophically different in terms of what you need to do to get in front of those kinds of problems for both structured and unstructured data.
David Sweenor 12:58 Okay. And maybe one last question is that you have… You had a study that said 3% of companies sort of met basic data quality standards. Right. Are we going to rerun this sucker? What would the results be if we were to? Are things improving, do you think?
Tom Redman 13:22 So let me just give a little background on what we do. So I do a bunch of teaching, and sometimes I teach in public settings and sometimes in private settings. And when I teach in public settings, then I give people assignments. And one of the assignments I ask them to do is a Friday afternoon workshop. And, you know, you can look it up and you’ll see what it is. But basically, it asks people to just look at the data and with a red pen identify the obvious errors. Right. And then count stuff. Okay. Okay. And usually I ask people, well, how good does your data need to be? And almost everybody says, well, these data are about health care. They need to be 100% or people die.
David Sweenor 14:03 These are bad money. Like there’s a lot of machismo around them. Right.
Tom Redman 14:07 Well, and, you know, and the average is coming in that’s sort of like, you know, only 50%. People say they need 100, and they’re getting 50. And nobody’s ever said, well, we need less than high 90s. So the 3% number came, I don’t know, maybe we had done a bunch of these things, and only 3% made that high 90. Okay, so that was what we published. Maybe we published that in Harvard Business Review 10 years ago. I continue to do that. Still reference it. Still reference it. By the way, it is the best we have, right? And a lot of things spring from it. But not yet publishable has been the following. And my teaching came to a close during COVID. But so far, without a rich enough data set to publish, things are no better. And maybe a little bit worse.
David Sweenor 15:05 But that’s not yet a publishable stat.
Tom Redman 15:08 And by the way, this is on the most basic problems.
David Sweenor 15:13 The marketing, sending bad leads to sales kind of stuff. That makes a lot of sense. Okay, maybe the final question here for today is, people are going to watch your session tomorrow. What’s the one thing you want them to walk away with or take immediate action on when they start their work day the next day?
Tom Redman 15:33 So too many people with data in their title, too many data professionals blame regular people for not getting with the program. And I want them to see that we’re not going anywhere without large numbers of regular people. And I want them to adjust their attitudes just a little bit and start thinking about how we can enroll them in our efforts and or we can join their efforts.
David Sweenor 16:06 Okay, so data people and regular people unite and talk. So Tom Redman, the data doc, thank you for joining me on the Data Faces podcast. Thanks for having me.
Tom Redman 16:18 Cheers. Okay.

