Alexis
Hi everyone. Welcome to Talk Tech with Data Dave. I am Alexis, your host of this podcast, and I’m here as always with my friend, Data Dave. Hi, Dave, how are you today?
Data Dave
I’m very well, Alexis. How are you?
Alexis
I’m good. I was just thinking about my little doggie as the weather is changing as we lead into the new season. And it’s kind of cold in my house, so I got my sweater on today. I’m feeling excited. But we’ve got a fun kind of philosophical question today. Are you ready to go for it?
Data Dave
Oh, absolutely. That sounds entertaining.
Alexis
That’s awesome.
I like to tell our listeners, if you have a question for Data Dave, please email it to us at talktech@d3clarity.com. You can send us a question right on the D3Clarity website or you can reach out to either one of us on LinkedIn.
Dave, that’s actually how we got this question today. One of your followers sent you a message on LinkedIn and they asked this question. I’ve done it before and I’ll do it again. I don’t want to mess up this person’s name [Ijaz Ahmad Ijaz], so I’m not going to share it on the podcast here. But we will add his name in the comments because I want to make sure that they’re getting credit for the question they asked. But I also don’t want to pronounce it incorrectly. So, my apologies if you really were hoping to get your name on the pod. Check the comments. It’ll be there.
Okay, here we go. Dave, here is the question. How do you see companies balancing AI, artificial intelligence, with data ethics over the next five years?
Data Dave
Okay. Wow, that’s an interesting question. That’s a pretty loaded question. There’s a lot baked into that.
Alexis
A little bit philosophical question of ethics. I like this. This is a good conversation for us.
Data Dave
This is going to be a very subjective answer because this is just my thoughts on the whole structure.
So let’s talk about data and ethics for a moment. Let’s start there. So, split this question up a little bit. We all make decisions frequently and use all kinds of data to make decisions because there isn’t necessarily a pool of data we can try to pretend that we don’t use all our knowledge to make decisions in everyday life, but simply your preference for the color blue that causes you to buy bathroom tiles might make it the less efficient bathroom tile or whatever.
So that is use of data that isn’t necessarily appropriate to the purchase of the tile. It’s not inappropriate, but it’s just a flippant example of where your preference and your history with the color blue might cause you to buy bathroom tiles, even though they’re not the hardest or the easiest to clean or whatever. So then you get the rationality of decision-making.
What data is appropriate for what decisions becomes a big ethical question. And that’s the foundation of some of the HIPAA legislation. It’s at the foundation of data access, data privacy, and that sort of thing. I have the right to be forgotten, GDRP and various legislations like that. I have the right for you not to use me in your decision-making process.
That becomes an interesting question. And we’ve got to evolve more in our sort of data lineage and in understanding of how we collect data, both as people and as organizations and as machines. As we look at the explosion of data and the explosion of cross-reference data, the ethics of what data should be allowed to be used, allowed to be seen, even allowed to know that it exists in order to make a decision is an interesting question. And I think we’ve got to really, as a society think about that and start to address the ethics of what data can we use to make a decision? What data should we use to make a decision?
And as we are all people, we’re all human, we’ve all got a slight bias. It might be a slight bias for the color blue, a bias for a rainy day, whatever it might be. We’ve all got a certain amount of bias within our decision-making, which is fine if it only impacts us, but it becomes less fine when it impacts other people and starts to violate other people’s rights, opinions, policies, thoughts, et cetera. That becomes an interesting conversation.
And so I think we’ve got to do a better job, certainly as organizations, certainly as society, to start to build this data lineage, this data governance framework, that can explain the lineage of this data and therefore the lineage of this decision. Makes sense?
Alexis
You keep using the word “lineage”. You’re talking about the lineage of the decision, the lineage of the data. Talk to me a little bit more about what you mean when you say that.
Data Dave
So. in the art world you use the word “providence”. What is the providence of this piece of art? Was it really painted by this painter? Who owned it throughout the entire providence and life cycle of that artwork? Right?
Alexis
So the history.
Data Dave
So the history behind it. Yeah, the lineage behind it. How did this data get to you? Where did it come from? And how did it get to be used within this context? Why is this piece of artwork worth a hundred million dollars?
Why is this data used in this decision? So that’s the lineage and how do I trust it?
I don’t want to go here into the sort of tenets of trust. How do you trust a piece of data? That’s probably for a different podcast, might be an interesting one, but I do think there’s the ethics of data usage and the way that we use data and the lineage of that data. So, we should be checking the providence of all our data because as we know, the propensity of the Internet to be less than truthful and the propensity of data to be less than trustable, shall we say, is very high. To a certain extent, it comes to the ethics of the decision-maker. If your intent is to deceive, then that is unethical. If your intent is to deceive, your intent is to cheat, your intent is to steal, then that is unethical by definition.
If you cheat, lie, or steal without intent because your data is flawed, is that any more ethical?
Alexis
That’s a very good question.
Data Dave
Now, we’ve got a responsibility to try and police ourselves to ensure that we have the right lineage and the right providence of decision so that we can be assured that our intent to not deceive hasn’t been violated by accident.
That’s where I think the data governance area comes in, the data providence. The lineage comes in to start to describe our data like that. And I think that’s a key point. With that intent, over the next five years, I think we’re going to see more and more controls and legislation and less tolerance of people not paying attention to their decision-making lineage.
Alexis
Yes.
Data Dave
So I think that is key.
Now let’s talk about AI for a minute. From an AI perspective…
Alexis
The initial question was how are we going to balance those two?
Data Dave
Exactly. So, AI gives us a challenge because AI, certainly public domain AI, does not give us that lineage of what data was used to make that decision.
If you ask AI to write you an email, you could probably specify which dictionary it used.
Is it US English? Is it English English?
Alexis
Right, English English.
Data Dave
The little grammatical changes and that sort of thing. You could probably specify that, but you don’t know when you first go there which one it’s going to use.
And do you think about asking it or constraining it enough to make sure it’s using only the right data?
Alexis
I don’t, I’ll tell you that much.
Data Dave
No, exactly. A lot of people won’t. If you play that data lineage and the decision lineage into AI, how can you be sure if you’re automating this decision, what knowledge it’s using as a basis for that decision?
Again, if you control all the data that goes into your AI engine, then you can be certain that it only used this carde of data. But if you expose that AI engine to the wrong audience, then you’re going to get decisions for that audience made with the wrong carde of data, potentially, or a biased carde of data. You also can’t necessarily constrain it after that, so you can’t trace through the AI engine back to the data, the knowledge that was used to create the decision.
Again, I think we’re going to see more and more data governance and people less tolerant of bad decisions, for want of a better phrase, or violating decisions, unethical decisions, unethical behavior.
How do you detect unethical behavior in that construct? How do you look at an image generated by AI and start to say this was generated with intent to deceive? How do you know that it was generated with intent to deceive?
You can generate an image and use it to deceive, but the AI engine didn’t know you were going to use it to deceive. That’s in the usage structure, not in the generation structure.
Alexis
Well, I want to stop you there, and I want to point something out just from my use of Gen AI.
So, I regularly go to ChatGPT, go to Dall-E and say, hey, create me an image to go along with this blog, to go along with this case study.
And when it puts people into the image, they’re almost all white, and they’re almost all males.
Data Dave
Okay?
Alexis
Because all of our blogs, all of our case studies are about technology. And for some reason ChatGPT, Dall-E believes that the only people who care about this are white males. And so, I’ve gotten to the point where I have to prompt it and say, hey, I want a diverse group of people in this image or I want women and people of color in this image as well as white men. I don’t know that that is it intentionally trying to deceive, but it does lead us to the conversation about bias that’s already in AI.
Data Dave
It absolutely does. And I think that’s notable on yourself because you’re actively correcting the bias by changing the way you ask the question if there is bias in that data set. So, I’m not going to state whether there’s bias in ChatGPT’s data set or not. I’m not going to address that question.
Alexis
I don’t have enough experience to truly give my own opinion on it. But I can tell you from my activities, this is what I’m noticing. Yeah.
Data Dave
Evidentially- from the evidence- it would suggest that it usually returns these. So, you are consciously making a decision to refine that for your own purposes. And I think that is a responsible use because what you’re essentially saying is, I’m going to use this tool, the tool is perfectly reasonable, but I’m also going to second guess the tool for my intent, because my intent is this and my intent is to be inclusive. And so, you’re second-guessing the tool with a human decision-making process, which I think is perfectly reasonable. So, are we going to see people start to put that kind of filter on top of AI as well?
Because realistically, you could say, okay, I’ve asked you a similar question the last 15 times in the last seven days. There should be this much disparity, right. Just randomly across that.
So we could put AI in to make that decision. But I think this is the point, the responsibility. I don’t believe we can delegate the responsibility for ethical behavior to AI or to legislation. We have to take that on board as individuals and say that we are using tools in order to behave ethically, and we will second guess that tool. I could buy a hammer to break a window. That is unethical behavior with a perfectly ethical tool.
Alexis
Right.
Data Dave
Or bad behavior with a perfectly reasonable tool. Right? Yeah.
Alexis
I mean, it depends on why you broke the window and where you broke the window. But yes, I’m there.
Data Dave
But it comes with intent. Right? We are responsible for behaving within an ethical framework, and we have a responsibility to do that and to present ourselves in an ethical way. And we have a responsibility, I think, therefore, to justify and be able to justify our decisions, if not to anybody else, then at least to ourselves.
Yeah, to me that’s kind of the root of it, which is now I want to know, if I didn’t make the decision, how do I trust that the decision that was made on my behalf is made within the framework of ethics that I believe in? This is why we have corporate values and why we have all kinds of things. It’s all about- can I trust the people that I’m entrusting with a decision, can I trust that their decision is made within the framework of trust that I would put out or the framework of ethics that I would put? And I think that’s the question we have to ask AI and look at AI and look at the data and look at that lineage of structure.
So through that data and lineage of structure, to get to a decision and second guess it and understand that we do trust it to be making that decision that we would make. And I think we’re going to see, when we put a time frame on that, the continued need for ever more data governance. We’re going to continue to generate more data. That’s a fact.
We’re going to use more data in our decision-making. Everybody demands more data in our decision-making. Everybody demands that their context is brought with them in everything. We like to be offered only the concert tickets for the bands that we like. We like it when people do that. I’m just making stuff up. But we’re demanding that. Consequently we have to pay attention to that structure as well and make sure we can understand that and change it. Because if my music tastes change and maybe I want to be invited to a different concert, I do actually want to know that those concerts exist, even though I don’t want to go to them.
I know these are flippant examples, but I think in the next five years we’re going to see more interest in this and more people introducing the idea of data governance, information governance, information applicability for decision making and that sort of thing, because you, you can’t afford for decisions to be made automatically that aren’t within your sphere of ethics, so to speak.
Alexis
I do want to put a little asterisk beside this podcast and say that, you know, we’re recording this in a time when GenAI is exploding, when we’re talking about it pretty much every day on every podcast and everything we do. But like, maybe in a year if we had this question again, we might feel differently or we might have a different answer. I don’t know that our opinion on data ethics is really going to change that much, but how we use AI could possibly change based on how the AI revolution continues.
Data Dave
Yeah, absolutely. It’s continuing to change and continuing to evolve, and we’re going to see legislation evolve, we’re going to see organizations evolve, we’re going to see a number of things evolve. So, I think that’s absolutely true. I just think that fundamentally, we are a gregarious species, and we’ve used every means of communication that has been presented to us throughout the history of time, and we’re going to continue to do so. What we have to be careful of, just generally, is that we take responsibility for that communication.
Alexis
Absolutely.
Data Dave
Whether it’s AI or cave painting or whatever. We don’t want to put anything offensive on it.
Alexis
Well, thank you Dave for answering that question. Thank you again to our listener who posed this question. This was a great conversation, and I think we have at least a couple of follow-up conversations that we’re going to need to have now, Dave, so I’ve got those on the list. I really appreciate you being with me here today.
For our listeners, please submit questions. We’d love to answer your questions. They can even be kind of philosophical, opinionated questions. If you don’t agree with us, we can always have you on the pod and we can have that conversation here. That would be amazing. Reach out to us at talktech@d3clarity.com on the D3Clarity website or reach out to either one of us on LinkedIn. We’re both there for you. Thanks Dave for being here today. I appreciate it.
Data Dave
Excellent. Thank you very much. Alexis, as always, a pleasure.