YouTube’s Angela Nakalembe on How to Build Compassionate AI Models
In this Techstrong.ai video interview, Angela Nakalembe, engineering program manager for trust and safety for the YouTube arm of Google, explains how making artificial intelligence (AI) models compassionate is actually achievable.
Transcript
ai Leadership Insights series. I'm your host, Mike bza. Today we're with Angela Nakae, who's an engineering program manager in the trust and safety division of YouTube.
And we're talking about, well, can we actually make AI compassionate? And is this something that's actually achievable? Angela, welcome to the show.
Yeah, thanks for having me, Mike. Glad to be here. I think a lot of this conversation stems from Jeffrey Hinton, who's considered the godfather of, uh, AI was on, I think 60 minutes talking about the need for more compassionate ai, and it seems to me maybe this is just a mere matter of programming or are there other things to consider here, but can we do this?
Yeah, I, I watched that interview and I mean, Dr. Hinton, a true godfather of ai, really, he does raise a profound question about, um, how we can ensure Superintelligent AI aligns with human wellbeing. And just from my experience in the work that we're seeing, the data that we're seeing, I would caution on the maternal instinct analogy, because while it is striking, I think it, um, it also underscores a deep desire, I guess, for AI that inherently like prioritizes humanity, though I don't think that encoding, um, like a biological social construct, um, is probably like the way to go about that.
But yeah, excited to dig in a bit more about that. Yeah, So what would be required, I mean, is this just a set of guardrails and policies that we need to develop and, um, like most guardrails and policies, wouldn't somebody just eventually end run them anyway? Yeah.
When it comes to, um, ai, anything like the maternal instinct of it all, I do wanna start by saying I did score AI is just, you know, a series of pattern recognition in prediction machines, right? They don't possess like the biological instincts or consciousness or like a human, like behaviors and emotions that would need to kind of have the empathy or compassion that seems to be communicated in, in that interview. Um, but what we can do, right, is focus on building and developing models that act in ways that, um, less about feeling compassion and more about aligning with like human centric values, right?
It's more about engineering. How do you engineer these models to have beneficial outcomes and predictable safety before deployment? And what we've, that's a really big thing that we've been doing, um, at YouTube, right?
And to your point, there is a lot of work that goes in from the research level, policy level, um, the engineering level testing is a really, really big thing. And then also having some sort of like break glass functionality, you know, just like a kill switch in case everything does go wrong, you know, you have to be able to pull the plug, right? Um, and so that's kind of, those are like the top level areas that we've kind of been, um, prioritizing as we build out our governance models for responsible ai.
And do I have this right? But it seems like maybe we've got the cart before the horse these days, but people are trying to apply guardrails for this thing after the AI model has been deployed, and maybe we wanna do that before the AI model is built, or how does that kind of work in your mind? Yeah, I, uh, it's easy to see that happen.
Um, you know, like to see all these issues that are happening where AI is, maybe it's hallucinating or it's, you know, talking about how rocks are healthy to eat. We've seen all these articles, um, come out and AI having to be models having to be reeled back and improved upon. But what I do wanna underscore is there's actually a lot of work that goes on, um, before all these models come out, right?
There's multiple layered defenses that not just at Google or YouTube, but multiple other, um, AI companies that we're trying to instill as the facto methodology for rolling out these new systems, right? And it starts with input filtering, right? When you're building out these models, how do you train these models to reject harmful prompts?
Um, and in the model itself, how do you train that to avoid generating harmful or misleading information? And once we get to that point that we feel good, that's when we usually deploy the models. However, this, you can't, despite our best efforts, you know, it's this, something's always going to like fall through the cracks.
And that's where we get the concept of output filtering, which is how do we scan and redact, um, harmful like outputs or post generations, um, for removal and make sure that they aren't replicated, right? And that's actually something that at least at YouTube we're doing using AI itself, right? So it's, if we're getting harmful, like misleading information put out by a model, how do we, um, train our classifier models to, we use human reviewers, actually, we use human reviewers to kind of, um, document to why certain responses were wrong and that information gets put back into the AI model and help train it to be better.
Um, and so it's, that's kind of like the cyclical process that we're talking about. And it actually brings me to a point that I'd forgotten to mention. It's, as we think about building these safety guardrails, one indistinguishable fact is, um, the human AI alliance is just like a non-negotiable, right?
It's, um, AI is fast becoming a lot more intelligent and you know, even like some of our smartest researchers, and we do, you know, see a point where it will be, you know, as Jeffrey was mentioning, like, it, it's, it's concerning, right? Um, thinking about what the future will look like. But one thing we know for sure is that there will always need to have some sort of human, um, dependence on the models that we're building, or like human AI collaboration, um, just to have that symbiotic loop to ensure that AI continues to serve in our best interests.
Yeah, I mean, to your point about that, otherwise we wind up using the data created by AI to train the next iteration of the AI model, and it just becomes a loop, right? Yeah, exactly. And that's where you get, you know, the AI slap on top of slop, you know, and that's what we're trying to avoid.
Yeah. Mm-hmm. Um, what is the definition of compassionate gonna be in your mind?
Because, you know, there's clearly compassionate in the sense of we're all gonna be nice to each other and have empathy, but sometimes, you know, um, we have scenarios where I don't know, you know, the horse is injured and we have to put the horse down, and that's considered compassionate because the horse can't have a life. But where do we kind of define that spectrum in a, in an AI context? Yeah, that's a good question.
I think the, the simplest way I can think about it, just think about explaining it, is compassionate has to mean human first and having AI models that serve in the best interest of, um, humanity and in the continuation and prosperity of humanity, right? Um, and one of the ways you kind of talked about, um, compassion could mean putting a horse down. I think that's kind of where we as humans have to think about the guardrails as we're building these tools, right?
We can build them with the best of intentions, but you never really know, um, what you don't know, right? And that's why I talked about the importance of having a kill switch or some sort of, um, emergency protocol that can be used to easily shut down a program, you know, and that's something that we'll have in our research labs at home. It can be, you know, something like parental controls that parents put on their kids', um, devices and such to limit what sort of content they can like search or conversate about with their, um, AI chat bots, that type of thing.
And so it's, it really at various levels, just always remembering that we do have to, you are interacting with a machine at the end of the day, and you do have to, um, you do have the power right to kind of step away, kinda like compassionately. Yeah. Like put an end to, um, the conversation or, um, whatever, like you're leveraging it for.
Mm-hmm. Um, it's pretty clear that AI malls have personalities and sometimes they're chantic and other times they're less so, um, you know, some people, like the ones that say, you know, boy, aren't you the most awesome person in the world? And others are kind of like, stop blowing, smoke up my skirt.
But, um, isn't compassion part of that personality and how do you kinda like weave that into the way that the AI model manifests its personnel? It's, yeah. That it is, uh, it's a delicate, it's like an art and a science right at that point, because you do, a large reason why we're putting these models that are like we're working on them and fine tuning them is to provide a need to fulfill a gap that we're seeing in the market, right?
There is that need for people to have, you know, somebody who we were talking earlier about, um, like parents and growing up and maybe not having had that person who believes in you or who supports you, right? And it's important to kind be able to, in encode that into our models, but then also encode the optionality for somebody to be able to, um, ask it to tone it down when that's not necessary, right? And so that's what we've been emphasizing and in whatever we're doing is the increasing the optionality for the models.
But one thing I'd love to talk about here is just encouraging the users to similarly educate themselves, you know, on what they can and cannot do with these models. Like a lot of people don't even know that there's a way they can prompt engineer their models to be less kohan, you know, or to avoid certain, um, habits or tendencies. And when you do like, kinda like inform them, you kind of like, it's a big mind blow.
Just like, oh, I didn't know I could actually like, you know, ask my model this question and then tag at the end, give me objective feedback or, you know, something like that that can just actually help. And so I feel like that's, that's kind of where it comes from, is, um, I think these models will always be developed with optionality to fulfill a bunch of different desires and, and like needs for our users. And it's really up to us to be able to, which is actually anything too, like, to know that you have the power to define and tweak how you want that model to engage with you.
Is there some way, as I consider what models I wanna use to evaluate the level of compassion and other personality traits associated with the model so that I can know this before I go build my app and not after? Mm-hmm. Um, that's a good question.
Nothing comes to mind at at the moment, but I'm fairly certain there must be some sort of scale that compares, um, models across, um, providers and also even just within the same company, like different iterations of models, right? And their level of compassion. And so I might have to get back to you on that one, but I'm fairly certain there should be something.
Yeah. Well, Clearly there's a benchmark that somebody should go out and build if they haven't already, right? I feel like that clearly you're pointing out a need right here, you know, so if that's not out there shoe, we might have a business idea on our hands.
Mike. There You go. Um, so what's your best advice to folks who are, um, first building models and then b the folks trying to consume that model?
What should they be thinking about as they go do that? You know, hopefully sooner than later? Um, I think the one thing, it's something we talked about earlier is that human AI connection, right?
And just establishing a robust team from get go of like a cross-functional team that you're gonna be building with and not just your engineers and your researchers. You want policy, you want legal, you want a diverse data set as well. And 'cause to help limit bias, right?
Um, of, of what the model's gonna be saying. That's a base level, right? Um, table stakes.
And then once it, once you go up, you want to think about transparency and explainability as you're building these models. You want to be able to build them in a way that they could easily explain, um, why they're giving certain responses as opposed to just giving a response. So you can kind of extract to a certain extent the, I guess like a level of intelligence and the thought process of these models, right?
And you can kind of work backwards if you need to. Um, I think another area that we've also been emphasizing too, that I cannot stress enough is auditing and testing. Um, a really big effort that I've been involved is, is adversarial testing, which is what I mentioned earlier, right?
Where you give your model a whole bunch of prompts to kind of just pressure test to see if it'll give you anything violative, um, or harmful. And then using human reviewers to kind of train the model to explain to it why this is bad so that it can be better and, and avoid sharing, um, results that you wouldn't like it to before deployment. I think just that pressure testing has been incredibly helpful on our end and we've gotten really good feedback, at least from the YouTube side and creators, um, being able to kinda like confidently use these models to generate content for their users or even internally, we we're using these AI systems to increasingly review content with like up to a 96% accuracy, right?
Which is fantastic if you're able to kind of, um, automate a lot of resistance. It frees up a lot of, um, human capital for us to redirect that to other efforts, you know, that we're building and growing our, our platform and our user base. So those are things that I, advice to anybody who is kind of like getting into building with AI or building models is just one, focusing on the team that you're building and making sure it's a very diverse team input from everywhere.
Very diverse data set, as much as you can get your hands on from diverse perspectives, um, that adversarial testing multiple layers all throughout the process before deployment is also important. And then just building that explainability in there just so you can be able to, um, defend where those answers or responses are coming from. Mm-hmm.
Yeah. Um, and then I think your other question had to be around from like a, a human standpoint, right? Like a everyday person.
Like how, what are things that we should be, um, cognizant of? Like what should we be thinking about? Um, I think a really big thing that I've been having conversations with, uh, with people in my life is about, uh, children, how do we keep children safe?
You know, um, this age of ai, a lot of Gen Z or Gen Alpha people are digitally native, right? And they know a lot more about this technology than we, you know, like their parents or their educators ever will, right? And I think that the first point is one, taking the time to actually educate yourself, um, about this.
It's been something that I've had to carve out time for every day and I work in this space, right? I've had to sit down and kinda sit and ask myself, okay, cool, so how does you know this new model from this company differ from what we have here? Um, what are, you know, like how are defects made and why are these a lot, you know, like kind of just being able to identify, you know, what all these defects are and explain to somebody why something might be dangerous.
Um, to, and not even just like kids, but like even older people in our lives, right? Like my grandparents have had to explain to them why certain, whether it's like political odd or something might not be accurate, just because, you know, explains to them what a deep fake is so that they know that certain things aren't accurate. And teaching them how to identify what this, what these are.
Um, and also just teaching people to be more careful about what information they're sharing online. Because again, all of this can be used to train models at Google, we have very, um, strict rules about the type of content that we're using, um, to train our models or data that we're using to train our models. Um, but that isn't the case everywhere else, right?
And so a lot of, you know, whether it's pictures or personal information that you're sharing could then be used to generate content that, you know, might not reflect well on you. So just being extra diligent right, is just one of the biggest pieces of advice. Um, I could, I could give out to people.
Um, but then also like, get your hands dirty. I think that's the, that's the good thing too. It's, it's, there's a lot of doom and gloom about ai, which is very rightfully so, right?
Um, but I think there's also a lot of good, and I don't want us to lose sight of that. There's a lot of, uh, boundless possibility coming out. So get out there and try to find, um, everyday ways that you can use AI to make your lives easier.
Um, test it out. Um, talk to the people in your life about how they're using it to make it easier while also, you know, protecting themselves and their data and, and yeah, just go at it. Well, let me ask you this one last question 'cause I'm been kicking it around in my mind lately, but of course I think, I think we're gonna wind up using multiple AI models for different tasks, but they might have different levels of, shall we say, compassion.
So how do I avoid a situation where essentially now I've got a bunch of AI models bickering in the back of the car and I'm screaming at them, don't make me pull this car over because they're bickering at each other. Um, I think it kind of goes back to what we, I was just talking about, you know, like getting your, getting in there and just trial and error, right? Trying out the different ones and seeing, um, which models you like better for certain tasks.
For example, and I hope it's okay for me to share this, there's, there's certain tasks that I prefer using, like Google's Gemini model for there certain tasks, I absolutely would never, you know, like questions I'd never ask, you know? Um, and I'd prefer using, say, Claude for coding or I'd use, you know, like chat GPT for something else, right? And so just being able to test it out and figure out, um, what task, whether it's engineering or whether it's, um, personal knowledge management or, um, just general like quick, quick queries, you know, figuring out what you, what works best in what situations so that you know which model which, which models to reference when you need certain questions or tasks addressed.
Does that answer your question? Is there like a, a specific piece to cover that? No, I think, um, I, I think parental supervision will always be required.
Yep. Uh, unfortunately that is the case. There is no like magic pill to kind of make it all go away, but yeah.
Yeah. There You go. Hey Angela, thanks for being on the show and sharing your insights.
Of course. Thanks for having me, Mike, and have a fantastic day. Alright.
And thank you all for watching the latest episode of the text on ai.