How Can We Trust Generative AI? | DevOps Experience 2023
AI is eating the world! And just like when software started to eat the world, the world isn’t quite ready. AI capabilities are being released at break neck speed and the world is starting to realize what’s at stake. From increased adversary opportunities to a shortage of water supply, is the industry really putting in the investment needed to make AI a durable capability? Trustworthy and durable AI is not backed by assuming legal responsibility; it is created through an effective value chain with all the -ilities in place.
Transcript
Hi, it's Shannon. Lets, uh, really looking forward to talking with you about AI today. A little bit about me.
I'm the c e o of third score, and I have been working in the industry for over 30 years. I consider myself a tech and security dinosaur, but with an emerging trend towards what's happening next and really thinking about how do we make things safe, capable, and really, um, solve for human needs with software and technology. So today I'm gonna talk to you a little bit about how can we trust generative ai.
Uh, it's a really awesome capability. I'm super psyched about chat, G b T, open ai, uh, claw bard, you kind of name it llamas. Uh, and nowadays I think there's so many coming up.
There's, there's hundreds of companies that are starting to really embrace and use these generative AI capabilities. So, so what do we need to know about it and how do we really think about how we trust generative AI and, and what can we trust it with? Um, in particular, you know, my research actually covers some of the ways in which we might measure trust.
How do we actually assess for trust and really how do we best leverage this technology for our future, and even in today's times? So I'm gonna talk to you a little bit about, um, you know, let's talk a little bit about Google and a little bit about the trends that have been happening. If you haven't caught up with some of the generative technology out there.
There's a thing called chat, g p t. It kind of became famous earlier in this 2023, uh, released a little bit in 2022. You know, generative AI became sort of a, a commodity capability quickly, um, in its introduction to the internet and within the first couple of months, uh, chat g p t a hundred million users.
So literally a rocket ship straight to the top of people really wanting to leverage and, and use this technology and, and like, uh, Google search, it's really just got that one bar. You ask it a question, you get back some really interesting answers and fascinatingly enough, um, you know, AI has started to exceed the term of chat G P T in Google Trends. I thought that was fascinating.
Around April timeframe, you sort of saw AI go back to being more interesting than just chat G P T. And I think that's because we're starting to see the emergence of not just chat G p T in this space, but a whole bunch of other capabilities and conversations around what is generative ai, the introduction to everyone in the world about what can we do with this new creative set of technologies. Um, fascinatingly, you know, I looked up, uh, artificial intelligence didn't, it wasn't even as popular as ai, but, um, and generative ai not popular at all, even though that's really what, uh, we consider chat, G P T and others like it, uh, to be so really interesting.
And so that's sort of the wake up call, the awareness call is that AI is here, it's here to stay, it's not going anywhere. It's gonna continue to be something that we have to think about as technologists. And really, the the next question is, well, you know, I've heard all these things.
There's a lot of rumblings, there's confusion, uh, a little bit of, um, uh, misunderstandings about what can this technology do for the community. And so I'm gonna talk to you about that today. So I'm gonna start with the thing that I said earlier, which is in February this year.
And you'll see I've outlined and given you a whole bunch of things to go read. So there's a little bit of takeaway in this deck as well, but, um, adoption was a hundred million users within two months of launch. And, um, you know, that is an interesting number because first of all, you know, how do you really calculate what a hundred million users is?
And there are billions of users that went to Chad g p t to understand what it was. There's a whole bunch of data and telemetry around, uh, how long does somebody use Chad g p t within minutes, you know, they might come ask a question and then they're gonna bounce. So adoption is a really popular number.
And why is adoption a really popular number? Well, investment commonly seeks for adoption of technology. So you'll see open AI has actually had quite a bit of investment within the community to advance its, um, G P T capabilities.
And that has led to a desire to really test out this product set and see what it could do. And within a couple of months to see a hundred million users, that's the fastest growth rate of any technology out there within that short period of time. And now actually they say that there's about 200 million users that, uh, leverage and use chat G P T on a recurring basis.
That's a, again, a fascinating number, really useful. So why is that number not enough to get us to think about trusting chat G P T? Just 'cause a hundred million people actually went out and leveraged it, um, have asked it a bunch of questions.
First of all, that means that this particular engine is actually getting trained and training is really important for the value of a, of an AI engine. Um, it means it being trained constantly and there's actually continual information being pumped into chat G P T to make it more and more valuable. 5 and chat G p T four.
And you know, thinking about what that means, there's a whole bunch of interesting information about the models and their versions and what goes into training them and how much more useful they are with chat G P T, uh, four o really being the cutting edge of the technology these days. So I also think that something interesting kind of comes out from the conversation. Once you start to see adoption, more and more researchers out there start to think about, well, what are the repercussions of this technology on humans?
And so the first thing that really kind of like caught my attention was just how much water is being used by say, chat G P T or other AI technologies. And so you see a whole bunch of researchers out there actually talking about water use when it comes to compute. Now this is not something that's, uh, out of the ordinary to see water being talked about as part of the compute that's out there.
The public cloud spawn this conversation. And so I think it's really an important factor. Did you know that, um, some of the researchers estimate that about five questions in, you've actually used somewhere in the neighborhood of 500 mils to 16 ounces of water per conversation, uh, that actually would have five questions associated with that.
I think that's really interesting because, um, that means that effectively we're seeing actually humans vying for water resources, the scarcity of that water, also being in direct competition, we say the compute of the world, and there's gotta be some level of resilience aspect to that. And so resilience against adoption, what's gonna kind of win out? So adoption tends to, to start the spike.
And then resilience really kind of comes in afterwards and people start to ask those questions of, well, is this a durable capability? And water kind of starts that conversation with ai. Um, and you can see here there's a little bit of, uh, information about how do data processing centers actually leverage water, electricity, uh, you know, the way in which they might chill, compute so that they can get the best, uh, performance out of the technology.
And so this is something again, you know, read on your own time, but really interesting and there's definitely estimation techniques. So this isn't like, uh, research that I would consider to be factual, but very interesting nonetheless. And that sort of brings with it, uh, the notion of, okay, so we know resilience has got a question mark, what about, um, you know, delivering malware to thousands?
We've heard that chat, G P T and uh, some of these other technologies might be able to be leveraged for malware. And so yeah, that's true because hey, with a hundred million adoption users, it has caught the attention of adversaries that are out there, but you wouldn't think about it in the same way that they do. What they actually have thought about is, well, it could actually be used for me to deliver against traffic that might be interesting in AI and interested in AI and therefore allowing me to actually capture their, um, use case.
And so because folks don't actually understand how the technology works, they've actually had malware infected installers out there that are using things like chat G B T and Bard and even some of the Facebook information to be able to distribute redline Steeler. And I thought that was another fascinating one because I don't think that anybody intended for this to happen that they'd have such a rocket ship of adoption and that would lead to folks actually leveraging their brands to be able to deliver malware to thousands. Another one, um, you know, there's a, a confirmed data breach that was actually a concern in May this year.
And, uh, it came from the fact that there was a technology that hadn't been vetted and all of a sudden, you know, raising those concerns about whether or not this technology has actually been tested effectively, uh, have there been threat actors out there trying to repurpose the, uh, this for malicious code delivery? And, and that's also another one to consider. What about failure?
So, you know, we know that Steven Schwartz here depicted on the right hand side. He actually used chat G P T to go and look up precedents for his, um, court case and actually got six court cases returned back to him. And unfortunately for him, they were all made up, uh, case law.
And unfortunately for him, he also hadn't checked the case law to determine if they were actually real or fake thinking that chat G P t had this information actually at his ready. And that's another one where I don't think anybody really intended for this to be used as technology to help lawyers to do their jobs at this level. If you look at some of the technology statements down and some of the, um, notifications that are in, uh, and warning labels that are actually being displayed by Chad G P T and others, you can see that, uh, there is no intention yet to support case law and actually determining what could be ca good case law.
So that's another one that pops up for errors. 7% traffic drop according to this article. And what's interesting is as part of this ZD net article that, uh, Chachi PT actually had a downward trend of traffic, and you saw that in the trends from Google as well, that there was a downward trend that's been started with chat G P T.
Now what's also interesting as well, there's not a lot of people searching for it. There's a whole bunch of people now searching for AI much more aggressively. 7, uh, percent traffic drop for resilience is an interesting trend to look out for because it means that the, the durability of technology is in question.
But at the same time, if you look at the numbers in August related to how many people are actually leveraging chat, G P T, they're now somewhere in the neighborhood of 180 million people actually leveraging chat G P T on a routine basis. Another one to think about is, uh, you know, whether or not we're actually thinking about training folks for AI and even these G P T technologies. I think this one was fascinating because as I've been thinking about ai, the question is who's actually training folks and setting expectations and actually creating the adoption patterns that are gonna be successful for business?
And right now, um, you know, gen Z in this particular article, it was noted, and I think it's a Fortune article was noted that one out 10 are offered, um, on the job formal AI training. And this is also something to consider that now when you're actually trying to recruit talent, what is your stance on AI's gonna be a question mark by the candidate and also by the company? So I think that having an AI training program is a, uh, next bold step that has to be taken by many of the companies that are out there.
Uh, if you haven't already, this is gonna be part of setting the trend for your company to have a competitive advantage. And so training is becoming more and more popular as we're starting to see the emergence and the mainstream use of ai. Now, what's really also interesting is that while Che G p T has kind of blown the lid open on how useful AI can be, it's been around for like decades.
And the truth is, is that there's a whole bunch of people out there that have formal AI training, uh, but we haven't really thought about them in terms of maybe leadership in some of those capabilities. So you may even have to take folks that have had that technical training of AI over the years and actually then augment their leadership capabilities as well. And so I think that we're gonna see durability as a question mark over the coming years because of folks not knowing how to best leverage the technology.
And that's something we're gonna have to invest in. So let's like compare generative AI capabilities, and this is another interesting Google trends. I I'm fascinated by all things Google trends these days.
And so Chad, G B t bar Dolly, mid journey co-pilot, there's by the way, a lot of different AI capabilities out there I could have put in here in terms of in, in terms. And I did kind of check them out like, Hey, is Islam a trending or any of this? Well, so the most, uh, competitive now with chat G P T is gonna be bared.
You can see that it's actually on a slow inclining rise. It's starting to take some level of popularity, but it's not really taking away specifically from chat G P T. And I didn't question mark it against AI specifically, but I think it's important to know that it's actually hit at least some level of popularity.
And so Bard is basically becoming, um, more of a challenger with chat G P T in the coming years. And, and that's gonna be an interesting trend to also keep track of because in March, uh, Bard was open to the public. And what's also gonna be interesting is sort of how they compete with one another and, and what does that look like and, and how are we really thinking about best leveraging things like open AIS capabilities, Dolly and some of these other things, uh, fascinating information, absolutely critical to understand.
So what I also thought was really interesting is looking at the velocity of these companies. You know, Bard actually has weekly feature releases according to what they actually put out on their updates. And so if you look nine 19 to 9 27, that's like a seven eight day release cycle associated with Bard at this point.
And if you are using Barge, you can see it every time you go in to use Barge. It's got more and more features. And so what's quite interesting is that these technologies are being developed with the community actually part of the journey and being able to give feedback directly to these engines and to the teams that are inspiring the creation of these capabilities.
Um, I think that that velocity rate is quite aggressive to be able to deliver so many different, um, aspects of a product release in that speedy amount of time. You can see that this is really becoming a focus of energy for the the broader community. So, you know, are we taking the steps in the right direction though, be able to see the question mark around resilience and adoption and velocity and errors?
I tend to believe that trust is built, um, from balancing resilience, adoption velocity and errors. So having low numbers of errors, having a budget associated with error rates, really putting, um, in place some of the capabilities to reduce errors so that your community's not finding them before you do having a high speed velocity, but balanced against the creation of errors and balance against adoption is really critical as long as you're not sacrificing your resilience. And so more and more what I'm seeing is that, uh, the question marks about resilience and errors are really kind of balancing the effects of adoption and velocity in this, in this AI race that's actually afoot at the moment.
So the steps in the right direction. Uh, the US National Security Agency just recently talked about, uh, having companies start to think about how they're gonna secure their AI models from theft or sabotage is, is basically one of the most key things that the US has to do. And especially around, um, uh, use of these capabilities when it comes to other countries and how they might actually respond to, uh, the development and the investment actually being made here in the us.
I think that's, uh, an interesting response. And the reason why I think that's interesting is because really the question around whether or not they're secured effectively, uh, is highlighted here. And being able to dedicate an entire center to helping with securing these capabilities, I think is a big commitment by the United States.
The other thing to think about is, um, in particular that open air, I recently actually invested in a bug bounty. And what's fascinating here is to see that only within this short window of time being able to release this open AI bug bounty on Bugcrowd, they've had 52 vulnerabilities discovered and rewarded during that period of time, and they have like a 75% of submissions are accepted or rejected within about 19 hours. Now that says that basically there are bugs that are getting released into the wild, into the public, and um, then found by the community and fixed by the community or, uh, said to be need to be fixed, uh, from a community perspective that then OpenAI goes back and actually makes the, the commitment and remediates those.
And you can see here within the last few months, they've actually had an average payout. Now it's pretty low average payout, but when it comes down to it, their maximum reward is $20,000. And you know, we've seen some things in, um, chat G P T in some of the open AI capabilities, a lot of them in scope for this program.
I think it's a great conversation to be had around resilience, but I wish that this particular company would actually start to invest in finding these before they're, uh, finding these vulnerabilities before they're released into the community. Another one, um, that's really been interesting worth noting is that the White House recently had a conversation with folks from seven companies about their commitments towards safety, security, and trust. And if you really read into what was published, it's about three pages.
You'll see the article links here, the safety, security and trust are something that have to be part of the D n A. And so the reason why I focused on resilience and adoption velocity and errors is as part of the trust algorithm, people really have to believe that these capabilities are going to be additive, that they're gonna be helpful, they're gonna be able to support somebody who's trying to make a living. And so the, the climate that's happening of releasing technology, it has to be tested for safety, it has to be tested for cybersecurity.
Uh, these companies are starting to realize that they have to make these commitments and volunteer to do them. So I thought this was a great landmark, uh, you know, movement towards actually producing greater resilience among our companies. And not just from a, we have to be cybersecurity safe, but also, you know, how do we best leverage this AI technology and what can it be used for that's gonna be not just popular among us, but actually useful and contributory.
And so this is really important. I also think the other thing to note here, and it's been on top of a lot of parents' minds, is, you know, how do you really shield your children from the harm or the potential harm of these, uh, technologies? What's being done to really make them safer for the community?
So, you know, from my perspective, there's some additional things to consider out there. So the fact that we've got the government involved and there's starting to be bug bounties and we're seeing safety start to arise as a resilience effort, that means that it's been out of balance for a while. And I think we are seeing more investment, more community aspects of needing to have higher resilience and lower errors in the technology that makes up ai.
And the reason why that's important is because you're starting to see some of that commitment kind of come out in the community. And so Anthropic actually has their constitutional ai, if you haven't actually caught up with this, it's quite fascinating what, uh, anthropics done and I think it's kind of built into Quad ai. But basically you'll see in here that there's a constitution that the technology that they're building has to subscribe to and that it actually has to understand the trade-offs from a human perspective so that it can be, um, more harmless and helpful in the way that it actually operates.
And so this one is one to check out. I think that's a really interesting article to look at as well as, and if you haven't seen this one, the Content Authenticity Initiative. This has been going on for quite some time.
Absolutely fascinating work. Um, that's being done. I'm really a big fan of this.
And the reason I'm a big fan of it is because effectively these companies that have come together have been thinking about the content provenance problem. Meaning how do we actually mark information about whether or not something's been generated? How do we actually trust that content to know that it's actually been authored by people that we expect to, to have been authored by?
How does somebody as a creator actually create their mark so that they can, uh, leverage and best utilize this technology? So this one, if you haven't checked it out, also really fascinating and I think gonna be one of the most important that we see as we continue with, um, ai, that anybody who really wants to produce AI technology or leverage leverage its effects is going to have to commit to doing content provenance, not only from the output perspective, which is where this actually focus focuses, but on the input side, how are we actually helping our creators to mark their, um, input data or their training data sets so that we can have true providence from creator to derivative work that's being produced by some of these generative tech technologies as well. And so I think true Providence is gonna have to take into consideration monetization and attribution and licensing.
And so I think it's gonna either originate from something like the Content Authenticity initiative or it's going to be as a derivative product to something like this. And then another one is gonna be the partnership on ai. If you haven't checked them out, they're absolutely an interesting group.
They are really about advancing the positive outcomes for people and society when it comes to ai. And in particular, one of theirs is safety critical ai, meaning, you know, know how is AI being used in automobiles and um, medical devices. And so this is another fascinating one to either get involved with or keep a watch on as you're starting to think about what could happen and how do we actually start to think about how we're gonna trust ai.
And I'm gonna leave you with, um, a little bit about this algorithm of rave. How do we best trust with resilience? Um, metrics and adoption, velocity and errors.
You know, how do we measure success with AI is I think a crucial part of as a community, how we're gonna advance the conversation. And so, um, this article's out on Medium, definitely worth checking out. There's a lot of metrics listed there.
I'm always interested in having the conversation with folks. And again, you can find me here at third Score if you're interested in furthering the conversation. Absolutely dedicated to seeing AI advance and being more part of the conversation about what can you trust and how can you trust it.
And thank you today, and hopefully you get a lot out of the resources that we have. I'm looking forward to continuing the conversation.





