AI and ML in Cybersecurity with Dilip Bachwani | Qualys QSC23
Dilip Bashwani, CTO for the Qualys Cloud Platform, delves into the intricacies of artificial intelligence (AI) and machine learning (ML) in the context of cybersecurity. The discussion highlights the critical role of data quality, diversity and volume in training reliable AI models.
Transcript
This is Textron tv. Hello, and welcome back to the Quas Security Conference in the Americas. We're here with Dilip Bai, who's CTO for the Qualis Cloud Platform.
Mm-Hmm. And we're talking about artificial intelligence. How you doing?
Good, how are you? All right. Welcome.
Um, at the risk of an oversimplification, in its simplest form, we have all these signals and data that we're collecting and putting into something that is a model, and then we're getting an output. How do we know that the output is trusted? What goes into that process?
Walk us through that a little bit. Okay. Uh, so, well, that's a big question.
Um, but, um, see, ultimately, when you are thinking of machine learning and ai, your models are trained on the data that you give it. And the outcome of that really is dependent on the quality of the data, the diversity of the data, and the volume of data that you have. So, if you have all of that and you are building the right kind of models, then you will have models that have a higher degree of Accuracy, right?
Because machine learning ultimately is predictive analytics. It's looking at the past, it's looking at trends and based on the data that we are feeding it. And then it's saying, here is what we can infer from it, and here is what we can predict, right?
Um, so the way we look at this is, today, across our platform, we have a lot of data. We have data across vulnerabilities. We have data across misconfigurations, cloud, SaaS, ot, iot identity.
And now with our enterprise tourist platform, we are also pulling in data from third parties. Because the vision here is to collect the data, to bring it together, and then to look at all of that combined and then to be able to say, what is your risk? Right?
So we are taking all this data. We also did an acquisition last year for Blue Hexagon, and we have a very talented team of MLN AI engineers that came on board along with some, you know, I'd say exceptional deep learning, AI model, building expertise. And we've invested significantly in our overall AI and ML efforts.
So we are taking this data, we have a very talented data science, uh, ai, MLT, we are building these models. We are training these models, and then we are saying, we are validating that. So to really, to answer your question, it's not just building the models.
You get it out, you look at the results, and then you are continually iterating and training these models, these models, until you get to a level of accuracy, and then you'll able to get that out in production. So it also seems that there's two fundamentally different kinds of ai, maybe more, but there's predictive, and then there's generative. Generative is the stuff that everybody seems to be excited about, but these two things need to work in conjunction with each other.
So how does that happen? Hey, Yeah. Yeah.
It's, uh, so generative AI is the new buzzword. Uh, I, I think we all know that as soon as OpenAI launched chat, GPT, it in many ways made AI more accessible to everybody. It was something that you can touch and feel and relate to.
And generative AI is built on, you know, what we, what we call foundation models. These are like models with billions of parameters, And they can solve problems that are not, that are more generic. So you can take the same model and you can use that model to answer all a diverse set of questions, uh, or have its applicability to a diverse set of domains.
Now, of course, you have to train that model more for your particular domain, but foundation models are something that will change how the industry works, generative models, right? Uh, earlier the way machine learning work was you would take a particular problem domain and you would say, well, let me build a model for this problem domain. And you may have a few hundred parameters versus the millions and billions that we are talking about.
I think when you also think about generative ai, I mean, as I said, it's certainly going through its hype cycle. Uh, I think the hype will settle down. We will find real use cases for it.
Uh, those are being identified for sure. Mm-Hmm. Today the use cases largely are how do you converse with it?
Uh, how do you chat with it? Or how do you take query language and convert that to natural language so users can interact with it, with your systems more easily? But I, I don't think that this is so much about convergence.
I think ultimately this is about what value are we providing, right? So, however you do machine learning and ai, ultimately it's the value that you're providing. And whether you are doing building purpose-built models, or whether you are using, you know, something which is more based on the foundation models and the l LMS on top of that, um, it's, it'll converge in terms of value ultimately, right?
So we get these outputs. Yeah. How do I operationalize that?
I mean, 'cause it seems like it's nice to have the models and it's nice to have a result, but I gotta do something with it. Yeah. Yeah.
So, you know, with a generative AI again, right? Um, I think when charge GPD came out, The only use case that people initially thought about was, here is a nice way of interacting with systems. And pretty much everyone jumped on it, and we saw them, right?
Most organizations very rapidly announced capabilities that were built on top of these large language models. Uh, that's a good use case, but there is a lot more that we can do with it, even in the cybersecurity domain. I think being able to use these large language models, and even in our use case yesterday we announced, uh, that we have a policy query language, which our customers use to query, uh, you know, their data to build really high quality dashboards.
But instead of users having to use a domain specific query language the way we do, you can ask questions using natural language, or you could ask questions using voice, right? And voice to text translation is a solved problem, right? And then you can get results more easy.
Now, I think, I really think what will happen is, especially with large language models and generative ai, the way, at least in some narrow scope that we are looking at right now, it will change how users interact with UI and UX right? Across industries. Um, and that will be a good use case.
But there is more. Uh, so when you look at threat hunting, when you look at, uh, doing extensive analysis of patterns, because in cybersecurity, there's also very large volumes of data right? Now, how do you analyze that data to predict what's coming out?
How do you analyze that data to say, here are things that the models can infer, but human analysts when they are looking at data, might not think about it along those lines. Machines think independently, right? They're trained differently.
So there will, there will be other insights that will come out of it that we will tap into. Uh, I think there is use usage in incident response. Uh, there is usage in, you know, threat hunting.
Uh, there are good use cases for generative ai. I think most people, when they play with this stuff, there's an initial wave of excitement. They're like, wow, this is gonna be great.
Awesome. And then there's this nagging little thing about like, is my cheese being moved here? Yeah.
So how will the job functions evolve, and how do you think this is gonna play out? So it's a good question, right? Uh, the, there is some sentiment out there that generative ai, especially, uh, and this wave that we are seeing with AI over the past 12 months will take jobs or will impact people.
Um, but I don't think that's the case. I I actually think this will make us more efficient. Uh, this will make us more productive in many ways, right?
So this is, these, these solutions are not something that will be your shopper. They're not gonna be the, they're not gonna drive the car for you, but it will be more like your copilot. It's with you.
It's enabling you to drive the car, but you are needed. And the reason I say that is there are certainly some areas where, you know, for more commodity use cases, you could replace them with automation, which is happening across many industries. But in, in, in a large majority of cases, there is a lot of business domain and internal domain knowledge that's needed to actually build the solutions and to provide the value to your customers.
And the generic model is not going to solve them. Now, of course, these models can be trained further to the narrow scope of your organization, but to do that, you need people, right? So I think it'll augment, it will compliment what we do.
I don't think it's replacing anything. I think there's, uh, maybe 4 million open jobs in the cybersecurity space. So is that gonna get shrunk down to 2 million?
I mean, not that we have 2 million, but I, given the shortage that we are seeing mm-Hmm. Uh, in cybersecurity, if we went from four to two, I'd argue we still have a shortage. True.
So, and this is what I'm saying, this will compliment, this will probably enable us to do better. Because when you, when you look at cybersecurity, the average number of tools that an organization has, cybersecurity tools, is anywhere from 15 to a hundred, averaging at about 60. And when you look at these tools and each of these tools generating their own set of data and security is a very complex, very, um, it's, it's a very busy marketplace.
There are too many vendors out there, and some of them, they do one thing, they do one thing really well. So there are organizations that have said, I'm going to go for this particular problem. I'm gonna go with this vendor for this other problem.
I'm gonna go with this other vendor. But ultimately, you have all these different pains of loss, right? And how do you unify them?
And the signals that are being generated are two men. So where machine learning, ai, generative ai, where this will help is you can do this analysis of this very large volume of data. You can look at past exploit trends, previous exploits, patterns, and then you can say, what can I infer out of that?
And what is a narrower set that I want to focus on? You could even take it further to say, what is a narrower set I want to focus on for my organization? Right?
Which is what we talk about in terms of enterprise true risk. As we are pivoting as an organization to be a risk driven organization, we will tell you what your risk is, right? And I think when you, when you really think about it, you know, these, these will help you there, right?
And then you benefit from that. And the folks that you have in the security industry, we don't have a lot of them, right? Or as well-trained, we can take that, the data that we have, and work off of that data to be more productive, to hopefully stay a step ahead from the adversaries.
Is this gonna accelerate a shift to the way we consume security as a service? Because a lot of people are still buying tools and installing them themselves, but it doesn't seem to me they'll ever be able to collect enough signals to drive an AI model. So will AI push more people to consume security more as a cloud type service?
I, I think what will happen is security vendors that are more platform driven versus single solution vendors will do well. Mm-Hmm. Um, there are use cases where you may have different security vendors, but I think the gap today is when you, when you talk to security leaders, when you talk to CISOs, they are asking, and the board is asking, what is your security risk?
What is your enterprise risk? What is your business risk? And it's hard to articulate that when you have 15, 20, 50, a hundred products, how do you know what your risk is?
I mean, your signals are all over the place. And, and the thought behind having a a security platform is let's bring all the signals into one place, whether you are the generator of those signals or whether you're integrating with others to bring those signals in. And then let's do analysis on that and then come up with a risk posture for you.
Mm-Hmm. I think that's where the industry will get to, and that's where the value will eventually come out of now to, to have a platform, it requires significant investment. And, uh, and the organizations that will do that, quass obviously is investing significantly into that.
We, we have an extensive security data lake platform. We've invested heavily in it for the past few years. And that's really why we are today at the point where we are able to say we have an enterprise tools platform where we can ingest data and signals not just from Qualys products, but from other products.
We bring that together. We combine all the data, we do our analysis on the data, and then we tell you what your true risk is, what should you focus on in your environment? I think that's where the convergence will come.
There are these people out there, there're generally considered bad guys. Yeah. And, you know, they have access to similar things.
Yes. So is AI gonna help them more or help the defenders more? You know, I, I, I think, think it's a, it's a double-edged sword for sure.
Mm-Hmm. Uh, you know, by, by some metrics, about 90 to 95% of breaches start with some form of social engineering. And, uh, I think we are all familiar with the Nigerian prints, uh, asking for help, uh, in exchange for a lot of dollars, is there A fax machine Or anything?
So, Uh, but I think especially with something like chat, GPT and generative AI and large language models, I think the Nigerian prince is certainly still around, but the persona has changed. The emails are no longer, you know, having typos and grammatical errors where you're thinking, well, this doesn't look right to me. Right?
Even if it finds its way in your inbox, what you are getting now is something that feels very authentic. And for many organizations, if you were to do a phishing test, you will find folks that click on these emails. So from that standpoint, it helps because social engineering attacks are already increasing.
We are seeing them. Yeah. Uh, another aspect is a lot of the adversaries, the bad guys, uh, there are nation state actors who are very talented.
Uh, and then there are others that are motivated more from a financial standpoint. Uh, they may not be software engineers, but now they have a tool that can generate code for them. So it gives them this sammo where what is generated is not perfect, but maybe good enough for them to explore and play with, uh, as a starting point.
So I think there is certainly areas where, you know, we are seeing that influence. But, um, but from a benefit standpoint, I mean, technology can be used in many ways. Uh, I think from a cyber security standpoint and across many industries, there is a lot that we can do on our side.
Uh, as an example, just here at Qualys, uh, we are deeply integrating ml ai, generative AI into our products and solutions. Um, the key for us to stay one step ahead, um, right, make sure that our systems, our products are secure, uh, but then also individuals, folks are trained on, you know, these social engineering, uh, kind of attacks, uh, which tend to be oftentimes the gateway into breaches. So, So have we essentially stumbled into an AI arms race?
Is is that where we are? You know, it's funny. Um, so our CEO, Samantha Kar, he often, uh, talks about this, that, um, we are getting to that point where if adversaries are using AI to get to you, the only way to counter that is ai, right?
Uh, so I, I'm not sure I would qualify that, you know, at this time as an arms race, but, uh, it definitely is a major tool in our, All right, folks, while you heard it here, love it or not, you're gonna need ai, so you might as well start figuring out how to make the use of it before the bad guys do the same. Hey, thanks for coming by. Thank you so much.
All right. Take care.





