Evolution of AIOps – Laurent Gil, CAST AI
Laurent Gil, chief product officer for CAST AI, explains how artificial intelligence for IT operations (AIOps) is rapidly evolving at a time when IT environments have never been more complex.
Transcript
This is Textron tv. Hey guys, thanks for the throw. We're here with Laurent Gill, who's chief product Officer for Cast ai, and we're talking about artificial intelligence in the enterprise, what's real, what's not real, and what people should expect.
Laurent, welcome to the show. Thank you. It's nice to, to, uh, to be there.
As usual. You guys have been an early pioneer in the whole space of AI ops, and you recently picked up an additional $20 million in financing. But what's real these days when it comes to AI and the enterprise?
'cause a lot of folks are, at least early on, we're skeptical. Is, is is the mindset changing? It's interesting you said that, Mike.
Um, it's, it's, uh, it's actually crazy what's happening. We, we are living the hype and going into reality and we are talking about ai, like something is wrong. It used to be like a few years ago, we, we were talking, oh yeah, we use ai.
It's so boring and normal that every, nobody caress about it anymore. And so that's the end of the hype, and now there's no hype anymore because it's, it happened. Uh, so yeah, we, we, it was a, it's a, it's a great time and the weird time at the same, at the same time.
So I'll tell you a few things. Uh, one is, um, the, there, there's a lot of, uh, uh, discussion. We all see chat, D p T and others, um, mostly chat, d p t, um, with, uh, the core things we can do as an assistant, especially for as human.
'cause that's the free tool we can access to on chat g pt, and I love it personally. Uh, however, there is a real, there is, there are real use of AI in operations, and we are one example. There are many others like us.
I'll, I'll give you an example of that. So essentially what we do at Cast ai, and it's called cast ai, okay? We, we had this before.
So what is, what we do at cast ai is, um, we automate completely the management of cloud infrastructure. That's what we do. We, uh, customer connect their cloud account to us, whether it's a w s, Google or Azure.
Then automatically our AI engine has been deployed and it's starting to look at all the cloud native application that you may have within these accounts. Cloud native applications means anything that use Kubernetes, anything that use containers. And immediately, as soon as we find them, we start to optimize their infrastructure.
So the AI engine has been trained to do, to look for a few things. One is it looks for over provisioning. This by far, the thing that happens all the time over provisioning means your application need a hundred CPUs, but it's provision on 350.
And why is it so? Well, they are there, they, you're not using them. Uh, you're paying for them, but they're not being used actively by the application.
And so our engine was trained to identify that, but also to reduce the amount of C P U in your account so that you only have what you need, nothing below it, otherwise you'll have performance impact and nothing above it. The reason why we use AI or machine learning techniques is because this is a very, very, very difficult to do by hand. And I'll give you a hint of why.
And this is where AI is interesting, right? It's not replacing the human. It makes us better because the human cannot do it otherwise.
So I'll, I'll, I'll, like, I'll show you some example. Uh, imagine that you are, um, a SaaS business. So typically as a SaaS company, your customers are sleeping at night.
So at night you don't have a lot of traffic. Traffic goes up during the day and then goes down again during the night. It's never predictable, meaning sometimes it goes up fast, sometimes it's very chatty.
So it goes like this, but globally it goes, it's a, it's still looking at the shape like that. Um, you need an engine that is extremely, extremely fast at adding machines as your usage goes up or your requested usage goes up. When you add a machine, what machine do you add?
Remember on a w s you have about 600 different type of compute, right? VMs, they're call C five, M five. There's C five, a m, M five, M five A.
There's a more memory, less memory. Some GPUs, sometimes no G P U, which one is the most effective at this instant so that you add that machine because you have a certain requirement of compute. That's what this fuzzy, non, non-linear machine learning techniques are very helpful for, because they will take that decision in a few milliseconds.
And that's effectively what we do. And this is why when we add the most cost efficient machine very, very fast, and then we delete them when there is no need for them very, very fast, uh, the machine learning techniques are useful. They tell us, or they find out by themself which one is the most cost efficient within a very non-linear set of problems, such as, do I need to add an on-demand machine, which is very expensive?
Can I use a spot instance, which is very discounted type of machine on the cloud provider? Which one should I add? 'cause there are many different types, you see?
So you take all these viable and that's when, uh, using machine learning technique is useful. And we see that impact as a substantial cost reduction with our client. So it's substantial, it's automated, and the human cannot do it themself.
This is where the magic of AI is useful. We're not replacing anybody, we're just making the DevOps function smarter. We make them, we make them analyst because they still have to tell how they want this to behave, but they don't have to do the job someone else does it.
What is the relationship in your mind then between AIOps and finops? Because we see a lot of folks these days trying to reduce costs programmatically, but maybe the mechanism for doing that is ai. Yeah, these two worlds are starting to merge between them.
So, um, the traditionally the finops function is a financial function where they, and they are very, very concerned with cost and budget. So they, they use very cool cost reporting techniques that allow you to allocate and manage budgets of teams and sometimes even budgets of application. So you can go very granular with a finops product.
AIOps is in the, uh, the job of AIOps is to automate smartly the operations. But if you think of AIOps and DevOps, it means your infrastructure becomes automated au autonomous. When you combine these two things, it just means that you build an AI engine that is very cost aware or even cost obsessed, and it will make decision based on cost optimization, infrastructure optimization, and it will feed the information to great reporting that the finops people will use.
So these two walls are colliding and merging into each other. Some finops, they will stay in the reporting structure because it's easy and simple and small finance function. Some device will keep the automation because that's what they, uh, they are there for.
They wanna make sure that everything works 24 by seven. The, the combination of these two things means that you use very effective machine learning techniques to automate your infrastructure and then to reduce your cost. Because as you automate, you may as well automate for costs optimization.
That's where these two all are combining to each other. Have we reached a point of complexity with it where a lot of IT professionals may not wanna work in an environment that doesn't give them access to AI to augment their capabilities. 'cause otherwise it's just too hard and too many mistakes will get made and everybody gets yelled at.
Yeah, uh, uh, um, the AI makes mis mistakes too. So, uh, it is not a one way or one street it they make, uh, it really does. Um, and maybe the, the person who trained the engine make the mistake and therefore the, uh, AI engine makes mistake by design.
You, we see a lot of things like this. Um, but it really, the, the debate I think is, uh, the debate, oh, is AI going to replace the human? It's a wrong debate in my mind.
It's more that the machine learning, the AI engine is doing things that the human cannot do anyway. And there, and when that happened, this is where we been all benefit from it, we become a better place. Like imagine you are, uh, uh, your founder of a SaaS business and certainly you use ai, it reduce your cost by 70%, so your cost of goods sold is dropping by 70%.
Then this founder, uh, or this C T O that runs this now should only have new budget to do more magical things, right? And that, and this is where we all benefit from it. So I think that, that that revolution publish unstoppable, and it's a good revolution because it makes us, it makes the pla the world a better place, you know, one SaaS application at a time.
We have seen the rise of machine learning algorithms, um, but we're also hearing a lot of people talk about generative ai. Now, how will these things come together and maybe be applied for AI ops specifically? Oh, it's easy.
We use the same techniques. It is a funny, funny, I didn't know that. Uh, we, we start to look at what, uh, what is, what are the, uh, uh, the techniques that chatt PT was using?
We find out, we actually were using the same thing, uh, uh, um, this, uh, this, uh, AI engine have a lot of trees and a lot of branches that goes in different direction, but the core of it was very similar. Uh, so yeah, it's, it's all about, it's, it, I don't, I'm I'm not saying it's the same engine, that's not true. Uh, it's not the same engine, but the principles and the way it's designed and the way it's been trained is very similar.
How do the algorithms keep up with the changes in the environment, especially in places where some folks are changing the code based daily? Yeah. How does the algorithm know or learn over time?
Our, our model change daily too. So one of the, one of the techniques we have is what we call a spot prediction. So we are, we're predicting with a few minutes in advance the evolution of the spot market, which is a kind of a low, uh, low cost machines on a W Ss and on on Google and Azure.
They, they have the same, um, and, uh, so we, we, we built the engine that it kind of tried to find out the prediction of this with a few minutes in advance. It's super cool, uh, because it means, you know, when the spot, uh, instance is going to disappear and you know that before it happens. Um, but that is so volatile to the situation of the moment that the engine is actually being feeded the real result over time.
So every day it becomes, it's not that it becomes better, it changes based on the situation, and it really does that, right? It's a full feedback loop based on what is predicted over what happened. And the difference, the delta between this is constantly fitting the engine, not so that the engine becomes better over time, but because the, uh, market is changing over time.
So it's almost like you are reading a newspaper article every day about news, and the news is never the same, but you have to read the article. So the articles are never the same, but you are, but the fact of reading them is similar. It it's the same for, um, this, so the, this evolution of, uh, the model, the evolution of the model is natural based on the environment.
What do you think people are still struggling with when it comes to AIOps though? Because we do hear a lot of instances where people are like, well, we deployed it, but we didn't get the full value out of it. Or we only got partially, or we couldn't make it work at all.
People are all over the spectrum. So, you know, in your experience, what have you seen customers who do use it do well, and what are those who are not maybe should be thinking about? Yeah, the best answer to this is trait.
And you'll see, um, and we, we say this to all our users, you know, when we talk to someone and we say, oh, we're gonna reduce your cost by substantially, right? By 50% maybe, or even even higher than that. Uh, yeah.
So you tell, how do you do that? Like, how is it possible? I know my thing, like it can can happen.
And so we told 'em, look, just, uh, click on a few button and wait for a minute and you will see what the engine is telling you. And very often the engines is predicting, uh, some really cool cost savings that you have. Then you have to activate it.
So there is a trust relationship that has to happen. But the, the first step is always, this is what will happen if you use it. And that's instant, like that takes a minute or two to get, we design it this way so that users can see by themselves what would happen if the engine was activated.
And then there's another button that says activate. And even when we design the activate function, um, it does nothing, meaning everything is turned off by default, even af after you activate. And then you as an analyst have to start activating features and functions of the engine, such as use spot instance with prediction.
And it's very step by step. And after, usually after a few days, they start to do that thing, they click on a few functions and features, and then they see the impact. It's very graphical impact usually is like this, right?
The costs start to go down and the curve of utilization, the curve of requested, uh, compute is, is getting close to each other and that's when the trust is built. And I, I tell you Mike, we don't have any users that don't use us in production and we don't have like, it, I don't think we have a case of, uh, uh, an organization using us for things that are not for production. It's already start by staging, by development because they wanna play with it.
Um, but then they immediately go to, to the next stage. So there's, there's this stress function that we, we, we show over time. How long does it take for the AIOps platform to effectively learn?
I think a lot of people think magic happens out of the box, but I think, you know, it takes a little while for the environment to become fully, um, described to a machine and then they can make smarter recommendations. Not so much like, uh, again, go look at chat g pt. If you ask a question to chat g pt chat, g PT doesn't know you, doesn't know who you are, but they will it, the engine will still answer to you something relevant to the question.
Hopefully it's the same thing for AIOps. AIOps is not so much linked to what is the application doing in our case, but it's only to what does this application request? And when you know what it requests the same way as you ask a question, the application ask a question.
The app, the question the application ask is, I need all this and the engine say, great, I will supply it to you. And the application say, well, what I need now is a bit different. The agent say, great, I will change it to supply it to you.
Every few milliseconds is like this. So the, the training is not so much application by application. The training is more about, I need something, the engine will try to find the most cost, cost effective way to supply it because it's being trained to supply something as soon as something is necessary.
So that's, that's how it works. What ultimately will be the impact AIOps will have on DevOps workflows? What's the, how do those two things come together in your mind?
Oh, the DevOps becomes, becomes much more interesting the same way as you ask a great question to G P T and you have a great answer. The same thing for us. So the, the, when as the DevOps becomes more interesting, it means first they have, they have a drop in the, in their cost so they can enjoy that to do something else.
Second is they become analyst. It really is like this, it's become analyst means, means the following. They start to, to craft the, the this, they start to influence how they want the model to behave in, in, in our world.
These are simple things, but uh, simple things can be, oh yeah, I've seen how spot instances behaving. It works really nicely. I'm going to increase the usage.
This is what the analyst is going to start to decide. Another, another one would be, oh, I see that my application, uh, at some point need more containers. I'm going.
And now that I know I can a supply enough compute whenever I need to, I'm going to use, I'm going to really use these auto scaling features that Kubernetes is giving us. It's a little bit technical, but these are the, the type of discussion usually DevOps engineers have with us. There's another one that says, oh, I thought this container need five CPUs to run, but the AI engine is telling me it never used five.
1. 1? 1 is a lot more stable than the five that we thought before.
So you, you are giving back the information to the human so that human can adjust and craft it's, uh, the way it want the engine to behave in a much better way. All right, folks, you heard it here. DevOps is gonna get more interesting because there's gonna be a lot less toiled and we're gonna have more fun.
That's a promise. Lauren, thanks for being on the show. Thank you, Mike.
All right. And back to you guys in the studio.