The Impact of AI on Infrastructure Management with TuxCare’s Joao Correia
Joao Correia, tech evangelist for TuxCare, discusses the transformative impact of AI on infrastructure management, highlighting challenges in integrating and securing complex AI systems without traditional downtime methods.
Transcript
This Is Textron tv. Hey guys, thanks for the throw. We're here with Johan Correa, who is tech evangelist for Tux Care, and we're talking about the rise of a new class of infrastructure, arguably for running AI workloads.
And how we manage that though may not be compatible with the way we currently manage infrastructure. Well, we'll jump into that in a minute. Yo, welcome to the show.
Thank you very much. Pleasure being here with you. What is the challenge we're gonna see because we're all trying to figure out how to operationalize AI and not all of it's gonna run in the cloud.
So, um, what are we not thinking through enough about the way we need to acquire, manage, and ultimately update all this AI infrastructure we're looking on getting? Well, if you take a look at the stuff that was presented, say at Compute X, uh, a few weeks back, um, you'll see solutions that take up in racks. Our sold is entire racks and they cost millions of dollars each.
So when you purchase something like that, that's so expensive that you then need to maintain and you need to deploy patches to and perform your traditional maintenance operations on every minute. That that thing is not wrenching numbers, it's wasted money that you spent and cannot recoup easily. So the way that we do our traditional operations on the data center, the way that we patch our systems, for example, the way that we've always done it unfortunately, is that okay, we schedule a maintenance window for some amount of time in the future.
We get approval for all of that, and then we take down our system, we apply our patches, we take down our systems for reboots, and they come back up. That entire process takes an inordinate amount of time. When you're talking about this type of equipments that are powering AI right now, these things have dedicated cooling, have specific power needs, have dedicated networking all built into the those racks.
So you can just take them down piecemeal and perform your operations and taking them all down at once becomes really expensive really quickly. So what's your best advice then for how to go and incorporate this stuff into our operations, which I guess we're gonna have to re-engineer in some way, but how? Okay, so traditional operations haven't been able to cope with the pace of cybersecurity threats for quite some time now, I'm talking years that it's not possible to respond to new threats in say, I'm gonna patch months from now for a critical vulnerability that has been disclosed today.
Those four weeks are gonna be really stressful for disease admin in charge of those systems and performing those operations because they know that at any moment those systems can be hacked. So if it's not possible to respond in that timeframe, we need to shorten the timeframe. One way to achieve this, and again, going back to these AI systems, all of them will be running some form of Linux because that's just what's powering all of these systems.
We have a solution for this and we have had the solution for this since 2006 when it was invented and brought into Linux, which is live patching. You can deploy live patch to your Linux systems. You can use live patch solutions to avoid all of the reboots and all of the downtime and all of that and be protected within minutes of the patches being out rather than weeks in the future that will shorten their, their response time.
Even for traditional servers. And this is even more important for systems that are as expensive as this ones because you will not incur any downtime. You will not stop your training sessions.
That can take weeks at the time. You can, you are not reducing the, the ability to perform inference on the models if you already have them trained running on the systems and you will be secure faster. And that's really important.
You're shortening your exposure time and you're not wasting money on systems that are being rebooted. To your point, we've had live patching for a while, but it seems to have been adopted unevenly. What is it that kind of holds people up from embracing live patching?
So this might be hot tech, but the IT industry is an incredibly conservative space. I don't mean this in the political aspect of it, I mean that it's really difficult to change the way that people are doing things in it. We love the best and greatest new technology.
We're really in love with everything that China knew, but then we shoehorn it into the way that we've always done our processes. You everybody doing it today. Every CIS admin out there has been brought into the field and has started doing this type of work with the notion that yeah, if you need to update your systems, you're gonna need to have to reboot them.
And that's what we're teaching people coming into the field and that's how they are learning and we're not really telling them why, it's just that we're doing it because we've always done it like that before and we're not changing that unfortunately, even though we have better and more effective solutions available to us, that's a problem in itself and that should be addressed. But it becomes more apparent when that downtime that is incurred in these operations translates directly into lost money and lost revenue. The live patching, is that gonna become the default or is that just buying me time before I have to at some point take the system down and kinda recalibrate everything?
Uh, um, you know, how much can I depend on live patching? That depends a lot on the live patching providers that you go with. Um, there are two classes of live patching providers out there.
There's the, the distribution, the Linux distribution vendors that will offer live patching solutions for their own distros and Los Oracle and Red Hat come out with solutions that say, yeah, we're gonna give you a stop gap solution and you should reboot in a month or two months or three months or something like that. But you should really reboot some time in the future. And then there's other alternatives.
Um, again, I'm going to be talking about tech care directly. I work for tech care, but our solution does not require that we have customers that have systems running for years that have perfectly secure kernels running all the latest and greatest security updates that have been released up to right now and the systems haven't been brought down in years. Years ago, if you saw a system with an uptime of two or three years, yeah, that's not a secure system in any shape we form that system is just asking to be, to be breached.
But it is possible to have systems that are receiving live patches that do not need to go down for many months and for many years. It's more likely if you're going with the live patching solution that you need to reboot and you need to take down a system to perform some type of hardware level replacement, replace a memory stick or add a new card into the system than it is to actually as a result of the life patching operation, there is no necessity to to perform a reboot. Are we just kinda scared of patching?
I mean we, to your point, it's a cultural issue. We have conditioned everybody that says, well, we can't apply a patch without a developer. So everybody just kind of says, all right, we'll wait for the developer to tell us that they're ready to patch something, but maybe we should be moving to a point where it operations people can patch things and kinda reduce our cybersecurity risks.
It obviously we is, we are conditioned to do it to get that way because when we started doing this 25 years ago, 30 years ago, yeah, many times you apply the patch and you broke something and that was a dangerous proposition. So I'm deploying a patch and trying to be more secure to get the latest and greatest version of the software that I'm running and I deployed that and I broke my application or it doesn't boot anymore or something like that. So we learn to fear that mm-hmm.
Microsoft might be to blame, therefore most of it. But still we learn to blame to, to fear the patching process. We learn to fear applying patches to our systems, but unfortunately we haven't learned that in the the time since things have improved dramatically, it's no longer as risky.
You have a proposition. Uh, plus with life patching, you have the ability to both deploy patches and remove patches. If you find it that uh, there's a performance hit that you're taking for a mitigation somewhere, then you don't want that mitigation 'cause you're addressing it somewhere else, say at the firewall level or something.
So you apply the patch, you're not happy with the result, you can remove the patch again and no problem. Wait, do you think that, I mean it seems to me at least that cybersecurity didn't provide enough of an incentive to embrace live patching. So do you think it's the economics of AI downtime that will finally push people in this direction?
I mean, is is economics a bigger issue than, uh, cybersecurity integrity? One place where we've seen life patching being a necessity is, for example, when you went to scale up your, your infrastructure past a certain number of servers, you still need to deploy updates to them. But it doesn't matter how many, how much staff you throw at the problem, how many people are handling the, the reboot processes and orchestrating all of that past some point, it becomes just impossible to do within a reasonable timeframe.
So many of our very large customers told us and recent customer in feedback interviews that we do that. Yeah, this has enabled us to go past say 5,000 servers, 10,000 servers, 15,000 servers because it would be unmanageable to do this traditionally. And life patching allowed us to grow past that with ai.
The angle might be something different. As we've seen recently, I believe it was JP Morgan that a couple of weeks ago, if not last week at all, came out with a statement saying that yeah, these investments in ai, they're so expensive then. And it's becoming really hard to see how we're gonna recoup the investment in any industry.
If you're adding to that, the downtime that those systems will suffer just from this alone, just from keeping them up to date and secure, it becomes even harder to justify the the initial investment. And when we're talking about this amounts, this is strategic investment that companies make. This is not just something that you do because say your customer base grew and you need to satisfy more customers at the same time or you need more performance for something, this is something that will enable you to create completely new solutions and completely new products on top of it.
So couple that with the cost. I believe that the only way to reasonably maintain and expect the systems to give you some type of return on investment in a reasonable timeframe is to keep them always running and to keep them always running and secure. Life patching is probably the best alternative.
Well, do you think maybe we might do better with smaller AI models that don't require as much infrastructure, but there may be a little more domain specific and uh, perhaps over time easier to manage than these large language models that people are running around that only a few organizations really can afford to kind of drive and that will ultimately be the thing we see in your average enterprise. That might be one way to approach it. Yes, you might have some more domain specific, more focused, uh, large language, large language models trained on specific domains that will give you specific answers on those things.
But there has been research on the AI field that tells you that the systems are still able to scale further and be further improved simply by adding more compute power to it by adding more stuff that they're trained on. So there are arguments against and in favor of both larger and smaller models. Um, it really depends on what you are trying to get out of them.
If you're satisfied with just a limited model that will only focus on a specific aspect and you can manage to train it just for that and it works fine for your use case, then by all means that's your solution. If you're trying to go for the, the generic artificial intelligence that some companies are trying to achieve out there, who knows how long in the future, um, yeah, then you went to throw lots of compute at it, you went to throw lots of resources at it and training and all of that, and that's going to require the expensive equipment and that's going to require all this that I've just mentioned before. You want to systems running a hundred percent of the time or else you're wasting resources.
Do you think that the economics of AI will force people to think through more of what types of models are running and what types of hardware because well, GPUs are hard to find and they're expensive. So will we see more inference engines running on other classes of processors and you know, suddenly, you know what hardware's running where will matter a lot more than it ever did. The economics will surely push that.
Scarcity will force us to look for alternatives because as you say, GPUs are getting scars, especially on the the high end things. So if you can't find those and you really want ai, you're going to have to look for an alternative elsewhere. So we, you're essentially being forced into an alternative.
Um, and this is something that Nvidia is pushing many players into because if they can't afford or they can't find NVIDIA gear, which right now is considered top of the the line for this, then they will have to look elsewhere for the, for that equipment. Are you seeing more IT teams taking over the management of this, this aspect of ai? And I asked the question because you know, historically the data science teams kind of tried to do it all.
They were building the models and then they were deploying them and trying to figure out what hardware to run where. And I guess is, is that the management of the actual deployment of the models? Is that moving back into the realm of the mainstream IT organization?
While the data science teams might be able to select the hardware and influence the decisions, the ones deploying the racks and maintaining the servers before it reaches the application stack where they are running the AI stuff will inevitably be the same people running the the servers that you have right now. Um, I don't think that will change first because the data, the science team don't want to get into the dirty and nasty stuff of server management and because the, the CIS admins very likely want to avoid touching all the math involved in the, the eye training and the inferencing and all of that because that's not really what they like to do. So yeah, and, and here I'm talking as a previous, his admin who I really went with who I really love when people came to me with, okay, I need the system to do this and I provided that system.
Okay, feel free to do whatever you want with it. I'll just deal with the, the system stuff. The application stack is yours.
What is your best advice to those teams then? Because there is a certain amount of irrational exuberance. Everybody gets excited about AI without thinking through all the implications.
So if you're running the infrastructure, what are you telling those folks today? Well you're gonna have a lot of very expensive ties to play with for a while, um, until everybody's expectations drop off a bit and everybody comes down a bit with AI and the hype passes, you're gonna have some very expensive equipment that you're needing to to maintain. But I really think that things will stabilize at some point and first those systems will become cheaper and more available.
And secondly, I don't know for how long we're gonna see this separation between AI dedicated systems that are just jam packed with GPUs and traditional systems that have more compute and more RAM and other type of peripherals and connectivity. I don't know for how long we're gonna see the, the two classes of systems being so separate. I believe there will be some type of convergence in maybe not the near future but midterm four or five years down the line.
Alright folks, you heard it here. There's something of an IT infrastructure renaissance coming and it's driven by ai. The only issue is like the way we managed infrastructure before may no longer be as applicable as it once was.
Hey buddy, thanks for being on the show. A pleasure. Anytime.
Thank you very much. And back to you guys in the studio.