Stop Wasting Your GPUs: The Secret to OpenAI-Level AI Hardware Utilization | Utilizing AI Episode 27
You might own the latest AI hardware, but are you actually using it? In this episode of Utilizing AI, host Stephen Foskett joins Tech Field Day delegates Gina Rosenthal and Frederic Van Haren to tackle the “utilization gap” plaguing modern enterprises. While hyperscalers like OpenAI achieve near-perfect efficiency with their GPU clusters, the reality for most businesses is far more wasteful. They explore the practical challenges of hardware systems design, why high-performance chips often sit idle, and how organizations can move past the hype of acquisition to achieve true, practical utilization of their AI investments.This and more on Utilizing AI, part of The Futurum Group Podcast Network.
Transcript
The word utilization refers to the practical use of a technology, and this has long been the goal of systems designers. This episode of "Utilizing AI," featuring Tech Field Day delegates Gina Rosenthal and Frederik Van Haren, focuses on the question of utilization of hardware resources. Welcome to "Utilizing AI," the podcast focused on practical applications of artificial intelligence from the Futurum Group.
Every Wednesday, we explore news and use cases of the ways in which AI is transforming enterprise IT and the industries it serves. I'm your host, Steven Foskett, President of the Tech Field Day business unit here at the Futurum Group. And before we dive into this discussion, let's meet who's on the panel today.
Well, it's great to be here. I'm Frederik Van Haren. I'm the founder and CTO of Hyfence, and we provide HPC and AI consulting services.
Hey there, I'm Gina Rosenthal, and I'm the founder of Digital Sunshine Solutions. And we supply fractional product marketing services to business-to-business high-tech companies. So Frederik and Gina are frequent faces at Tech Field Day events, including the AI and AI Infrastructure Field Day events.
AI Field Day was last week as you're listening to this, and AI Infrastructure Field Day, I think is in two weeks when you're listening to this. Frederik and Gina will both be at AI Infrastructure Field Day, along with Alistair Cook, who's running that one. I am not going to be there.
But these are topics of great importance that have come up in a lot of the events. For example, the three of us were together at the recent Qlik Connect event. As I said, we've been together at AI Field Day, AI Infrastructure Field Day.
And again and again, the conversation comes almost back to the title of this very podcast. So we have spoken at length about how to make productive use of AI technology, how to utilize AI. But that being said, there's another use of this word, utilization, that I think is quite relevant, and this is something that has come up again and again with us.
So, Frederik, I'm going to kick it to you because I think that you're the one who brought it up, maybe at Qlik Connect. Utilization of resources, the percentage of time that various resources in a system are used. This is something that's been of interest to sysadmins for a long time, to FinOps, to CFOs, and CIOs.
What does utilization look like when it comes to AI hardware? Well, to be honest, it's pretty ugly nowadays. There was an article that was talking about xAI, which is Elon Musk's AI initiative, and they were saying that the utilization was at 11%, which is not that great.
You wouldn't expect it to be 100%, but at least a little bit closer. I think one of the challenges is really about the complexity of AI. You need networking, you need storage, you need compute.
All of this has to work together like a symphony, and in most cases, it isn't, right? And so some of those challenges are due to the fact that the market is growing really, really fast, and that sometimes it's more important to deliver something than to do it efficiently. But it's an important point, right?
If you look at the supply chain today, it's because those same organizations with low efficiency are asking for more and more hardware. So it's almost like an infinite loop that we have to address. Yeah.
I think it's going to be interesting going forward, too, and I love how it all comes back to infrastructure again, and how you use it. And I love how you talk about all the time that you can't just throw hardware at the problem. It's actually a systems problem, and you have to tune the system to work properly.
" We've been talking about the xAI-- well, the Musk and the OpenAI lawsuit, and lots of questions around, okay, one of the fundamental questions of that lawsuit is you were supposed to be a charity and a nonprofit doing AI work for the good of everybody. All of a sudden now, you've got this spinoff and you're spending lots of money on really good people, and all the rest of it. And of course, Musk's numbers came out at that 11% utilization, and OpenAI has a 97% utilization.
So such a big disparity, and are they telling the truth that they needed that money to go buy the right people to tune their systems? Are they getting places faster? Are they building things faster than xAI is?
Is there a problem with that if they're doing it the right way? But that seems to be a key point now of what's going on. It adds to that lawsuit.
So it's not just about the systems running really well, it's about getting to those answers faster. It's having the people that know how to do it, and I don't know, Frederik, is it that hard to do it? Yeah, it is.
Simply because the market is moving so quickly that you really come up with a solution, and then the market changes, or your workflow changes. And so you need to have a methodology in order to continuously improve your efficiency. I think the other problem is also due to the fact that traditionally, people have a tendency to throw more hardware by default at a problem when it occurs, as opposed to looking closely at what the real problem is and fix it from that perspective.
Yeah, that's a very traditional IT situation, and the three of us have been in enterprise tech for long enough to know, well, first off, that this is not an AI problem necessarily. This is a systems problem, as you said, Gina, that this is pretty common that we've seen for many years. Frederic as well, there's always been a temptation to just throw hardware at the problem.
And of course, vendors have always leaned into that too. I remember my days back at banks and energy companies and so on, and when systems were running slow or looking taxed, the easy answer was always get a bigger system, get a bigger box, get more storage, get more memory, get more whatever. The harder answer was try to figure out why the system is taxed, where the bottleneck is, and address that bottleneck.
Because, as we've said for many years, system tuning is all about moving bottlenecks. When it comes to AI, though, there is a question, and this is also a question when it comes to other kinds of system tuning, which is, what if there isn't a bottleneck? What if there just isn't demand in various locations?
I think that it's no surprise to hear that a company like OpenAI that lives and dies by AI inferencing on hardware would have incredibly high utilization of their hardware resources. I would imagine that they could probably tune it somewhat, but the point is, they're very focused on doing exactly that, making the system run as hard as possible. Whereas, a lot of the enterprise companies that you both may work with may find themselves really not in the position of being able to tune it, being able to resolve bottlenecks, but also early enough now that we may not even have that kind of demand internally on our systems.
What's your reaction to that, Gina? Because, yeah, I think you see that definitely OpenAI is in the business of providing that service, so they're going to make sure it's tuned properly. But you look at someplace like xAI, they're using xAI to run all sorts of stuff on the Musk and the Musk family of products.
So why has it got such low utilization? And specifically, I think these are all talking about the GPUs being utilized, right? Completely going all the time.
So I think that what we haven't seen-- We hear a lot about making things work by themselves, and you have to have GPUs to run AI, but we don't hear a lot about matching the usability of the hardware and the software. We don't talk about matching that to a use case and looking at a whole pipeline of an AI pipeline of the different places where you're going to have different types of compute, number one, and then how to optimize the hardware you do have for those use cases. Coming from storage, that's what we saw a lot.
Any new version of Oracle, any new version of Exchange, there was already, when it was time to launch, you had so many, well, if you're running it like this, or do you want it to do this, you had all sorts of architectures you could look at and watch the tuning and where it was. How do you use what you have without having to buy the latest GPU? How do you do that?
I don't know if it's because it's still so early, or maybe OpenAI has figured that out. But what is your workload that you're tuning it to? I don't see any conversations about that.
All I see is about the type of GPU and how much it can do things. You don't see anybody tuning them that I see. I don't know.
Right. I do think that it's also important to kind of separate the training components from the inference. Right?
Training organizations have more control over the workflow because it's more batch. Right? Latency is not that important.
It's more about how much can you process. So I think for organizations that have 11% where they do training, they have very little excuse as far as I'm concerned. On the inference side, because control is more driven by usage patterns, I think 11% is still very bad, but I give them some credit for the difficulties of dealing with that.
And I think the other thing that is worthwhile mentioning is that today, NVIDIA and other organizations are kind of pushing hardware that is supposedly agnostic to the workload. In other words, they're pushing GPUs to say, well, you can do some training, you can do some inference. And so now customers are kind of building systems that kind of are half and half, if you wish, which is a whole different segment.
But I think nevertheless, I think the push for hardware continuously as the answer is a concern. I do think that organizations like NVIDIA is investing, quote-unquote, acquiring companies that provide these services, or at least through software. But there's still a long way to go.
Yeah. It is an interesting situation because NVIDIA, and this is not an accusation of them by any means, but they benefit when their hardware is sold, whether it's sold for practical use or not. They benefit when it's sold to be sitting in a crate on the loading dock as well.
But ultimately, I think that NVIDIA is very much aware that they don't really benefit when their stuff isn't powered on and running and making practical use. Because the short term, yeah, they made the sale, but everybody that I've talked to in the industry recognizes that in the long termHaving resources sitting idle or sitting in the box is bad for everyone, and it is going to come back and haunt us. So, I know, Frederik, you spend a lot of time with these companies getting those GPUs out of the box, powered on, and in production.
So I think that we have to consider various phases here of the utilization question. Certainly, there's the basically putting it in production aspect, then there's the making better use of it aspect, then there's the what are we doing with it aspect, and then there's sort of the ultimate system tuning where we want to see, okay, if we're getting high-level use of these hardware resources, is there a way that we can get even better or more use? And also the question, as you just brought up in terms of training versus inferencing, whether the same hardware will be useful long term for both points.
So, let's zoom in on that for a second, Frederik. What is your feeling about that? I know that you don't have the answer, but ultimately, if companies are spending money on training hardware versus inferencing hardware, are they going to be able to use that hardware interchangeably?
That's their goal, right? Maybe we should look at it a little bit historically. Historically, if you go back a few years ago, organizations were really doing end-to-end work, meaning they were responsible for training, and then in the end, there would be a model, and that model would then be deployed.
So they were kind of responsible for the training and the inference side. And some organizations were doing training on-prem and inference in the public cloud, and the public cloud made a lot of sense because of the elasticity. Today, it's kind of a shift where the large language models from OpenAI and Meta are being used as a base model.
And so now we're doing fine-tuning, and fine-tuning is kind of in the middle between training and inference. And so organizations are kind of looking at going from two separate systems, meaning training and inference, to something that, from a hardware perspective and a cost efficiency, makes a lot more sense. And NVIDIA is playing right into that game.
The challenge is that it's not just about buying the hardware, it's also kind of customers understanding that they have to approach the problem differently. It's not just training, just inference. And I think that's one of the reasons why there is a lot of concern about efficiency, because people have no understanding on how to efficiently tune and switch between the training and inference.
And so that makes it very challenging for them to define efficiency. The problem with saying you only have 11% efficiency, it's kind of challenging because everybody can measure things differently, right? So 11% measured by one company might be 24% by somebody else, depending on what exactly they're measuring.
You could, for example, use metrics where you use memory consumption of a GPU as a metric as opposed to cores being used. It's definitely a challenge. Are there any guidelines out there or any matrices out there, or is it a one-on-one every time you go to a new organization, it's, "Okay.
" But are there guidelines any place that people can look at to know, okay, I know I don't want to be 11%, but is 40-something percent enough? Do I need to be 97%? Where can I reasonably expect this hardware to take me?
So first of all, no. As far as I'm concerned, there's no standard. Just to give you an example, with CPUs, you can debate about core usage and how efficiently those cores or threads are being utilized.
GPUs is a little bit different because the main focus with GPUs is memory. You really allocate memory, and then with the memory, there are cores. So it's not like you can compare apples to apples.
And if you look at a system efficiency, there's a lot going on, right? Maybe your data is too slow, maybe your network is the bottleneck. And the problem is that the bottleneck always changes.
So I could come up with a metric that basically says, "I'm going to measure the memory usage of my GPUs," and maybe that's really high, but it might mask the fact that my network is dog slow or has latency issues and is masking another problem. So is it standard? Are there standard metrics?
Probably. You can probably have MLPerf and other organizations do have, or MLCommons have ways to provide a metric, but in the end, it all comes down to organizations using the same metric, and then you end up with an apple-to-apple comparison, which is definitely not the case today. Yeah, and we've spent years in the storage admin, the storage pundit world, trying to come up with a taxonomy to measure utilization and adjusting that according to the different ways that storage could be deployed and utilized.
And of course, the same is true ofGPU processing power, and so on. You look at the work that's being done by MLCommons, for example, with their benchmarking. I would say they're firmly in the middle of coming up with taxonomies and understanding of system performance, and yet, well, I don't know, maybe I'm not doing them justice.
We'll see. But it seems as though that's something that is really actively being developed with the various benchmarks they have. But yet, as you're saying, different workloads may have different results.
Your mileage may vary. Which is actually one thing that I really love about some of the MLCommons benchmarks is that they're very practically based. They're not based on arbitrary flops or MIPS or whatever.
They're based on, here's a workload, run the workload, see what happens. Which is, I think, a much better benchmark. But of course, we've seen over the years with every kind of system benchmark, every time you do that, the goalposts tend to change, and you end up, for years and years, running some old benchmark that doesn't make sense anymore.
Look at Geekbench, for example, where it has had trouble keeping up with multi-core systems, for example, and accurately reflecting those. So certainly, that's an area that we need to think about. So the taxonomy and the measurement and the kind of benchmarks that we use.
And we can make a case, I think, that system utilization isn't as low as it looks because hand-wavy reasons. But ultimately, the question comes down to it, are we making good practical use of these systems? So let's, I guess, move forward a little bit from the question of percentage of utilization.
Is it practical? Are we getting good value from these systems? What do you think of that, Gina?
I feel like that's a trick question. It might be. Since you threw it to me first.
I think there's all sorts of ways to get practical use from this, and I think there's a lot of applications that are just coming out that are fantastic. If you look at the medical field, for instance, there's crazy stuff going on. I worry about some of the things that we might be going to and using this technology for gambling or things like that.
But there's so much good that can come from it, and I'll leave my normal naysaying behind. I just keep thinking about when you asked that question, my first thought went back to something Frederic said about the users. I think the way that perhaps that it's being deployed where it's like, "Everybody try it and see what you can come up with," probably does make it really hard from the utilization side, because you probably get somebody figures out something cool, and everybody jumps on board in an organization, and that's just hammering the infrastructure.
So I'm not sure that that's the best way to go about things. At some point, this has to stop being a science experiment, let's see what we can figure out. And there are tons of companies that are running it this way.
I think about the startups that can, looking through all of the medical documentation it takes to get an insurance company to approve a treatment, and trying to make that easier and faster so people can actually get their treatments versus waiting and waiting and waiting. There's so many ways, and there's probably other great ways. That's just the ones coming to mind.
I feel like we need to start concentrating on this, stop looking at it as my friend, the AI, that's going to write my homework for me or write my blog post for me or whatever, and start really thinking of how do we use these important resources, and that would probably help on the inference side for the hardware to be able to keep up. Yeah. I think no matter what, the AI infrastructure has to support the business, no matter the efficiency, right?
Even if you have a very efficient system from an infrastructure perspective, but it doesn't support the business, then you're not really helping yourself. I think the way we actually talk to our customers is that, one, they have to understand their business and then come up with a use case that works for them, and then provide an efficient solution that also is very adaptive. Right?
AI changes or moves so much quicker than traditionally workloads, such that it's measure, fix, and measure again. Right? That's really the whole goal.
If you're at 11% today, maybe that's not great, but at least make it such that tomorrow you're improving that 11%. And I think that's the other thing where I'm seeing is that people are not looking at improving. They're just making the situation much more difficult by adding more hardware to the equation.
Yeah, that's actually a really interesting point because you can really compound the problem by continuing to deploy more and more. And this is the same with kind of almost anything. If you're not managing what you have well already, then adding more, I guess, in a way, it seems to solve the problem, but it doesn't necessarily really solve it.
In fact, it doesn't really solve it, and it makes it worse because then you just have more to not manage properly and to try to justify. And I guess you also-- I want to get back to a point that Frederic made at the very beginning, which is sort of this, let's call it a CYA factor of folks who have purchased... " But yet, I don't know that people have that in them to really address the root of the issue.
Do you think? Well, I think you end up like xAI, if that happens, and you find some startup that wants to rent out your equipment to do their work. So that's a way to go with it.
Yeah. Well, that was the origin story, for example, for AWS. They were trying to build a flexible and scalable cloud platform, and Andy Jassy realized that what they were building could be useful by other people, and also that it could help them offset the cost of the investment they were making.
" Yeah. Is that practical? Is that something that could happen where a company could end up becoming a service provider instead of just using the stuff themselves?
Yeah, I think that's what the neo clouds, in my book, are trying to achieve. Right? Is to provide you an ecosystem, and they're saying...
It's the concept of AWS, but then more specifically for AI and workflows, basically saying, "Hey, we have a complete stack for you, and it's multi-tenant, so consume as much as you can. " At least that's what I'm seeing now, certainly with the supply chain issues, where a lot of people going to the neo clouds, and kind of do what they used to do with the AWS traditional workloads many years ago. I wonder if there's an opportunity for software there as well.
There's certainly a lot of companies that are working on AI compute platforms for enterprise use. I wonder if there would be a potential for a compute platform that also allows others to use your hardware, or if that would even be desirable. I don't know.
That's kind of what xAI has, right? Because they have the compute platform, and then you can use their model, which is right there on that platform. Is that what you're saying, or?
Yeah. Well, I'm thinking an end user company. I don't want to get necessarily down that.
Gina, you've got experience with these companies. Do you think that there would be any appetite at a bank or a manufacturer or whatever to participate in some sort of use my hardware scheme? Maybe, and I think they do it already.
Frederick would probably know better than me. But I think they're used to the cloud model right now, and so I think if it made sense within their business model and their security model, of course, then I think they'd consider it, at least. Yeah.
I think in the life science world, this already happens, where they have centralized clusters where competitors actually are running on the same environment because their main goal is to deliver products and improve products for society. And so in those cases, they do share amongst competitors, which is an interesting model, which I've not seen being exploited beyond life sciences. Well, so I guess in summary, let me ask you the question that I asked you at the beginning, or that I posed at the beginning, which is, when we talk about utilizing AI, we've often talked more about how are we going to use this technology to do something practical and functional for the business.
But we should also be thinking of how are we going to be able to use this technology in an effective and practical way. What's your reaction to this question? Do you think that this is a new question?
Do you think this is something that people aren't yet taking into consideration? And do you think that there's going to be a path forward for this? So Frederick, I'll throw it to you first.
Now that we've had this conversation, what's your take on utilization of AI? Yeah. So my take is that it's definitely being underutilized.
However, I think there needs to be a mentality change with end users where they realize that just buying more hardware is not the solution, and that they have to look at more efficiency. It's a challenging mentality shift. " It's difficult.
I don't see it happening overnight, but it's definitely the only way to get out of this supply chain and efficiency issue. I think it's a maturity issue as well. And so understanding the infrastructure and understanding your business case and what you're trying to do, as Frederick said, is essential.
So it's no different than any other type of computing we've ever had. But at the moment, it's so expensive, and the parts of pieces are so expensive, that it's critical to make sure you're utilizing this in a way that makes sense for your business. " Unless apparently if they work at OpenAI, in which case, they're doing great.
So, I guess, thank you so much, both of you, for joining us, and sort of sharing the type of behind-the-scenes conversations that happen at Tech Field Day and at industry events all the time. I think that's really what, for me, is interesting about conversations like this is that this is what we talk about. We have these sort of big philosophical discussions.
We think about how things could be improved, and it's great to share that with the world. So before we go, where can people connect with you and continue conversations like this with you? Frederic?
Yeah. So first of all, I will be joining AI Infrastructure in a few weeks. You can find me on LinkedIn as Frederic V.
com. Yeah, and the same for me. com.
And as for me, you'll find me here on the weekly "Utilizing AI" podcast. You'll find me Mondays on Techstrong Gang. com to learn more about those.
Thank you for listening to this episode of the "Utilizing AI" podcast. If you enjoyed this discussion, please do subscribe on YouTube or in your favorite podcast application, and consider giving us a rating and a nice review. This podcast was brought to you by the analysts and experts from the Futurum Group, where insights meet AI.
ai, the "Utilizing AI" YouTube channel, or the Techstrong TV app. Thanks for listening, and we'll catch you next week.