OverClock Labs CEO Greg Osuri on Solving AI GPU Shortages with Distributed Computing
In this Techstrong.ai Leadership Insights video, OverClock Labs CEO Greg Osuri explains why a mismatch between the demand for the latest graphic processor units (GPUs) and actual supply is creating an imbalance in the artificial intelligence (AI) ecosystem that might be better addressed by relying more on distributed computing.
Transcript
ai Leadership Insights series. I'm your host, Mike Bazar. Today we're with Greg asi, who's the CEO for Overclock Labs, and we're talking about AI data centers, and whether or not we have enough capacity or that capacity for that matter is in the right place.
Greg, welcome to the show. Great to be here. Mike, thanks so much for inviting us.
What's your assessment of what's going on here? I think on, maybe in the short term, we're a little bit ahead of our skis, and maybe we have more capacity than we need, but then long term people are saying, we're gonna need more data centers than you can fill the sky with. So from your perspective, what's, what's in the middle here, or what seems to be driving this conversation?
Uh, we're in an interesting place, uh, depending on the week or the month rather. Our own, you know, coming from our own numbers, uh, from demand supply standpoint, we seem to be oscillating, uh, in terms of demand and supply. Um, for some reason this month, like our utilization rates for let's say H one hundreds, which are a pretty good, um, GPUs is right now around 98% last month.
This particular chip had, I don't know, 60%, and the month before it was, you know, fairly high as well. Um, we seem to be the problem so far, at least on the surface, seems to be misalignment more than in that type of compute, more than, um, like demand for everything, right? So 20, 23, 24 timeframe.
I mean, if you had a GPU, it was just coin, you know what I mean? Like, so you had like, especially the, the, the higher density ones like H two hundreds or H one hundreds back then, uh, you could not get anything on demand. Uh, people were paying, you know, uh, I mean, anything that can, that the cloud provider would ask them, because those enormous shortage of supply, um, and, and really wrapped up their H 100 production from making, um, um, half a million units to about 2 million units.
And that seemed to like, at least ease the demand, uh, ease the supply constraints a little bit. But now we are seeing every bottle of that for the H one hundreds again. Um, and I think my thesis is because it takes about 18 months to two years to reconfigure a fab to produce a new chip.
And, uh, you know, when you reconfigure the fab, you want to take advantage of it, but in order for you to take advantage of that older chip, the market moves away to a new chip, and then NVIDIA is moving. Its, I mean, the design is such moving significantly faster than production. So every time Nvidia announces a new chip, the demand for the newer chip, you know, is much higher than the older chip, but the production for the older chip is a lot more than their demand exists or something.
So there seems to be misalignment, at least from our observation in terms of, you know, NVIDIAs rapid, uh, growth and a lot of the growth, you know, uh, ensure it has to do with the current demand, but also they have to keep producing new models to keep the market happy. Um, when every time they introduce a new model, the demand for the newer model goes up and the supply is not quite there yet, and the demand for the older models come down. So that's why you see, uh, cloud companies that own a lot of hardware, uh, and not seeing the same amount of demand for the older hardware.
Right? And we looked at price fluctuations, particularly for, uh, something like H 100, and we try to predict what the future price is going to be. Uh, and we look at historical prices, right?
Uh, what we've noticed is over a five year term for A GPU, the average drop, we're talking about training grade GPUs, right? H 100, C 100 V 100, whatnot. Uh, the, the average, uh, annual drop for price is about 20%.
The biggest drop you see is in the, the first to second year drop, right? So if you are someone like a cloud provider that gets access to the latest and the greatest chips you make money on year one, year two, year three, you, you lose. But if you're someone that gets, you know, compute for that model in year two, year three, um, uh, the you, you don't quite get the cost benefit that you're, you're, you're usually getting when you, when you have year one, and everybody wants year one, right?
Like B three hundreds are out, or, or, or at least they will be out very soon by end of the year. Um, and everybody wants B three hundreds and no one wants H two hundreds anymore, even though H two hundreds a very, very good platform, right? So, uh, I think that's how you're seeing fluctuations in the short term.
Why? Because our, there are several key challenges for AI to, to be flexible in terms of, in, in terms of what kind of chip, uh, they can use for both, for training as well as infants. Training in particular is very homogenous in the sense like, if you train on, let's say H two hundreds or H one hundreds, you have to stick with the same chip.
You cannot, you know, uh, be heterogeneous in terms of mixing and matching what chips that you need. So if your data center has lots of a one hundreds and some H one hundreds, if you train on H 100, you cannot leverage all the un underutilized a hundreds. So there's a mismatch and everybody wants to train on the latest and the greatest because the price performance is, is drastically different when you have the latest chip versus even one generation later, right?
And at scale, that makes a huge difference, like massive difference. And particularly we have another concern, which is energy, right? The, the energy prices.
And there's a whole, like, all that, I think people don't understand that very well. Um, and, uh, so ultimately the kilowatt hour to performance to the flops is what I call the real metric, right? Latest.
And the greatest chips, like B three hundreds have significantly better, you know, uh, ratio how many flops they can extract out of, uh, one kilowatt hour compared to just a previous generation. So, uh, and when you're a hyperscaler, that means anything you have, if you have data centers over 50 megawatt capacity, your energy costs are enormous, right? Because, um, it's very, very hard, if not impossible in the west to get large amounts of concentrated power in a single place, um, because of our energy problems, right?
Like we, I mean, it's easier to get more distributed energy, smaller quantities in lots of different places, but it's very hard to get large quantity in a single place. In fact, um, and if you look at the energy demands for, for training clusters, it's doubling every two years, right? So, uh, like for the, for producing the state of the art model from, you compare from, you know, your G three to g PT four, G five, we don't have the data on GP P five, at least from like graph two, graph three.
Um, if you look at historical, uh, uh, I dunno if you can hear my dog, but he's going crazy because he's a squirrel outside, okay? Uh, the, the, um, the amount of, uh, uh, the energy and stock capacity required is doubling. So opening a data center, for example, and in Albany, Texas is currently scheduled to come online with about 300 megawatt capacity.
The GR data center, the XA data center in Memphis, uh, has about one 50, uh, megawatt capacity. And if you look at the GR data center, they're only drawing about seven megawatts from the grid. The rest of the energy they're getting from burning LNG, we're talking about Tesla, or like Elon Musk's companies are actually burning, uh, LNG because there's no reliable way to get energy.
So if that's doubling every two years now by, you know, 20, 20, 27, 28 timeframe, we are looking at 600 megawatts. And by 2030, we'll need a minimum gigawatt, uh, capacity to train the stay the year model only thing that can produce a gigawatt capacity. Um, uh, or the best way to purchase gigawatt capacity is nuclear, right?
Um, and in us, we, the last nuclear reactor we built took about 14 years. All the nuclear reactors that we have currently are utilized pretty much full time. So we cannot produce new energy or new source of energy.
And that means there's higher demand for energy. That means the prices are gonna go up. That means you want to use the latest and the greatest chips to reduce the tho those or improve your price performance.
So a lot of factors that get into hyperscale economics, uh, for the first time, you're seeing a reversal in, uh, e econom economists at scale working, right? Like it's almost cheaper. And residential is significantly cheaper than commercial right now, uh, even though the commercial demand is kind of eating into residential prices as well, but, uh, you know, in areas like Virginia and, and whatnot.
But, but at least even now, if you can like tap into a residential grade somehow, um, uh, that's the best way to to, to reduce the cost. Anyway, long story short, there are a lot of factors that are going into the demand supply and the economics we're seeing. Uh, and most of it comes down to, uh, you know, energy per flop.
Are we, uh, Mimi too obsessed with the latest and greatest? And are there AI models that could be trained or run perfectly well on older processors and systems? And maybe we should just be smarter about which ones we use for what?
Well, what's perfectly well is the question here, right? So for example, um, a 100 has about 80 meg, 80 gigawatts, uh, of virtual memory. So the gig, the virtual memory is what determines how big of a model you can do.
You can run, um, if you were to run, let's say, a deep seek, which is about six 75 billion parameter model, you know, you will need, um, you know, the 80 gigawatt, 80 GB chips, about at least eight of them clustered together, um, and serve inference for a deeps, right? Uh, now it can run okay on 80, uh, you know, I know HA 100, but the inference performance is gonna be significantly slower compared to, um, something like, um, the H one hundreds or H two hundreds and above. So what that means is your response time is gonna be slower.
So depending on the kind of application you're looking for, yes, it's possible to run on all, all the chips, but your response is gonna be, um, like not optimal for your use case. So that comes down to it. Or can we do a small model?
Can we do necessary experts? And looks like that's where the world is trying to go to, is widely believed. J GT five was supposed to be that, but we all know the, the, the quality of charge GB five is much.
In fact, it, it hasn't improved. It, it, it, it, it, it detrimental. I mean, they actually went down.
So, um, the, the, the state of like, I mean, we know that bigger models, uh, that use a lot of energy are usually, I mean, much, much better than the old models. We know that, uh, but it's not sure about the economics of charge G five gb, I mean, GB five, um, and, and because it's closed, so not too many, there's a lot of speculation as to what kind of model or what kind of, uh, you know, training or what sizes this model is. But, uh, it's widely believed that it is an agent tech model that, uh, you know, that, uh, that, that, you know, leverages a lot of small models.
Are we also maybe miscalculating how efficient the LMS might get over time? And we're kind of making projections for build outs based on what we think will be the size of an LLM, but there could be ways to more efficiently build, train, and run them. And might we wind up with another situation where, you know, we overprovision fiber and we wound up with all this dark fiber, Right?
So LLMs are getting efficient. I mean, at least like comparing from G PT three to GPT four alone, um, you know, of course the algorithms are getting much better, but also like the, the chips, uh, contributing to the efficiencies is huge. 5 is widely believed to, uh, I think each pro each, um, um, each prompt was on an average consuming about, uh, 10 watts per pro, 10 watt hours per prompt, um, which is quite a bit.
Um, and now, uh, and now we're, uh, wait 10 watt hours, well, three by, yeah. Now the thing is like 10 times cheaper than that in terms of energy, right? With HH one hundreds.
Um, um, so we are almost like, so if, even if you have like a whole order of magnitude different from every GPU model, another Berkeley, you know, I mean the, the Lawrence lab at Berkeley, uh, in collaboration with D Department of Energy did a very, very deep study about energy, uh, you know, consumption for ai, including the efficiency gains. Uh, and they look at several efficiency gain models, and it's very deep, deep, deep, uh, deep, uh, analysis, including that they assume or they, they concluded rather, um, about 12% of US energy will be consumed by AI data centers. Um, and this is a very conservative estimate.
The reality, if you look at Gartner's world or some of the other, you know, analysts, um, the, the reality is most likely about 30% of us, uh, energy production will be used for ai, right? Including inefficiency gains. Um, so, which is not totally on the, on the, on the, you know, outfield, considering that most people still don't have access to basic ai.
I mean, we are in, I gpi tragedy is relatively expensive, very expensive for, for common usage, right? But most people don't even have access to that. Uh, and we are like in barely scratch the surface about how much inference demand is gonna come through, right?
I mean, all the way from medical, I mean, medical medicine hasn't even, um, you know, uh, hasn't really like embraced AI as much as they could, um, including every piece of transcription, automatic billing, uh, second eyes on, on diagnosis, and a whole lot of incredible, incredible applications in, in medicine alone that hasn't been unlocked because of regulatory concerns. I mean, there's a lot of, like, a lot of areas that are still early in evaluation that could mean quite a lot. I mean, looking at medical expense alone is about 20% of American economy, right?
Like something crazy numbers, the Medicaid. So we have massive like blocks of, uh, econom, I mean, sub economies in America that are very early stage evaluating ai and all that is gonna get unlocked in the next few years. Uh, even models are advancing at such speed that the market hasn't quite caught up in, in terms of demand, right?
Like by the time you, you know, Chad G three three is a great model, right? Like, uh, you know, like the moderate advancement we saw Chad GB three took years of AI into, into perspective. And if you look at medicine, it's not even a, you know, able to leverage cha GB three, like for basic diagnosis.
So a lot of areas, that product hasn't quite caught up because experience hasn't, people haven't figured out what's the best way to deliver this experience to some of these, um, these traditional fields, right? Um, um, and we, we are slowly seeing a lot of like normies, I, I call it normies, like my, for example, my AV guy or my electrician now using charge GBT to answer some of the questions, right? I mean, we are just starting, starting to see that.
But imagine, um, you know, because I introduced him to chat GPD and he had no idea that he could actually take a photo of, uh, of any issue he has and just, you know, have chat GPD recommend solutions for him, because that makes him a much better electrician, right? Simple things like that. People haven't figured out, uh, basic like usage strategy GBT.
So we're not, we're not there yet. I think the, they're we're severely underestimating the demand. Mm-hmm.
Uh, I guess depending who you talk to, you talk to Eric Schmidt, he'll say like, 99% of the world's GDP will be spent on ai, which is a, a little stretch. Uh, it's non-zero, but that's a little stretch, right? But we're talking even government reports, like DOA reports, uh, that are very conservative talking about it.
And I testified to Congress recently, um, uh, before Congress about two months ago on this particular subject. There, there is hope though. There is hope, uh, it's not all view mean like, oh, um, or all going in a negative direction.
Um, the hope that is, is if we can figure out how to leverage nitrogenous GPUs, right? Like my gaming machine at home, which has a 50 90 or, and a 40 90 could very well take part in a training run, but it cannot right now, because, you know, training is asynchronous. That means, I mean, synchronous and, and we have the is is only as fast as the slowest node in the, in the, in the training batch, basically.
So if you have H one hundreds or a one hundreds, even the H one hundreds would complete the task fairly quickly. It still had to wait for the A one hundreds to complete before it can synchronize all layers. So it doesn't matter if you have, you know, all H one hundreds and a few a one hundreds, the batch will still be determined by the size, by the speed of the A 100.
So, uh, there are researchers now, um, that, you know, have figured out how to leverage, uh, heterogeneous GPUs. And I, I was recently at ICML, which is International Conference for Machine Learning, which is the most, it's very academic, most prestigious, uh, machine learning conference in the world. Um, and we pay attention to what kind of subjects are being discussed.
And a big part of this year's ICMO was distributed training, um, and we had like six papers on distributed training, uh, all addressing different, different parts of the problem. Uh, but asynchronous training is seemed to be extremely exciting, uh, asynchronous in the sense, like even it can, it can tolerate faults better. Uh, like right now, if you have a training run that's run into multiple batches, if one batch, uh, if there's a fault in a single GPU in a single node, in a single batch, you have restart the whole thing.
Whereas asynchronous means you don't have to, uh, and you can do hydrogenous GPUs. That means like you don't have to wait for the slowest machine, uh, to, uh, to complete a, to complete a batch. And so you can do asynchronous.
We have, um, um, you know, uh, a distributor for example. There are, there are, you know, aspects where, uh, so, so the reason why you want a lot of GPUs in a single place is because, um, there's something called gradient synchronization in, in training where you had to, you had to synchronize all the, all the gradients, uh, periodic periodically in order to train. Uh, that means these nodes are very chatty.
You need, you know, end by end, like the, the speed is determined by the complexity of the nodes. So more nodes there are, it gets lower the network, right? Because they need to communicate pretty, pretty often, um, uh, and their algorithms now that, that, that reduce the amount of synchronization, make these nodes less chatty or the training less chatty by using several techniques.
And a lot of companies are working on different techniques on how to reduce the amount of communication. So they're called locum. So locum algorithms, uh, Google DeepMind is pretty famous for their paper on Dial Co.
Um, and news research, uh, distro, these are very popular algorithms that are taking a lot more attention now. Uh, so yeah, there's a lot of hope and like reducing, um, uh, uh, communications, improving asynchrony, improving heterogeneity, improving fall tolerance. Um, and on top of that, you can add incentives for someone to participate.
I think that's a missing piece where all these technologies come together and you can actually have a, a viable open alternative training, uh, uh, mechanism compared to what we have, which is a very closed and very, um, you know, controlled, uh, environment. Hmm. So we only have a couple of minutes left, but if you had to summarize it, what's your best advice to folks in terms of how they should approach this whole space?
Is there something that they should be thinking about more than just say, having patience, Buy GPUs and buy solar? All right, there you go. That's the advice, because I'm, I cannot emphasize enough on the importance of energy crisis.
Uh, right now we are seeing, but GI alone, um, the utility companies in Virginia are paying 20% more for the same amount of energy for last. They, they, they, they got last year, it's growing significantly past faster than the CPI, uh, energy prices. Um, and there's no hope.
So you may want to think about how the future energy market is gonna look like, and having solar at home alone is, solves a lot of problems. But if you add a GPU on top of it, the very high chance that that GPU could partake in a training run tomorrow and in exchange give you some of the inference revenue. So there's quite a lot of, um, and I'm, I'm very hopeful and I'm very excited about this future.
Uh, and also it's better, right? Like right now, single data centers are targets. Uh, you know, these are like just sitting there.
We have about 500 hyperscale data centers in America. Everybody knows where these data centers are primarily in like the, the data central allies, the, you know, the Virginias of the world and in, in, uh, in, uh, in New Jersey of the world. Um, and that's not good from a security posture, right?
So we want a more distributed, more decentralized, uh, data center, uh, market. Um, and the way to achieve that is to have home ownership of GPUs and, and energy. That's really the only, only way.
All right, folks. You heard it here. I think at least even in the age of ai, uh, failing the plan is planning to fail.
So we should think about long term and what we're after here, because if we don't, we won't have any AI or not enough of it to go around. That's for sure. Um, Greg, thanks for being on the show.
Thanks so much, Mike. All right. Thank you all for watching the latest episode of the Techstrong AI Leadership Inside series.
You can find this episode and others on our website. We invite you to check them all out. Until then, we'll see you next time.