AI hypercomputer and GPU acceleration with Google Cloud
Dennis Liu, a Product Manager at Google Cloud specializing in GPUs, presented on AI hypercomputer and GPU acceleration with Google Cloud. Liu covered Google Cloud’s AI hypercomputer, from consumption models to purpose-built hardware. Focus was given to Google’s cluster director for managing GPU fleets.
Dennis then moved to the hardware aspect of Google Cloud’s AI infrastructure, discussing current and upcoming GPU systems. Available systems include A3 Ultra (H200 GPUs), A4 (B200 GPUs), and A4X (GB200 systems), which are built on Rocky on CX-7. Also discussed were two systems coming in 2025, the NVIDIA RTX Pro 6000 and a GB300 system, offering advancements in memory and networking.
The presentation also featured performance projections for LLM training, with A4 offering approximately 2x the performance of H100s. The A4 was described as a Goldilocks solution due to its balance of price and performance. There was also discussion on whether Hopper-generation GPUs would decrease in price because of newer generations of hardware.
Presented by Dennis Liu, Product Manager, Google Cloud. Recorded live in Santa Clara, California, on April 22, 2025, as part of AI Infrastructure Field Day. Watch the entire presentation at https://techfieldday.com/appearance/google-cloud-presents-at-ai-infrastructure-field-day-2/ or https://techfieldday.com/event/aiifd2/ for more information.
Transcript
Hi, uh, My name's Dennis Liu. I'm a product manager here at GCCI. In particular, I'm on the GPU team.
Uh, I own a little VM called a four with a B 200 GPUs. So when I show you the slides, and I prefer a four, it's not necessarily a technical reasons, it's just, it's my baby. Uh, two sections on the slides.
One is like, sort of the setup stuff. Um, we'll talk about the, uh, management layer for a little bit, and then we're gonna talk about the hardware. That's what my team does.
Um, like I said, that's where we get to the part of this particular day. Four, you've probably seen this slide like six times today, but, you know, we do all of these things literally across the entire stack of AI stuff. All right.
Uh oh. So it's just slow, I think. Okay.
I gotta click. Um, so we, you've probably heard about ai, hybrid computer that goes from all the way up the top. Includes our consumption models, very flexible, all the different kinds to meet all different kinds of needs and price points, as they always say, all the way down to the bottom where we talk about purpose built hardware, which again, is me.
So, uh, I'll be talking mostly about focusing on that area. Um, in the middle, we have our entire software stack, you know, we're experts at that too. I, I don't know how much you know, users are GKE non GK users, but, you know, we we're really good at all that stuff.
Um, any, if you, like I said, if you have any questions about it, you stop me and, and ask a question. This, uh, slide did not paste over very well. At the top it's supposed to say cluster director, sorry, it's, it's in white on a white background.
Um, so Cluster Director, uh, I think you had a session about that early in the morning. It's a good time for us to reinforce it 'cause we literally built it for the GPU ML use case. Um, cluster Director is about managing your fleet of, uh, GPUs all together at once as a cluster.
It gives you some visibility, gives you, makes it easier to manage your workloads, uh, insight into your performance and your locations and, you know, all the good things. Um, it's, it's, you know, new, uh, so I'm sure there are short shortcomings. So as you read about Cluster Director, as you hear about Cluster director and, and things that people are using it for, we would love to get feedback on that.
Right. Uh, my colleague actually works on it, and he's always looking for feedback. We're looking for feedback on everything, right?
Um, we are in the middle, you can see we highlight the different orchestrators. We are trying to be agnostic and be, be bring your own kind of things, things wherever we can. So make sure that, uh, you know, you, we, we get feedback about that too.
It does. Yep. What, Uh, guys are available An orchestrator.
Uh, I'm not an expert in that area, so, uh, we, we will have to talk to somebody else afterwards. Uh, I think the, in general, right? The orchestrator needs to know things like where my hosts are, so I know to place the workloads on the right VM so that they're talking to each other.
Those APIs are available. You don't even need an orchestrator for that. Like, that's just a baseline, API that we expose.
And you could do it like through command line if you wanted to, but nobody would. So all of the things that, that you like would expect to need that are available through GKE or available through s LM are also available through like the underlying CLI or API. And then you could script whatever you want on top of it too.
Thanks. In, in terms of the naming you mentioned, so if we stop, uh, we, we, we've stopped now, uh, the conversation, we're at the director level. We're then moving into the kind of the manager's level, then we get to the orchestrator levels, and then above that, uh, or below that, sorry, it's the foundations levels, um, above the director level, or is that like something central, or what would be that theme above Director?
Whew. Um, Okay. I'm, I'm gonna, I'm gonna crack a little joke.
Okay. We are being recorded. So, so this is just a joke.
People, um, that has been discussed internally a lot about our positioning regarding this different layers and using these different words that sort of give people, uh, the impression that something else is going to happen or is coming, is gonna have like general name or something like on top of that. Um, it is, it, it, we don't have anything at that level. Let's say right now, cluster director is the highest level object that we're trying to like, build out conceptually.
And from like the slides, you'll see ai, hyper computer, right? Ai, hyper computer com encompasses everything. Mm-hmm.
Um, we are hoping that you view cluster Director as like the highest level thing you need to manage, like an AI workload. Thank you. Cool.
Uh, so this is the part where we get to the actual hardware. So these three are almost available, uh, in, in the sense that the third one says it's in preview. Again, my baby's in the middle.
It's a four. Um, so going left to right, just in case we, you haven't already covered this, I gotta do it. Uh, a three ULTRA with H two hundreds, a four with B two hundreds and a four x with the GB 200 systems, right?
Uh, if anybody remembers or has used our H 100 systems called a three high and a three mega, those systems were built on something we call GPU Direct, TCPX, fast Trek, uh, I think that is our official name. Uh, we are now moving to Rocky on CX seven for all three of these systems. Uh, that should make it slightly easier to bring up your oss, bring up your systems, bring up your workloads.
Um, performance should be less slightly better. We've also doubled the bandwidth for GPU to GPU communication. 6.
8. What does that mean? Uh, it really means 400 gigabits per GPU hang up.
Um, any questions about these three systems? Because you might not be surprised, but there's two more coming up on the right hand side that we're gonna talk about in a second. Two, one.
Sounds good. So the two systems that are coming soon, sometime in 2025 are, uh, we didn't write the names on them, so I'm not gonna say what those names are, but the Nvidia RTX Pro 6,000, which previously before Nvidia announced, that name would've been called something like L 40 s next, for example. It is the new mid range card.
Um, I believe the specs have been announced, uh, about the memory. Okay, cool. Cool.
Go ahead. Um, I just really afraid, if it's not in the slide, I don't don't wanna accidentally say it. Um, that one is coming.
Uh, if, if it, it's, it's actually very popular. We, we've already gotten a lot of interest from like inference users, uh, global users where, because this GPU is gonna be able to go into many different places that are other ones weren't necessarily able to. Uh, and then also for graphics users, right?
Like, graphics is going to be really awesome on this. Uh, you can't do that on any of the 100 series, right? So, uh, the, the last one on the right, again, there's no name attached to this, so I'm not gonna use that name, but it is the GB 300 system, which is basically just the GB 200 more memory, more networking, everything's faster, uh, takes up more power.
You can't even imagine. 6, so double and it uses a CX eights. Any questions about any of these systems that I can answer?
So just to co uh, cover very briefly, we, we, we already have a fours available today. I just found out, uh, that we can do spot VMs. So if anybody is interested in trying them out, you can actually go selfer quota increase and then create a spot VM all on your own.
No need to interact with anybody anymore. Dennis, I actually do have a question. Paul, our paradigm Technica, um, can you maybe give us like a, a, a cheat sheet about the level performance differences between do like a four is two x-ray three or something like that?
That is a great question. I'm going to jump a few slides ahead so that we can talk about it. How's that?
Okay, so you asked about performance. These are performance projections, I must say projections because all of the data that backs all of this is not real world testing. It is projections.
I will make a few more comments addendums to this slide that I think are safe to say. Uh, but I will also say we also have a lot of data under NDA that like we do share with customers. And I'm not supposed to talk about that right now.
Uh, but, you know, you know, if, if it gets to that situation, we do have a lot more numbers for that. Okay? So, uh, the basic, basic approximation I use is that, uh, a three ultra slightly better, you know, uh, very, very low double digits, like 10% better.
A four is two x better compared to H one hundreds. Um, in our real world testing that is, uh, if I say something, you're not gonna hold me to it, I hope, because I'm not promising anything. Sometimes it is better than two x, how much better.
And, and on which situations are, it's very specific. Sometimes it's better than two x, most of the time it is two x. Uh, I should also say that, uh, the headline here is very specific.
It says projections for LLM training. Okay? So if you are looking at non LLM use cases, uh, I care very much about that use case.
But it is harder to get performance improvements of that scale necessarily on these new GPUs because all of the research and the work and the effort and the transistors and everything is going towards improving LLM training and inference, right? Uh, on the right hand side is what you'll see the difference between, say the A four with the B two hundreds and the A four x with the GB two hundreds. Um, you'll notice that the, so the top right of the right graph says MOE training speed up, right?
So when you get to these mixture of experts, models, large ones especially, you're going to be a get a much faster speed up on the GB 200 system. Which one presumes will extend to GB 300. But we don't have performance projections for that.
That's why it's not on the graph. Uh, so my very biased personal opinion is that, uh, for most people, what you're end up using is a four. With the B two hundreds.
That's where you'll get your most bang for the book. Um, you can see on the left hand graph, that's LLM training dense, LLM training. Uh, there's a couple here, notes here at the bottom.
I'll read out loud 'cause you might not be able to see 'em. But the left hand side is 4K GPU scale. The right graph is 32 KGPU scale, right?
So yes, there are customers using 30 2K GPUs all at once in a training job. That's a very small number. And you got the gold for the Goldilocks, right?
Uh, yes. Um, that is maybe a coincidence, I dunno. Um, so, so, and, and the price, right?
If you go on our website and looked at the prices between a three mega, a three ULTRA and a four, you're going to see that the price for a four is not double the A three mega, it is not double the A three ultra price, for example. Uh, I'm not gonna say what it is. You can see it on the website.
I don't remember off the top of my head. Uh, so you're gonna see a lot of performance price perf improvement. We don't have the graph for inference.
Uh, nobody asks, but I'm gonna talk about it. So the graph for the inference is very similar. Uh, so actually with smaller LLMs, you're not gonna get much of a speed up going from a four to a four x because if you've got, you know, 70 billion parameters that's gonna fit in the GPUs, you don't really need the NVL 72 of the GB 200 system.
Dennis, this is okay to publish, right? I see next on the bottom. Uh, yes, I, what's going, I didn't get a chance to ask my boss, but I took it from his slide deck.
So the answer should be yes. Okay. Thank you.
Slides are public, so he knows. I mean, there's literally being broadcast on the internet. They're, they're now, really?
We didn't know that. Did you? You're live, you're live on the internet.
Oh my God, It's from next. Next, yeah. Thank you.
I haven't presented that next before. Everything So far has been the toothpaste out of the tube. And, um, I, I don't know if there's been any roadmap discussion so far.
So there's, okay, we talked about what's above now, what's, what would be to the right in your mind? Uh, okay, so, um, uh, I'm very used to presenting to customers under NDAs. We are not under NDAs here.
So there is a line, and I don't know exactly where the line is supposed to be drawn. What I do can tell you is just that on the previous slide where we showed the, uh, uh, you know, all the systems in a line mm-hmm. GB 300 is on there, okay?
And what that means specifically is we are going to ship something with a GB 300 in it, what it's gonna look like, what size performance. So there's this very interesting thing that the first performance numbers that we're gonna get are obviously from our partner provider that gives us the GPUs. Uh, we don't have them in hand.
We don't even have the final specs. It's, we do try to do our own internal projections, um, but often it comes from them. So I would probably just be regurgitating whatever they had published, you know, on the internet already.
So, any other questions? Okay. We might actually get to the backup slides.
Just Dave, if you guys don't keep asking questions, the backup side's a super boring. So, um, okay. So, So at, uh, this Keith Townsend RUM group, Jensen had this provocative statement at GTC that said all of these hopper, those H two hundreds you show showed earlier, you won't be able to give them away.
So who's using them? Um, I have a, I have a teammate who's the PM on the, on the hoppers. And, uh, he had kind of a sad day, the, you know, Um, and, and especially the H 200 pm because, uh, those, you know, you know, we're trying, we're trying to, you know, we have them, let's put it that way.
So actually, um, even after that statement, so that's been a month or so, right? We had with GTC mid, mid-March, and then we had next, uh, demand has, uh, increased, if anything, even on the hopper generation. There are basically two reasons for that.
And now we're getting towards, like what I say is not gospel, right? What I say is, is often also my only informed opinion. Um, B 200, we have it available, we have it available at a certain scale, right?
I can't tell you what that is, but it's, it is not millions of GPUs in a single cluster, okay? Hoppers on the other hand, we have a much larger scale. You can get more of them sooner in the places you want larger clusters right now.
And there are certain customers and, and companies and users that, uh, value that over waiting X number of months. Uh, the other way to say that is we are the only, okay, knock knock on wood, I I haven't read the news today. Um, maybe we're still the only large, uh, cloud provider that has B two hundreds in production, uh, GA wise.
So even if you were waiting for that, like who knows how long you'd have to wait, right? Uh, now on the GB 200 side, it's, uh, it's a also an interesting situation where there have been providers that have announced that they were going ga um, we announced a limited preview. I think, uh, oh, hope I didn't say that wrong.
But again, availability of GB two hundreds limited, it's going to take x months X period of time before it becomes available. And pe your business can't wait, basically, right? So we have a lot of people using Hopper.
We have a lot of people using H one hundreds. It, I mean, this is not to say too much. We have a lot of people using previous generations, even older generations because price perf is there.
And, uh, you know, not everybody is trying to run trillion billion, you know, trillion parameter LLM models, right? If you are running a smaller model recommendation model, image generation model, and you get your hands on an H 100 or an A 100, or even a L four, right? You could get, you might be able to get better price perf and you can get at that today.
You know? Yeah. That's not what we, we saw this, uh, even a few months ago when some of the more, um, uh, weird CEOs announced that they were building a new, uh, AI data center and they built them with H one hundreds and everyone's scratching their head like, why are you building it with H one hundreds?
And I think we're seeing it from a practical deme, uh, availability demand and what's useful, uh, from like, what's available if, if you, sure, I can eke out way more performance out of a B 200 or GB 200, but if they're available in one region or one area and my data is somewhere else, uh, well, you know, So that's, that's definitely a factor. Uh, we get that from directly from customers. Also, uh, another thing to think about is a, uh, recent overseas company that I want, it's not recent anymore, announced a very well performing model that they managed to train on a smaller number of GPUs, uh, and they did it by optimizing more, right?
And so if your code is severely optimized for hopper, even H one hundreds, not H two hundreds, you might be able to get better perf per dollar than moving to H two hundreds because you have to do all this reinvestment into redoing everything in the stack, right? Um, even if you move from one company's H one hundreds to another company's, H one hundreds data center per network performance is different. Uh, you know, the software stack is different.
You, it's, it's, it's really interesting and simultaneously annoying to me as a pm no offense, um, I wish everybody's workloads just worked the same optimization. You mean quantitation, right? Uh, no, I mean, lots of different things, right?
So like in that one particular case, you might manually write cuda kernels that are going to do very specific calculations at a very low level, and you might be able to speed those up like two x. And if you could speed that up by two x, wow, you've just got a huge, huge win, right? Quantization is definitely a huge factor, and FP four is supported on, on the, uh, Blackwell generation, but we don't have a lot of research or, or, uh, feedback from customers on that one yet.
Like, it's like, yes, I wanna do it. No, I'm not sure whether, you know, it's gonna, you know, be good enough. Am I right in thinking though, that, like from a guy courier, futureum group, I, I would think to the point of Keith's question, there's maybe three main factors here.
Um, uh, one is, um, uh, I think what you refer to, which is, you know, sometimes called availability. I think of it as accessibility. Is it in the right region with the right proximity, uh, to your data?
Or can you get your data there or whatever it is. So that's one factor. Um, uh, a second factor is, you know, what, what performance, you know, do you need?
How fast you need to train, what size of mile, what have you? That'll that can dictate, you know, a a a B 200 outta the gate. But then the third one is your pricing.
Google Clouds price it, you just said earlier that, you know, you're getting paraphrasing, you're getting two x the performance, but it's not two x the cost. Um, so where's that coming from? Because it seems to me the H one hundreds that you have, like, you're not gonna burn 'em.
So, uh, but I don't know that, you know, you might price 'em to sell. That could make a big difference in terms of where folks are going. And I'm not, obviously, I'm not asking you to reveal where the prices come from unless you feel like it on this non NDA broadcast, I will reveal my prices by telling you the ones on the website are the ones that I would talk about.
Uh, but more realistically. Uh, but I mean that what, what what I'm saying is your, that that's the third leg of that stool. And it's a, it's a really important one because people are trying to do practical things with the models that they're training themselves.
And those practical things are not always have it in a month. So They, they, they'll wait for something useful and productive that they spend less on and, you know, so yeah. So you're, your, your your, it's like an artificial injection into the, into the system.
So, so whenever something new comes in, you know, the law of free market economics says that the old thing that is slightly worse will get cheaper, just it's more of a question of when, right? And, uh, it could be a while, right? So we were talking, we were talking about B two hundreds are not of not, are, are not large scale, not available everywhere, GB two hundreds, not large scale, not available everywhere.
So there might not, uh, this is I, yes or no, I cannot confirm. There may or may not be a price pressure on the hopper side yet, right? Yes.
They might get cheaper eventually, but the, the reality is there are customers when they sign a long-term contract, right? You sign it for X number of years, whatever, it's one year, three years is what we offer in standard. You signed it at that time with, and you were like happy with that decision, right?
And sometimes for a business, if you made a decision and you're happy and you continue to be happy with it, your migration costs to do anything else, literally anything else, right? Move to a different GPU, move to a different GPU generation, move to a different cloud, move to whatever the risks it means for your business is just not worth it, right? So even if you could save 5% on like your inference cost for, by moving from H two hundreds on Google Cloud to H one hundreds on another provider, five percent's not worth throwing your business away, right?
Like I, if you have like a downtime, if you have a whatever, transition costs, all of that factors in. And obviously I can't speak for every customer, right? But every customer is always balancing those things out as well.
They don't always make the right choice either, right? Like they're, they're, they're trying their best to make a good choice, uh, on what they use. I just think there's a lot of gravity.
It's not just data. There's a lot of gravity towards, uh, something you've invested some in, and there's, there's, there's a lot of sort of market pressure and messaging pressure, bigger models faster, like all this performance, and it sounds all really cool and everything, but if you're making good business decisions out of it, um, you know, it's, it's like asking for a generic instead of a instead of a brand name. Yeah.
So, so, um, so we launched a four, it went GA in mid-March, for example, right? Uh, we had lots of customers coming on board testing out their stuff, right? It's like, I wanna do a one in week POC, I wanna do a two week POC, I wanna do a one one month POC with like eight nodes, right?
Evening, go on, this is good. Go. Okay.
Uh, this should not be too bad. Um, uh, but what happened was, you know, we did have customers come and say, I did not see the performance improvement to justify moving over, right? Or at least not yet, right?
And whether that's, They'll, they'll get there, they'll get there, they wanna see how it's gonna, they wanna take it for a spin, the Right, right? So, so, so they might have seen some percentage, right? And the price perf is actually the same right now.
So they don't wanna like, okay, well I, if it was gonna give a 20% gain immediately with no work, right, they would've taken it. But if it's gonna take them some work to extract that value out, it's like, we'll wait. Right?
We'll wait a little bit. And even those customers, you know, they might say, they might look at that situation and say, that's, that's fine, right? Like, it's not like they were expecting something and then didn't get out of it.
They were like, I'm gonna test out this new thing, right? Is this is not the right time for me to move over. My existing solution works great, right?
Every customer is gonna be in a slightly different situation. Some are gonna be, say like, we had another customer tweet about like how great they were experience they were having on the B two hundreds, how it helped them jump to the leaderboards of their particular metric. And so that one was like gung-ho, right?
They, they wanna get on top of it as soon as possible. Uh, but you know, it's, customers run the gamut. So again, you said you were having the internal debate about what to name above the director level, and I feel like, uh, you know, guy Kimberly, uh, probably at different points in time brought this like, when is switching cost?
You know, when is that juice worth the squeeze? And so there seems like there has to be some type of a, you know, Google cloud advisory thing at the top of it, and then I think that's probably a human right now. Um, but you do probably have the scenarios.
Um, is it, is it switching costs that are really, is that an internal debate here? Or is it that's more of a customer and customer success and account management and sales and pre-sales and post-sales? Where is that conversation around switching costs?
Great. So, so underlying like hardware costs and time to switch over, that's easy calculation. Like we would have calculators internally, our field folks would, would, would be able to like punch that in mm-hmm.
And get Youku number. The problem comes to like, we don't know, or, and we can't know every customer's switching costs on the software stack, right? On moving your data from one data center to another, unless you're telling us it's, you know, however many bytes you have.
So we can't really just go directly to a customer, it's like, switch here and it's gonna save you X and it cost you y and net win or whatever, right? Um, the same way that we can't go to any customer and say, make the switch to Jacks over from PyTorch because it'll sort of get you these value and this value, and it doesn't cost you that much. You will, some customers are like, yes, right?
I wanna be forward looking, I wanna be, you know, uh, on the most cutting edge I want be agnostic to hardware. I'm gonna go do that change. But a lot of them are just gonna say, as you said, it's like, I, what, what I have works.
And so the follow up question is, does, does that domain of knowledge does, is that residing within Google or does Google want to see the partner ecosystem, uh, with competencies helping that client plus Google conversation move in the correct direction? Or is, is it Google's position that's really a Google Cloud discussion with that particular customer? A partner is not required or included in that conversation?
You wanna get me in trouble, don't you? Uh, What, what I would say is what we love our partners, right? Like we can't scale out that far.
We just don't have that many people. We don't have the same expertise that partners do. Yeah.
Um, so in, in the situations where the partners have either the knowledge or the skills or whatever and can make our lives easier and the customer's lives easier, why would we want to take that off, right? Like, why would we want to interrupt that and get, get in the way? That's not a promise.
I, but, you know, uh, we have two minutes left. So, uh, I, we covered most of the slides that, that I really wanted to do. I've been just showing this one because again, my baby, so my favorite, um, so, uh, Forex bandwidth, let's see, non-blocking cluster.
Oh yeah. We, we, we are trying to aim for a large scale. So in our a three high and a three mega systems, we are basically had like, I think it's 700 or, or something GPUs that are in a single non-blocking cluster.
We're trying to get to 10 K. So, you know, we can match the scale of a lot of our competitors. Um, again, this says two x faster training on an LLM, and if you go to some of our blogs and some of our other things, it might have a number that's slightly higher than two x.
Um, it's not gonna be that much, it's not gonna like say five x or anything like that, but, um, a little higher than two x for sure. Uh, the follow on question to, uh, one that Jack asked earlier, um, I, I'm Andy Banton. Uh, so in, uh, unnamed version one and unnamed version two, uh, what have you as a product manager asked for in terms of, uh, performance improvements there?
Oh, wow. Um, I, I mean, and I, I realize that you, you can't answer that question correctly, but you, you can answer, but ballpark what, what order of magnitude? Um, we are, there is a, there's an like an Optum optimal place that you're already trying to chase.
And most of the time that's the partner's specs, right? The partner publishes some specs and they publish some, here is a roof line model for performance of this item against model training against inference, against whatever. And that's usually what we're trying to chase.
Um, under NDA circumstances, those numbers could be talked about a little more, um, and how close we are to them, for example. Um, that's the best I could give you right now though, so, Okay. And another question that I, I have to ask, even though you only have five seconds left.
Yes. Is, uh, what is the power consumption differences between the various different models? Uh, it, it's per GPU, it's like low single digit percentages difference between a four and a four x.
So it's not much. Um, if you look at the NVIDIA specs for the GPUs, that's probably tell you the power consumption difference. Okay.