Secure and optimize AI and ML workloads with the Cross-Cloud Network with Google Cloud
Vaibhav Katkade, a Product Manager at Google Cloud Networking, presented on infrastructure enhancements in cloud networking for secure, optimized AI/ML workloads. Focusing on the lifecycle of AI/ML, encompassing training, fine-tuning, and serving/inference, and the corresponding network imperatives for each stage. Data ingestion relies on fast, secure connectivity to on-premises environments via interconnect and cross-cloud interconnect, facilitating high-speed data transfer. GKE clusters now support up to 65,000 nodes, a significant increase in scale, enabling the training of large models like Gemini. Improvements to cloud load balancing enhance performance, particularly for LLM workloads.
A key component discussed was the GKE inference gateway, which optimizes LLM serving. It leverages inference metrics from model servers like VLM, Triton, Dynamo, and Google’s Jetstream to perform load balancing based on KV cache utilization, improving performance. The gateway also supports autoscaling based on model server metrics, dynamically adjusting compute allocation based on request load and GPU utilization. Additionally, it enables multiplexing and loading multiple model use cases on a single base model using LoRa fine-tuned adapters, increasing model serving density and efficient use of accelerators. The gateway supports multi-region capacity chasing and integration with security tools like Google’s Model Armor, Palo Alto Networks, and NVIDIA’s Nemo Guardrails.
Katkade also covered key considerations for running inference at scale on Kubernetes. One significant challenge addressed is the constrained availability of GPU/TPU capacity across regions. Google’s solution allows routing to regions with available capacity through a single inference gateway, streamlining operations and improving capacity utilization. Platform and infrastructure teams gain centralized control and consistent baseline coverage across all models by integrating AI security tools directly at the gateway level. Further discussion included load balancing optimization based on KV cache utilization, achieving up to 60% lower latency and 40% higher throughput. The gateway supports model name-based routing and prioritization, compliant with the OpenAI API spec, and allows for different autoscaling thresholds for production and development workloads.
Presented by Vaibhav Katkade, Product Manager, Google Cloud Networking , Google Cloud. Recorded live in Santa Clara, California, on April 22, 2025, as part of AI Infrastructure Field Day. Watch the entire presentation at https://techfieldday.com/appearance/google-cloud-presents-at-ai-infrastructure-field-day-2/ or https://techfieldday.com/event/aiifd2/ for more information.
Transcript
Yeah, so, hi everyone. Uh, my name is Vibe h I'm a product manager in, uh, Google Cloud's, uh, networking product management team. Uh, been working on infrastructure enhancements and cloud networking for secure optimized, um, yeah, ML workloads.
All right, so we'll jump a little bit into, uh, that for, alright, so just laying down the, uh, lifecycle of, uh, AI ml. You've got your, uh, training workloads, uh, fine tuning, and then finally serving an inference, right? And if you look at it from a network perspective, I just wanted to like map it out to, uh, what are the imperatives for each one of them, right?
From your data ingestion side of the house. Um, especially as, uh, there's a tremendous amount of data still setting on premise. Um, you know, your fast secure connectivity to your on premises is critical, especially for, you know, training and fine tuning jobs.
Then when it comes to, uh, your training jobs, and this is where we've made a number of enhancements in the GK networking stack, uh, enhancing the scale as well as the performance. Uh, and, and that's again, contributing towards, uh, accelerating your training jobs. And then we'll proceed to a bunch of innovations we've made on the inference side of the house, uh, primarily also on gk.
Okay? So on data ingestion, we have, uh, interconnect to, for your secure fast connectivity to your on-premise environment, as well as cross-cloud interconnect for, uh, you know, direct connectivity into third party clouds. So wherever your data is residing, if you wanna like, uh, train your models, fine tune your models, leveraging that data, that sort of, uh, interconnect facilitates that high speed rate transfer movement.
Next is on top, uh, coming to your GK clusters. Uh, the GK cluster scale now is up to 65,000, uh, nodes in a single GK cluster. Uh, effectively almost like think five to 10 times what you have.
This an open source, and this is the same kind of infrastructure that is being used to train this Gemini models or even some of the leading frontier models, um, in supporting that sort of scale. And then finally, uh, we'll talk a bunch of improvements that we've done in the, uh, cloud load balancing side of the house to deliver improved performance, especially for LLM workloads. So cross cloud interconnect and interconnect.
Uh, we are today at a hundred gigs, but, uh, soon going up to 400 gigs. Um, and also we now are able to provide application awareness so that you can even apply some sort of priority to the type of workloads that need to get, uh, classified as high priority and ensure that you've got, uh, the fast data movement between, uh, on-premise and cloud. On on GKE.
Uh, we now have a purpose-built RDMA VPCs. This is, um, you know, essentially using NVIDIA's reference architecture embedded within our, uh, Google Cloud network itself. 2 terabytes of non-blocking GPU to GPU, uh, connectivity.
Uh, this is valuable not only for training jobs, and again, you talking about scale of training, speak about the 65,000 nodes that we can support. Uh, so this is not just valuable from a training perspective, but even some of the large models, right? Some of the deep seek or even LAMA falls, like some of these, um, uh, 600, uh, billion parameter models are actually running across multiple nodes.
And even when you're running an inference task, uh, the sort of, uh, fast RDMA, uh, VPC connectivity improves performance. So now coming over to, uh, so that was like kind of the, uh, some of the infrastructure benefits that we had across from, uh, you know, just both training and inference. But in particular, when it comes to inference, there's a few unique challenges, uh, when customers are trying to run inference on scale, especially on Kubernetes infrastructure, right?
So the first one is that, uh, your accelerator, G-P-U-T-P-U capacity today is still constrained across various regions, right? And so customers are often scrambling to, uh, obtain G-P-U-T-P capacity from different regions, different providers wherever they can. Second one is how you distribute traffic across your nodes, across your pods is still an important consideration.
And your traditional techniques don't exactly, uh, are not exactly fine tuned towards the workload profiles of your lms. And that's another area leading to, uh, uneven traffic distribution that again, as a result impacts performance. And then finally, you know, how much of compute do you allocate for a certain model?
Uh, it's a very dynamic, uh, problem, right? Essentially it depends on the volume of requests that you're getting, the size of the model, the underlying GPU, the types of GPUs that you are, as well as what your latency objectives are for that workload. And so allocating that amount of compute is still an, uh, kind of a uphill task for our platform operators that we are looking to address.
So we are looking at all of these problems and looking at it from a GK gateway perspective, and we have enhanced GK gateway and the underlying load balancing capabilities as, uh, therein adapting them towards some of the unique LLM workload characteristics and addressing some of these problems. So with that, we have a GK inference gateway. This is your same GK gateway, but now optimized using certain extensions for optimized LLM serving.
So what do we have? The first thing is we are load balancing rather than just doing it based of, you know, round robin or number of request pending, we actually look at inference metrics coming from model servers such as your VMs or Triton, uh, or even Dynamo coming up, or even Google's own JetStream. And we've done a whole bunch of benchmarking and we found out was that instead of just doing round robin or looking at just the number of requests pending or number of sessions, what actually gives you the optimal performance is looking at your KB cache utilization of your model servers and your GPUs, essentially your inference process.
The decode process is often limited by, you know, how well you can reuse your KB cache. And if you're able to load balance using those metrics, it can actually give you, uh, improved performance. And we'll go a little bit deeper into that.
Not only do we load bands, but we can actually even autoscale based on, uh, you know, these model server metrics. So going back to the problem of how much compute do you allocate, you know, being able to autoscale based on a certain threshold can actually dynamically allow you based on the request load as well as, you know, sort of your utilization of GPU infrastructure, uh, and make much better use of your, um, GPU infrastructure efficiently scale without impacting performance. Secondly, uh, we spoke about how your G PTUs, uh, is, are still constrained.
And so typically, uh, what customers end up doing is that they deploy multiple instances of models, uh, with on dedicated G-P-U-T-P-U infrastructure. However, what they can also do is use a common base model, but use these LARA finetune adapters to actually multiplex and load multiple model use cases on a single base model on a single G-P-U-T-P-U uh, accelerator. So a practical use case of that, right?
Imagine, uh, you have a global customer, you're serving models in multiple different languages for multiple different, uh, localities, right? Let's say you've got like German, Spanish, English. So typically if you've got, uh, fine tuned models for each one of them, you would have to, let's say deploy three of you would require, let's say three GPU or TPUs.
But now imagine that you've got a single base model. Let's say it's a LAMA for, you know, Gemma three model, but then I've got these fine tuned adapters for each language model. I can actually, you know, better bend back three of these, uh, models on a single GPU and thereby make more efficient music.
My infrastructure and our load balancing algorithms are now able to, you know, actually recognize the fine tune models and actually route to them with a certain level of priority as well. And this essentially helps you increase the density of your model serving deployment and make more efficient use of your, uh, accelerators. The other thing also is, uh, multi-region capacity chasing.
So oftentimes, um, you know, you might have a surge of load in your primary region or your primary region might be out of axi capacity for whatever reason. In that case, rather than have multiple dedicated clusters per region and have to have, uh, to individually manage these regional stacks and the low distribution between them, that again, is not a very effective use of your operations or of your capacity. You could have the single inference gateway that can route to regions wherever the capacity is available in case, uh, you know, your primary region is out of capacity.
So this is another way that you can have, uh, you know, more efficient, uh, use of your capacity and also use it across multiple Google, uh, pull it from multiple Google Cloud regions and use it as, uh, as part of the single gateway. And then finally, we've seen that, you know, these, uh, LLMs, uh, they themselves are sort of an attack surface, right? They themselves can be jailbroken.
You can have pro ejection techniques and also, uh, safety and security guardrails, right? So for example, how do you, uh, not serve harmful content? How do you have certain moderation policies and uh, you know, also like just the observability around it.
And so with that, uh, we have integrated, uh, these sort of AI security, uh, tools, including Google's own model armor. We are partnering also with leading providers like Palo Alto Networks and NVIDIA's Nemo guardrails to embed these AI safety security tools, right at the gateway itself. So traditionally, your developers would have to embed these as part of their code, and that obviously is fragmented approach, uh, has its own challenges from a central security operations standpoint.
But by integrating it right at the gateway, the platform and infrastructure teams now have centralized control over this and also provides a consistent level of baseline coverage across all models, uh, running in your cluster, in your environment. Right? So that was, um, that I just wanna take a quick pause if there's a, a round of Yeah.
Questions. Yeah. Jim Rinky, CDC, uh, so basically you'd be able to intercept questionable internet traffic at this point by, by doing it this way, for example.
Yes. Uh, because I'm looking at the, uh, model armor sounds quite fascinating. Yeah, I'm not, it's Already mentioned that.
Or can you elucidate or is it coming a little bit more on that, how it works? Yeah, So we are, uh, screening for re requests as well as responses. Okay.
And it is look, and we've got different categories of filters that you can apply at different levels of sensitivities. Okay. So there's, you know, filters around, uh, you know, let's say racial content or, you know, hate speech and you know, so on and so forth.
And you can also define like the level of sensitivity against which you want to be able to apply each one of them. Okay? And you can also do it on a poor model or a poor use case basis.
So you, let's say I've got a Gemma workload, um, or I've got versus a LAMA workload, I can, you know, on a have, like on a per model basis, I can specify what model, uh, policies I wanna apply. That could be even interesting intra language or between different languages, right? Something that could be offensive in one language Yes.
Or vice versa, right? Yeah, that's right. Yeah.
There's different, uh, aspects of uh, language support as well. Uh, yeah, wouldn't have to look at, you know, how those are specifically supported per language basis, but yes, there's a language to it too. Wow.
That would be impressive. It sort of auto, uh, overkill to do like armor versus some of the open source guardrail tools. I mean, I'd rather use those because they're emerging faster instead of using like a Palo Alto.
Yeah. And then the second question is, do you, does the gateway provide, um, authentication and authorization? 'cause some of the other providers, the customers I'm working with, that's a key thing they wanna hook up like Okta to the gateway and then that's the only way you get in within the organization.
So is that sort of built into this as well? Sure. Yeah.
So, and first question. So our integrations are actually open in the sense that you can actually take one of your open source favorite open source tools and it is a load balance service extension. And so, you know, just by writing a little bit of, uh, go code, you can actually, uh, integrate as part of, uh, this service extension call out.
And so, yeah, it's not, um, limited to these providers and we are looking at, uh, more contributions to the ecosystem. Um, on your second question, um, sorry, I just lost that. Could you repeat that please?
You're gonna make me remember it now. Oh, yeah. Does the, um, does the gateway include like in integration to, so like Okta for authentication authorization?
Yeah, So one of the ways that, uh, so there's two techniques that we've heard quite a lot of interest in. One is OAuth and the second one is using API keys. Uh, especially from a user facing perspective.
Yes, we've heard requests for OAuth flows and then, you know, oftentimes like lot of the, uh, open providers, uh, managed providers are using API keys to authenticate to models, and that's kind of request. So that's one of the integrations that we have on our roadmap coming Okay. Is to be able to support, you know, whether it's OAuth, whether its jaw token or API Keys, I mean clients on OAuth and yeah.
And the other thing is that they're having to build their own with the other providers. In other words, they're having to create their own front end gateway, right? And then do the hack to do that.
So it'd just be nice if you already have the gateway. Yes, exactly. So, you know, our idea is that we just kind of, uh, federate the authentication to one of the existing IDPs providers and the existing mechanisms.
Yeah. Yeah. So, you know, this is like something that we have like use case for deploying both for internal workloads as well as for like if you want to like, uh, if you're a model as a service provider and you want avail of all of these benefits, uh, and optimize serving and more efficient use, if your infrastructure, even you as an inference service provider can use this and uh, expose your models out to the public internet.
Alright? Yeah. So, you know, just going a little bit deeper into, uh, the optimized load balancing that we have, right?
So essentially, um, think about your G-P-U-T-P-U infrastructure as, uh, you know, think about almost as cattle versus pets, right? Uh, in your traditional workloads, uh, your web service requests, think about them as cattle. You could simply like spray them best effort round robin.
Uh, but with respect to your inference requests for these gen AI workloads, every single gen AI workload request is like almost like 10 to six, uh, you know, six orders of magnitude as computationally intensive as a traditional web serving or even traditional infant request. And how do you as a result, land that every single request on the least loaded GPU has a very big impact on the efficiency as well as, you know, not just for your own workload, but even as, as an aggregate, right? And so, because also each of these lms uh, requests could have such a highly variable processing times in the order of like seconds to minutes was a strategy request that completing the order of milliseconds, if you simply go with round robin load balancing, you might end up on a GPU that is highly overly loaded and that affects your serving latency, uh, throughput and the effective end user experience, right?
Especially if you're using like a chat bot that where the user experience really matters. And so by looking at, you know, which is the least loaded GPU, not just in terms of the GP utilization, but actually the KB cash utilization, right? Which is the most computationally intensive aspect in an LLM, um, you know, inference decode process.
Actually looking at that particular metric and looking at the least loaded model server to route to actually helps you, uh, get the best, uh, uh, performance in our own benchmarking, uh, we found out, uh, we are able to achieve 60% lower latency and 40% low, uh, uh, uh, higher throughput of inference serving using these, uh, optimized load balancing mechanisms. The optimization is based on, uh, model like the, the, the G-P-U-T-P utilization by model. Yes.
Right? Yes. What about, what about in, you know, in situ like metrics?
Because you know, something like, something like more than 50%, maybe 60, 70% of models used in a lot of organizations are trials or like there's, there's a great, it could be the same model exactly. Used for different applications and have much more variable usage patterns and also different priority levels. Uh, you know, high latency may not matter for a lot of lab experiments.
Yeah. Or I mean lab experiments on the model, not necessarily scientific. Is that accounted for in this?
Yeah, absolutely. So what we're able to do is like, you know, every single model deployment, you can have it assigned a unique model name. Uh, according to the, and this is again, all based on the open AI API spec, and the gateway actually now has been enhanced to actually route looking at the model name in the body of the requests and based on the name of the model, uh, which again could be mapped into like, hey, you know, this is a production model for a certain use case was a development trial even for the same sort of base model, right?
You've got like LAMA four for dev test versus LAMA for like some production chat bot, you could name it differently. And the client request that comes in, we look at the body of the request and based on that we can prioritize it differently and we can also like, um, autoscale it independently. That's good.
Yeah, That's good. And so yeah, there's a serving priority element, there's a model name of aware routing that is also compliant with the open AI a p spec. That's a key enhancement we've made.
And then we've got the autoscaling element and of course, you know, uh, underlying whatever is the no pool, the c no pool, you can assign that differently. Okay, great. Yeah, And this again goes back to your other point is like, you know, you can set different autoscaling thresholds for your production codes with dev test and make sure that, you know, you've got like, uh, let's say your production workloads are always like autoscaling at 80% utilization, uh, to ensure that they never get fully saturated versus you can maybe oversaturate your government workloads a bit more.
The only thing is this is roughly a tagging process. I mean, obviously prioritization, that's, you know, a human Yeah. Quality, but there's not a really a, uh, uh, if someone forgets to, to name a dev test.
Sure. Um, so you're not, you're not capturing live telemetry, let's say, to figure out how it's being used. You're not shifting it that way.
It's based on the Very much so on the deployment infrastructure and Yeah. You know, Okay. But, but still that's keywords There.
Yeah. That's, that's pretty important because the variability in usage, um, has a lot to do with all the experimentation that's going on right now. Exactly.
As opposed to production large scale, you know, which can happen on its own as Well. Exactly. Okay.
Exactly. Yeah. So that's one use case.
The second key we use case we heard is that, you know, customers have got batch workloads or like from a popup queue that is not so latency sensitive, like jobs that can wait on like, let's say a, you know, a certain queue and you know, can, uh, return back to the user offline versus, you know, real time online jobs like chat bots and stuff. Right. And so, you know, you can assign a chat bot like a more, uh, latency sensitive critical priority versus if it's like a reasoning job, an agent or an offline job or a bad job, you can assign it like a standard or a shareable priority.
Yeah. And Mitch, Nick Ashley with fu I'm curious, your, you, your notes here, the infant gateway, is that, is what you're running I Kubernetes the gateway standalone across all of those nodes? Are you mixing any the application code in those nodes as well, or those strictly the infrastructure layer that other nodes?
It's, yeah, it's really strictly the infrastructure. In fact, it's really nothing. But, uh, you know, the gateway, API is the instantiation.
That's how customers instantiate it. But underneath the covers, it's a cloud load balancer, and it's the cloud load balancer routing to your pods and your GK cluster. Okay.
You use the autoscaling reference earlier and, but a prior slide you also indicated the openness integrating. And so you, you had like a Palo Alto logo I saw. Really small, but it was there.
Yeah. And so I, I'm wondering, is there an autoscaling and licensing implication as that autoscale occurs to maintain that secure AI by design, I think is what they would call it, right? Um, is that autoscaling handled by, um, you Yes.
The autoscaling is part of the GK infrastructure, the horizontal pod autoscaler and the load balancer essentially then, you know, autoscales to all of the new pods that are spun up. And uh, yeah. You know, as far as the openness concerned, like that is an existing cloud load balancing feature.
It's called service extension. It's a data plane, uh, extensively pattern, uh, in, actually based off of Envoy, our cloud load bands actually are based off of, uh, open source Envoy. Yeah.
And is the same open source extensibility, service extension data plane extensibility that we are using and integrated as part of Inference Gateway. And so there would be no lift on the customer's part to have that Palo Alto, uh, instantiated and managed and maintained that that would be handled by Google. Uh, no.
So Palo Alto's, the, uh, AI runtime security is still a SaaS service. Oh, and so it's still, yeah, it's, it's a SaaS consumption model, but essentially it's A-G-R-P-C call out to their, uh, SaaS service, uh, the inspection response. Yeah.
Yeah. Yes. 65.
Um, looks like the inference gateway supports, uh, a lot of NVIDIA metrics, uh, from DCGM. Yes. Uh, does it also support, um, metrics from your TPUs as well?
Yes. Uh, so two key model servers that we are supporting, um, Nvidia Triton and V-L-L-M-V-L-L-M, um, most popular open source project probably, uh, in, in the inference space. Uh, and also like, you know, one of the most popular model servers out there.
And so with VLM, we recently introduced, uh, server sup, uh, t uh, TPU support. And so the same extensions, uh, same optimizations are now also applicable to TPUs. And then we are also looking at, so Google also has their own, uh, JetStream model server.
And so we are also looking at supporting JetStream, which will support TPUs. All right. Yeah, and I think, um, you might have seen this slide earlier, uh, but yeah, I was just kind of retreading the point of, you know, essentially, uh, load balancing based on KV cache utilization.
And so on the left hand side, you see like a lot of variability in the utilization across your model servers. Some of them are getting saturated. What that means as a result, as requests are to get queued up, uh, leading to higher latency, versus when we try to do, uh, load balancing based on KB cash utilization, uh, under the same load conditions, we actually see no queing.
And you see that, you know, your response times, your latency remains consistent even in higher load conditions. Yeah. And this is again, just, uh, highlighting how we are able to more densely pack multiple models on common set of g PTUs, and at the same time, how do we also fair share capacity across all of them by applying, you know, the, uh, serving priority we just touched upon earlier.