AI Inferencing Sizing Considerations on Nutanix Enterprise AI
Jesse Gonzales, Staff Solution Architect, offers sizing guidance for AI inferencing based on real-world experience. The presentation focuses on the critical aspect of appropriately sizing AI infrastructure, particularly for inferencing workloads. Gonzales emphasized the need to understand model requirements, GPU device types, and the role of inference engines. He walks the audience through considerations like CPU and memory requirements based on the selected inference engine, and how this directly impacts the resources needed on Kubernetes worker nodes. The discussion also touches on the importance of accounting for administrative overhead and high availability when deploying LLM endpoints, offering a practical guide to managing resources within a Kubernetes cluster.
The presentation highlights the value of the Nutanix Enterprise AI’s pre-validated models, offering recommendations on the specific resources needed to run a model in a production-ready environment. Gonzales discussed the shift in customer focus from proof-of-concept to centralized systems that allow for sharing large models. The discussion also underscores the importance of accounting for factors like planned maintenance and ensuring sufficient capacity for pod migration. Gonzales explained the sizing process, starting with model selection, GPU device identification, and determining GPU count, followed by calculating CPU and memory needs.
Throughout the presentation, Gonzales addresses critical aspects like FinOps and cost management, highlighting the forthcoming integration of metrics for request counts, latency, and eventually, token-based consumption. He addressed questions about the deployment and licensing options for Nutanix Enterprise AI (NAI), offering different scenarios for on-premises, bare metal, and cloud deployments, depending on the customer’s existing infrastructure. Nutanix’s approach revolves around flexibility, supporting various choices in infrastructure, virtualization, and Kubernetes distributions. The presentation demonstrates how the company streamlines AI deployment and management, making it easier for customers to navigate the complexities of AI infrastructure and scale as needed.
Presented by Jesse Gonzales, Staff Solution Architect, Nutanix. Recorded live in Santa Clara, California, on April 24, 2025, as part of AI Infrastructure Field Day. Watch the entire presentation at https://techfieldday.com/appearance/nutanix-presents-at-ai-infrastructure-field-day-2/ or https://techfieldday.com/event/aiifd2/ for more information.
Transcript
When, when we talk about these inferencing endpoint requirements, kind of going back to this, this idea of what's it mean, you know, when a customer is saying that they want to deploy a model, I'm sorry, should I not move over there? Right in your box. Excuse this, I gotta stand over this box.
I apologize. So, um, you know, what, what does it mean to, to, um, to, to deploy a model? You know, obviously if a customer has already given you an idea of what model, uh, they're looking to leverage, chances are you, you have a rough idea on what GPU device type they'll need to actually run that model and how many they would need on each node.
So in this case, again, you know, within our own documentation, we provide reference for exactly the minimum number of GPUs you would need. And, and given a GPU device type, you know, how many GPUs you would need for that particular model. And then the next step, again, a lot of this stuff is already baked into ui, but the idea is that, you know, you don't have a ui, but you have to do some forecasting from a sizing perspective.
So having a little bit of an understanding of this, you know, this level of, of breakdown will help with some of the forecasting and sizing that you're looking to do. Okay. So, so, you know, we have a model.
We have a device type, we have an understanding GPU count, we have an inference engine that will, will most likely, likely leverage. And again, chances are, because we're leveraging a pre-validated model that's from hugging face, and we're looking to leverage an inferencing engine from hugging face, uh, that's, you know, because we've partnered with them, that's our preferred inferencing engine, we then would effectively, uh, be able to help you calculate exactly how much CPU and memory you would need for that particular inference engine to run that model. So in this scenario, you can, you can see that we have two GPUs.
We have inferencing engine type of TGI, and because we have two GPUs, the actual amount of CPU in memory is, is, you know, there's a, there's a, a different multiplication factor for each inference engine type. But in this case, for TGI, the, the minimum amount would normally been four by 16 for a single GPU. But in this case, because we're leveraging two GPUs, we have to double that.
Now, the, the next phase of this becomes, you know, where, where I think a lot of folks kind of struggle is when you're deploying that LLM endpoint, knowing that it is a pod that's running within Kubernetes, the minimum requirements for that particular pod to run is eight CPUs, 24 gigs of ram and two GPUs. That's the single, for a single instance of that LLM endpoint. When it goes to be deployed, the allocatable resources that are available on your Kubernetes worker nodes have to match that.
And naturally, if you have a worker node that only has eight CPU 24 gigs and two GPUs, that is not enough to account for the administrative overhead that that may also be running on top of that cluster. Other pods that may be running demon sets for, let's say flu D for logging and so forth. So this is just the minimum requirements needed to just run that one endpoint, that one instance inside of a single worker node plus, uh, administrative minus, uh, administrative overhead.
And then the number of replicas will tell you if you needed to, to run your LLM endpoint in an high, in, in a high availability fashion. Now you're deploying two LLM instances across two worker nodes that are gonna be distributed across two worker nodes. So this is just kind of give you an idea of some of the different factors that you have to take into account when, when, uh, when deploying your pods inside of a, a Kubernetes cluster.
And this is just some different scenarios that I want to touch on, where we have a, a worker node pool. There's four different, uh, worker nodes. Each one basically has two L 40 s, GPUs, 16 CPUs, and 32 gigs of memory.
Very top. The very first use case, it's just a chat bot application that's integrated with retrieval, augmented generation pipeline. And in this scenario, the customer is looking to leverage 3 1 8 B.
The single, uh, L 40 s, uh, uh, DPU, that's, that is required leveraging TGI minimum requirements is four, memory is 12, and we have a replica count of one. So maybe it's just not that important that that customer is running this, this application with this model in a, in a, in a HA fashion. So from a scheduling perspective, you'll see on this last, uh, node, this kind of visual of where that pod actually ends up being deployed.
And the Kubernetes scheduler is making that determination. Yes, hi, Denise Donahue, and I promise I won't cause this much trouble as before. It just is surprising to see me memory of 32 gig, excuse me, for the last two days we've been hearing about terabytes of memory in little tiny form factors.
And then you're going, yeah, well, two GPUs and 32 gig, no sweat. So that's a typical use case. So It depends on the model.
That's what it all comes down to. And that's actually beautiful. Kind of one of the, the, the value adds, if you will, of n ai.
It's when we have a pre-validated model, we're giving you kind of the recommendation of what would it need, what you would need to run that model single instance in, in, in your, in your platform. And it'd be quote unquote production ready. So you guys have been out just Kimberly based again, um, return group.
You guys have been out with this since November. So by this time you've been configured and talked to customers and everything else about what does a typical configuration look like that's going out the door, or is get getting proposed at this time? I mean, I'm sure you've got a few installations, but a broader view on your sales pipeline and kind of what that looks like and where is it going?
Because it seems like this is a, you, exactly. What you're trying to do is an easy button for somebody that wants to roll out with the chat bot or the copilot kind of thing, or whatever it is. So they don't wanna have to figure all this stuff out.
And so I've got a mid-size enterprise, you know, I'm a billion dollar firm or whatever, and I don't have the money to set up an entire AI factory, if you will. Mm-hmm. Yeah.
Um, I think we're seeing it's, uh, it's definitely shifted over the past year. It started with most people trying to figure out what's the use case that they wanna solve within their organization. And again, when I say what's the use case, it's like, yes, we wanna consume generative ai, but what do we do with it?
Right? Like, how do we get ROI for our organizations from this? So that's where we started.
And at that point it was mostly small clusters. It was like, we wanna build a POC and we wanna prove that there is value, right? And, um, that was basically actually one of the points where, uh, Jesse, basically the, the models that he's showing you, these are all like super small ones.
Mm-hmm. It's like 4 billion, 1 billion, um, not something that you would use for a really, you know, high quality production use case. But I think this is in the phase where people are like, let's build A POC and let's say, let's show that it works.
But then since then, we've seen evolution where people are like, I think it's becoming more and more real where people are like, I don't think this is a hype cycle. So they're like, okay, so this is gonna roll out. And the the shadow IT problem is basically causing people to think about, okay, I think it makes sense to have a centralized system, right?
So the easy button, you're absolutely right. But the centralized system is what is kind of resonating more because they're like, we need that centralized system where we can actually bring up this one big model, but then it's shared. So we're not worried about, oh, to run a 70 billion model, maybe I need two L 40 S's or four L 40 S's, but that's all that I need.
And then everybody leverages that model and that endpoint. That's where we're going to now. So yeah, so as we do get into kind of a larger model use case, you'll see here that Cole Lama 34 B does require a minimum of two GPUs for an L 40 s.
All of this has been calculated for you, but it's also documented so that you can do forecasting and sizing. And then naturally the, the UI will already account for the number of GPUs. And because of that inference engine type, the amount of CPU and memory that you would require, in this particular scenario, they actually want a high availability.
So they have two replicas. So you can kind of see really quickly what's happening under the covers. When Kubernetes goes to schedule that guy, he's deploying two different pots, consuming the two GPUs that are available on each node.
And you some, some, you know, if you're familiar with Kubernetes, some, some of this stuff may seem very simple, but it's, it's a, you know, it, it's amazing how many, how many folks are challenged with understanding how the scheduler works, especially when GPUs are in the mix. Mm-hmm. And then the last use case, oh, actually fail over.
The one example that I also wanna cover is the idea that, again, as you're doing your sizing, you're wanting to account for and plus one use cases. So from a scenario where a worker node may have had some planned maintenance and that, that, um, node is going down, you have enough capacity to ensure that those pods can move somewhere else and have the actual DPU available. So again, things that people don't take into account when they're trying to build out a true production ready LLM inferencing endpoint.
Do You warn, uh, Brian Signal 65? Do you warn or prevent overscheduling of GPUs? We are hyper dependent on the Kubernetes environment from so, so as a matter of fact, in this particular EKS cluster, it's a persistent cluster that's running compute nodes only for 90% of the time.
When you guys saw me in the back trying to do some preparation, all I basically was doing was taking two different node pools that were already scaled down to zero and just increasing the number. As those nodes start to come up, the GPU operator detects that those nodes have GPUs on them, and it starts installing the device drivers once those guys are up and AI is of, is ready to start scheduling. Thanks.
And the last use case is really just this idea of leveraging, uh, compute only for the very small model. So again, if you have end users who are just looking to do some development, some PLC testing of these smaller models, the idea that, you know, we're accounting for, uh, kind of later generation, uh, CPU models and taking advantage of them so that we can effectively run those models in, in an efficient way so that you can actually see some, some higher degree of performance, lower latency and, and still get some value add out of it. And that's, that's pretty much it from an inferencing perspective that I wanted to cover, just to make sure everyone's is, is aware of this.
So there are obviously costs associated with all these. Would I be able to get out the back end of this? Uh, you're gonna implement what kind of calculations of Right.
You're gonna, you're gonna use how many, how much memory, how much CPU? Yes. Yeah.
Right. Yeah. You know, because again, if you're a citizen data scientist or someone like that, that's new to this and you have no idea how expensive this could be.
Yeah. You know, if you were, if you were, uh, you know, whatever you're gonna be achi attempting to achieve, right? It could run for quite a bit of time or consume an enormous amount of resources.
Would that be something that you already have in here and we just haven't gotten to yet? Or is that the next phase? I, I, I just wanna say that I think that almost, uh, it's more like a behavioral thing, if you will, when a customer, uh, a lot of times when they wanna run these extra large models mm-hmm.
Um, and going into the documentation, they'll see which GPU device types are supportive. So let's say a, uh, 4 0 5 B model requires, uh, uh, eight H 100, all of a sudden that becomes kind of a, a, a, a factor for the customer to kind of revisit A, their budget B, whether or not they really need that large model. Because yeah, if you're running large models, you know, you want a, a super fast GPU so that you can ensure that you have lower latency.
So, um, a lot of times they end up choosing a smaller model and in many cases, they won't even choose an H 100 for those smaller models. They'll actually go with L 40 s because, you know, less cost prohibitive. And that becomes kind of the target.
And that's, again, going back to my first slide. Now I know my model, I know my GPU device type, I know how many GPUs I need, so now I can start kind of the process of, you know, sizing. Yeah.
Right. I could see somebody who's trying to learn how to become a better prompt engineer, just as an example, right? Yeah.
Or someone that's brand new to that. Yeah. Right.
And is just saying, you know, how big is space versus something very specific. Yeah. Right.
Okay. I'm just curious. No, Uh, uh, I think that's a great question, right?
And, and, and the way Jesse, um, was kind of walking through this, this is the kind of conversation we would have with our customer before they commit to the product. I got it. Right.
Uh, and so, uh, we are having these conversations essentially to prep them and say, okay, you know, everybody does want the latest in the greatest, you probably want LAMA four, but this is what it means, right? And, and this is the calculation where it's like, okay, so you are going to need, you know, H 100, so you are, and this is gonna be the budget, right? Um, and, and so that's kind of the offline piece of, um, the entire calculation.
But then, uh, to your point, I think there is also a, a piece that comes after you bought the GPU, right? Where you're like, um, people are trying to prompt engineer, they're trying to try these new models, and then they wanna know, okay, now to run this, you know, as I start, maybe I start with two users and then I expand to 15 users. Right?
Uh, what does that look like? Like in terms of consumption, right? Yeah.
And, um, that's where I think philosophy, as a philosophy, Nutanix has always about been about start small scale as you find the need. And I think we've carried that forward with AI as well. Yeah.
Okay. Because we are like, uh, we don't wanna force people to have to commit to having 200 users when they start because it is a new field. And to your point, it's like until people have used it to a certain degree, they don't know the, the trade offs in the model.
Right? Right. Uh, I can't name the customer, but I can tell you that, uh, just, um, as Jesse was talking about how some people change their mind after they figure out what it needs to run the model.
Yeah. Some of the customers have actually chosen to run quantized models where bring down the level, right? They, they go to a four bit or an eight bit quantized model.
And interestingly, for some use cases, they find it just works. Yeah. Like it gives them what they need.
So they, they're not worried about using a quantized model. So I was gonna say, I think the, the concept of, um, an easy button, as Kimberly said, AI in a box is really nice. Um, it's maybe not targeted to a highly technical user, somebody who doesn't know Kubernetes, which is great.
There's a whole lot of people that would benefit from this. I see. A good customer base for it, um, is it's not dependent on Nutanix hardware.
Correct? Yes. Okay.
And I think I, we saw you running it in the cloud Yes. As your demo Yes. Assumption that customers would actually be doing the same thing.
Yes. Yeah. I, I can speak to that, to that.
Yes. No, No. So you, you intend this to be an on-premises deployment?
Not, No, we, we believe in choice. Uh, but, but to answer your question about do most customers choose the cloud, the honest answer is completely, Uh, completely different question. Which Oh, sorry.
Much more about is if it's in the cloud or Yeah. Not even not in the cloud. Um, you mentioned something about sort of offline in terms of cost, right?
Because you're basically saying, if you're using us and you're using this particular model mm-hmm. Then you need to have a cluster size like this. Mm-hmm.
Great. Mm-hmm. Now, when you're running this, particularly if you're running it in cloud, do you have any, um, metrics, uh, finops, any of that baked in at the moment?
Mm-hmm. Okay. Got it.
Got it. Um, we don't do finops for the clouds, um, honestly, because we're also cloud agnostic, right? Right.
We're not saying, Hey, run on EKS or a KS, we're like, run on any Kubernetes of your choice. It could be, um, anywhere, hosted anywhere. Uh, but, but, but the piece that I will say is, uh, this one of course is an offline calculation to your point about cluster, but there was something that we do in product mm-hmm.
Which will work even on any cloud on EKS or E Ks. Right? Right.
And the in-product piece is once you've purchased, so from AWS, you've decided to consume some GP. Mm-hmm. Right?
But then you still need to bring up an endpoint mm-hmm. For a model. Right.
And you need to resource that endpoint and you need to either manually calculate that resourcing or you need somebody to do it. Right. Right.
And that's the, that's what the product will, so the Nutanix enterprise in, that's What's the Product? Bake in? Bake in.
Thank you. So it'll be do like, it, it'll basically be autopopulated, right. When you try to do the endpoint page.
Yeah. I think it's just when you get to the realm of the non, the, the, the non platform engineers then having some metrics capability that says basically, um, every time you do this, oh yes. It's kind of run and it's gonna take some significant computation and we've tracked your usage over time.
Oh, okay. Yeah. And you know, you guys, you, we've got an easy button, right?
So you have, you have a non platform engineer who's using this. They will be not technical, they will just have convinced somebody to go spend a hundred thousand dollars on a couple of inferencing hardware stacks that they're now using your solution on. Yep.
Right now money is now in here rolling around big dollars, not little dollars. Yep. Being able to figure out what's going that utilization, all of that information is gonna be very useful to continue, you know, you know, if you're playing around trying to decide which LLM am I gonna use Yes.
To put in my chat bot to make sure it doesn't give away free airplane tickets. Yeah. Right.
Then cost is part of the calculation. ROI is part of the calculation usage, you know, performance, some of that. So yeah.
Is that in there yet? Or, yes. Yes.
Uh, so we have a little bit in there. Uh, if we could go back to the demo we could show you, but, um, that's Alright. But you can have it, we'll, We'll we'll bring up the demo again.
But essentially what we have today is, uh, we, we have a request count and a latency mm-hmm. Calculation, which is baked in per endpoint. Mm-hmm.
But there's also a dashboard view for the admin. Right. So to your point, if there is a particular endpoint, which is basically seeing a lot of consumption, it's evident to the admin.
Okay. Right. Um, because it's also consuming more resources, but it's also showing the counter in terms of usage.
There are pieces which are coming very soon, which actually augment those metrics to actually add, um, the dimension that you're talking about. Not necessarily cost, but just in terms of, we know that the whole paradigm with LLM is now not just about API calls. Right, right.
It's about the whole tokens and the, so, so essentially we'll bring all that in. Uh, it's definitely coming soon and Okay, great. In the roadmap.
Thank you. Thank you. I mean, that's really all, all I had to Be honest.
We had a quick question here that's actually coming from Max. Um, the, uh, question on, is this running on a HV? Yeah, I Can.
It is running on a HVI can answer that. So, uh, again, I think we're confusing everyone because of all the choices that we have. So I'm gonna break it up into, we call, we have a full stack offering, which is Nutanix Enterprise ai, NKP, which, um, Jesse just spoke about at the Kubernetes layer, the A HV, which is our virtualization platform.
Mm-hmm. Nutanix Unified Storage, which is object store and file storage, which Nutanix provides. And, um, and our database offering where we, we, we are, we have the ability to manage databases.
So Postgres, for example, is a vector database, which we have a managed offering of it where NDB is a manager. So this entire set, we bundle it together, we call it GPT in a box. And we're like, this is full stack if you want.
So, uh, so to your question about is it on-prem for a customer who's saying, Hey, I want on-prem, give me everything. Right? This would be everything.
Yeah. It's like everything that you need to get up and running and do your thing. Now, the second kind of customer says, Hey, I want bare metal.
I am very specific. I do not want the virtualization layer. We are saying, okay, we, we trust your choice.
So you go with bare metal, but you can still do our Kubernetes layer and you can do the AI layer with the storage. Right. Or without the storage, like completely their choice.
And then the third kind of customer is what Jesse was talking about where, hey, I want a really quick POC, right? I'm not committing yet to anything. I want a really quick POC, I wanna just do this on the cloud.
Maybe it doesn't use my data as yet. Maybe my data is on-prem, but I just wanna try something. And in that case, you can have just the AI layer without anything.
So did I confuse No, no, not at all. You just, I I just that, that when we started out, I was asking, I asked the question about the difference, But it's running on a HP Well that you have something called Nutanix ai Enterprise ai. Enterprise ai.
And I was just trying to figure out what is that? Yeah, That, that's the exact ai Leo. That's the ai Got it.
The ai Leo, Can I jump on the question from Kimberly on the, 'cause you, you were showing free, you are kind of talking about free use cases. What about the fourth use case, which is you have an enterprise customer, which has a significant Nutanix footprint running whether you know VMware or a HD, and they want to kind of extend that to what these, you know, these features, these capabilities, right? So what do they, do they need to buy, uh, these uh, uh, kind of, uh, GPT in the box solution to run alongside?
Or can they just add whatever nodes with GPUs and run just the software stack on top of that? So, yeah, so if I, if I could give it a shot ba basically the first answer they have to, to, to nail down is which Kubernetes distribution they're gonna run. So if they're running HV VMware on top of Nutanix, um, they, they have to first determine what's my Kubernetes distribution?
And obviously, you know, nkp are, are preferred part of our portfolio, but there's other distributions that also run on top of HHV and VMware. So once you've defined that Kubernetes distribution, at that point, n AI would be deployed as containers on top of that distribution. Yeah.
And the only, the only other factor that, that I cut out is storage. So it becomes a question of am I leveraging Nutanix objects S3 and or Nutanix files? And naturally, if you're leveraging Nutanix files, Nutanix volumes, you need rewrite once or rewrite many, uh, you know, block or, or file storage.
You would leverage the Nutanix CSI driver. And in the case of EKS, if you're leveraging their EBS and EFS storage, you would leverage their CSI driver and so forth for GKE and a KS. So it's, it's kind of a almost, uh, a la carte based on Yeah.
Mm-hmm. Where you are with your infrastructure. Right.
And there is one other use case, which is, I already have an HCI infrastructure. I have tons of storage, I have tons of compute, I need GPUs, I don't need any more storage. You could run compute only nodes.
Maybe they're, they have higher dense number of GPUs per node, and you can effectively run that in its own tier. And your GPU node pool would be deployed onto your compute only nodes that have plenty of GPUs, and you would make that part of your Kubernetes cluster, and now you can deploy your LLM endpoints on your compute only nodes. Once again, I'll, I'll step back in here.
I'm Mike Armi, I'm product marketing manager. And add to what Jesse says. No, that was great.
I think what we're talking about, to answer your question very cleanly is that we have licensing options for N AI that includes all those things. So from that perspective, we align to exactly how the customer setup is, and the case exactly how you mentioned they would probably purchase NAI and NKP on top of their current stack and add that to their current distribution at that point.