Let’s take a look at Nutanix Enterprise AI
Ashwini Vasanth presented Nutanix Enterprise AI, which simplifies the complexities of adopting and deploying GenAI models and addresses common customer challenges. The product, launched in November 2024, focuses on providing a curated and validated approach to model selection, deployment, and security. The presentation highlighted the “cold start” problem, acknowledging the overwhelming number of available models and the need for a user-friendly starting point for IT or AI admins.
Nutanix Enterprise AI offers a curated list of validated models through partnerships with Hugging Face and NVIDIA to address these challenges, providing a “small, medium, and large” selection. This approach aims to simplify model selection and ensure reliable operation. Additionally, the platform handles GPU selection, inference engine choices, and security complexities, incorporating dynamic endpoint creation to streamline the deployment process. Key to Nutanix’s offering is the integrated security, where Nutanix security experts perform scans for vulnerabilities, eliminating the need for customers to manage their security efforts.
Beyond the mechanics of model deployment, Vasanth discussed the need for on-premises deployment, choice of environments, and addressing the “shadow IT” problem through centralized resource management and monitoring dashboards. The presentation underscored Nutanix’s strategic move into the AI space, leveraging its existing infrastructure expertise, including its Kubernetes platform, storage solutions, and the core principles of simplifying infrastructure. The company’s approach has evolved from a solutions-based offering to a full-fledged product based on the need for a pre-integrated AI platform.
Presented by Ashwini Vasanth, Principal Product Manager, Nutanix. Recorded live in Santa Clara, California, on April 24, 2025, as part of AI Infrastructure Field Day. Watch the entire presentation at https://techfieldday.com/appearance/nutanix-presents-at-ai-infrastructure-field-day-2/ or https://techfieldday.com/event/aiifd2/ for more information.
Transcript
Hello everyone. My name's Ashwini Vasant, I'm the product manager for Nutanix Enterprise AI that Mike just spoke about. First of all, I wanna say it's, it's just simply amazing to be here.
This is such an accomplished group, so thank you. Uh, I know I'm, I'm not trying to appease you. So, um, just going into, um, what Nutanix Enterprise AI is, um, the reason I wanna kind of talk to this from a customer lens is I wanna be upfront, uh, we released our product in November last year, right?
So, I'm not gonna say it has everything it solves for all the challenges in the world. It's, it's basically a new product, but it is a very opinionated take. And the way our opinion is shaped is essentially through our customer lens.
And so I'm kind of gonna talk about it from point of view. What's the customer talking to us as a challenge, and then how are we addressing it? And then we'll go into a little bit of detail on how we address it.
So these are broadly the themes that have come up. So one of the pieces that has come up is around Cold start. So where do I begin, right?
Um, so I think, as we all know, with generative ai, the starting point kind of shifted in terms of AI consumption. It isn't about, Hey, you need to be an expert in AI to be able to consume ai, but then to be able to consume AI internally, you, you still need a strategy. You still need a mechanism to deploy this, to make it work.
And that's something that many of our customers came to us and said, Hey, there is a cold start problem as an IT or an AI admin. I don't know how to get this thing up and running. Everybody wants to use Chad GPT, right?
Um, and so these are some of the specific challenges that came up. So things like, yeah, sure, there are open source models, which is, which is a great benefit because you now have open source models as a starting point instead of having data and then having to build the model. 6 million models, and by the way, this is three days ago, so it's probably much more now on hugging face, right?
6 million models do I consume, right? Um, and also it's about, of course, I mean, this is a group which is of course, very aware of all the security concerns. You don't know that all of them are legitimate.
You have no idea how to pick it. How are you gonna curate it? So what we do in product, or the way we decided to answer this question is we decided we needed a two-fold approach.
6 million models and validate all of them, like figure out if it's legit or not. Not it's, it's a losing battle, right? Uh, but we actually decided to partner.
So we have two partners. We have hugging face as a partner, and we have Nvidia as a partner. But the way we partner is different.
We actually say we have a deep integration. So we work with our partners to identify what are these models, you know, what are the models that are having the highest traction? What are the models that you think are most useful?
What are the models that are actually solving for specific niche problems, right? And that's essentially the curation part of it. So we have a curated list of models.
So also don't want, we also don't believe one size fits all because it's like the people who have many, they, they have like H 100 GPUs. There are people who have CPUs, they have, you have the entire range. And so a curation basically has a small, medium, large approach.
So if you have, if you just want a small model, you're doing a quick POC, you can, you, you'll find something there. And if you want a really good model, we'll have like the latest and greatest out there with like 70 billion and, uh, more, that's the curation aspect. But there's also a validation aspect that comes into picture.
Because to run these models, of course you need to run them on GPUs, but you need to make sure that they run reliably. Like it's, it, it's very different running it in A POC where you just say, Hey, I have this up and running for this particular situation, right? And the moment you increase the token length, you just start sending more data, it just crashes.
We can't have that happen on our validated set. So that's the validation piece that we actually do. The second thing that's come up, I'm, I'm sure, I mean, everyone in the industry is talking about agents now.
So it was, it was all about chatbots. It was about rag, and now it's about agents, but it's about how do you get all the pieces that you need to build these agents, right? So what, what are all these different model types?
So we kind of make sure that we support all of the types of models that you need to build an agent. So there is embedding and re-ranking, which of course has been used in the context of rag, but it's also applicable in the agent world because now people want to use it as a tool, right? And then they just say, oh, rags a tool.
And then the agent is more intelligent. So now the agents become, uh, of course they have safety, which is like front and center in terms of guardrail models. So they need to have that incorporated.
They wanna have vision models where you're able to deal with images, uh, as input. And then there is of course, the whole reasoning wave that's taken us over with reasoning models. Now, we, um, we do not officially list, for example, deep seek as a validated model because it's, we don't have legal clearance to be honest, right?
Um, and so we do, however, support customers who wanna run deep seek with our product, with the same level of ease. So we do something called a custom upload. So you can actually custom upload a model.
You can custom upload fine tune models. So this is for the power user who wants to go in test's, the, um, the greatest model out there, The other piece of it. So now we figured out the model piece, but the next challenge that comes up is, Hey, I wanna end point, I actually need to be able to inference this.
I need to be able to give this to a developer and say, Hey, go write your application against it. And to do that, I need to do a bunch of things, right? I need to first identify what's the GPO type that I need, how many GPUs do I need?
And then what's the BCPU and the memory consumption that's gonna be required to run this reliably. And then inferencing engines, I mean, we have a bunch out there. So, uh, depending on which inferencing engine you pick, your requirements are basically gonna change to run this, um, to run the endpoint.
And then to make things more complicated, the model doesn't necessarily fit in memory for A GPU, right? So you have like the whole tensor parallelism that you have to deal with. So for somebody who's starting out to actually deal with all of these challenges and build it out is, is a huge task, right?
So, um, so what we do is we do it on their behalf, right? So we hide that complexity and we just say, okay, so you pick one of our validated models and you'll see a bunch of this is auto-populated for you, right? So you have a great starting point, but of course, we honor the fact that there are people who wanna customize this even further.
So if they do, did wanna play around with any of this that's still accepted. Um, in terms of inference engines, we support TGI from hugging face, and we support VLLM, which is an really popular open source project, of course. And then we also support NVIDIA's, uh, TRT engine.
Now, the other interesting, um, thing over here is this is what is done dynamically. So every time you wanna bring up an endpoint with a new model, you would have to do this, right? And it's all taken care of, it's dynamic.
All of that is great. But before this, you probably need to buy the GPUs just based on figuring out, okay, these are the models, like which GPUs do I buy and how many do I buy? Um, so that one, we actually have guidance where we actually tell people upfront in our install guides what they should do and what they can expect.
Now, this section, we are going to dive much deeper into this. We're gonna have Jesse kind of go deep on the sizing later on in this presentation. But, um, this is more of an overview.
The other thing that we've heard is, apart from the models, just the inferencing engine and all the pieces that you need to get a model up and running, they could have security vulnerabilities, right? Especially when you're piecing together a bunch of these open source components. So what do you do about that?
You can have an in-house security expert, you can do all of these different scannings, you can make sure that it passes that. And of course, it's not a one-time thing. Every time a library is upgraded, you'd have to go through this process, right?
Um, what happens with an ai, with our product, with, um, the Enter enterprise AI product is we do all of this because we have an in-house security expert. We do the scanning for every release. And so this is kind of baked in and taken care of.
Now, this is the problem of velocity. Uh, I mean, every, every other day there's a new model out there. And I think what we've heard is there is a section of an, of the organization which is interested in adopting the latest and greatest, I'm sorry.
So Kimberly based term, on the prior slide that you had up here, I got confused. You're use, you internally are using the conference tree scanning, the harbor registry scanning, and the black duck scanning on the libraries that you're recommending that they use and do the scanning or, or I I got confused. Yeah.
I'm sorry. What's being done there? Yeah, Let me, let me actually clarify.
Um, so I think, I'm trying to say ours is actually a product. So we don't expect people to go and download libraries and install them. Okay?
Instead, we have a full fledged a turnkey product. So when you actually install Nutanix Enterprise ai, everything's taken care of behind the scenes. Now, the product itself, because it's an enterprise product, right?
It comes from Nutanix with support with all our legal and security clearance, you're not worried about these things. Okay? So you're, you've, you're, you've put these names up here to say, this is what we've already done for you.
Yes, that's Got it. That's right. Yeah.
Can you, can you explain a bit on what you mean by security clearance? I'm not really clear on what you're trying to say here. Sure.
So, um, so essentially every time you, uh, release any enterprise product, you would need to make sure that you have, at least most enterprises have this process where you need to do some kind of vulnerability analysis. You need to make sure that there are no, uh, CVD CVEs or defects. You also need to do a bunch of other analysis with the code itself.
Um, and so there is a very standard process before you can certify something as worthy of a release, right? Uh, and this is a process which is not baked into an open source project, right? Because for, for obvious reasons.
So, uh, every time somebody consumes directly from open source versus consuming from an enterprise, that's the additional effort that they would need to put in to make it happen. Got it. So, so, so it's your own process to validate and provide a commercial solution.
It's not a security clearance in the sense of some extra, you know, requirements for like NISD validations or DOD uh, certification for military and government stuff, Right? This is not that. This is basically just talking about our product certification.
Thank you. Thank You. Uh, so coming, so coming back to the, uh, part of the organization that is ahead of the rest of the organization, right?
And wants to experiment with, let's say, deep Sea a LAMA four, which just got released, there is a little bit of a chicken and egg problem because you can either be really an early adopter, take it as soon as it's there, or you can be like, let me actually test this validated, make it production grade. And so we, we wanna make this choice a little simpler. So what we say is, we'll give you a mechanism where the admin is in charge, can still monitor and observe, but then give people, you know, we, we have role-based access control.
So you can basically invite a user and say, Hey, go ahead, deploy the latest and greatest and test it out, but you know what, I'm never gonna call that production grade. Hmm, right? So do your POC and then it'll get baked in.
And then our validated models, on the other hand, can be used for production without actually, uh, worrying about the consequences. The next thing that I've, uh, we've heard is basically around safety. Now, of course, this has so many different connotations, uh, but I think the way we've heard it is people are concerned over not having control, um, now in different ways.
So if you have a hosted provider, then you have absolutely no control because you, you basically, uh, at the mercy of their, you know, uh, whatever they decide to do with your data. And essentially there is no way to verify whether your data is being used to train another model or anything of that sort. Um, and then there's the, uh, next level where you could, you could always run into a situation where you use a hosted provider and then the policies change and it isn't transparent enough, and then you're left with no choice.
So that's where I think we've heard a lot of asks around having an on-prem deployment, which we do support. So Nutanix Enterprise AI is on-prem, it can also run on the cloud. So we, so we believe in choice.
So we believe in, you know, whatever works for a particular use case, and an organization will cater to that. But with on-prem, we also have the ability to cater to air gapped environments. So for people who are in very highly regulated industries, the air gap deployment has been really critical.
So we support that too. And then there's, uh, the shadow IT problem that we've heard. And, and I think this comes down to everybody is eager to adopt an experiment, but then there is no mechanism or oversight around everything that's happening in an organization.
And it's, it's, it's also really hard to track because, um, I think just that typically organizations don't monitor egress traffic. They don't monitor things which are kind of going out, uh, from an AI perspective. And so, um, we basically offer a centralized system.
3, right? And the other person in another team wants to leverage it, but is just not aware of it. 3.
Now you've kind of duplicated the resources, you've duplicated, you know, the effort, all of that stuff, right? So the centralized model helps with that. The second one is more around, once you've downloaded the model, you could have any number of endpoints, but you're not actually reusing you, you're just reusing the model.
So you're not consuming more storage. It's still kind of just that one model that you've downloaded. So that's where the centralized mechanism actually really helps.
And then the next part of it is, of course, the actual monitoring and the dashboards. Um, all of this will become a lot more clear when we walk through our demo, when we actually show you the dashboards. But being able to keep tabs on which endpoints are actually being used, maybe there is an endpoint that is up and running and nobody's using it, then it's kind of a clear indication that the resources are not actually being utilized and you can turn it off.
So, um, that level of centralized monitoring is another thing we bring into the picture. So you might have, yeah, I think you went through it quickly, and you might have started in the beginning with the intro, but how did Nutanix get into this space? Sure.
Obviously it is a group you work with, but how specifically, what was the, uh, what was the event or the realization that got you into putting together the, uh, Nutanix platform for ai? Yeah, tha thank you so much for that question. I should have started with that, to be honest.
But, um, I think a couple of things happened. One is, uh, we've always been a platform where we believe in, uh, I think we have, we say this internally a lot. We wanna make infrastructure invisible, that we wanna just make it easy to consume.
Now, when, um, when the whole generative AI wave came into the picture, we had our first version, which was not a product, it was actually a solution. 0 if, uh, anyone remembers it. But, um, it was just a collection of parts, right?
Uh, and that was our first take. We were like, okay, let's see. You know, we, we don't wanna build something and then realize that nobody wants it.
So we were like, let's start with, you know, let's, let's do the solutions approach first. So that's, did You just aggregate what you had already from other things and sort of put it together Frankenstein style to begin with? No, so, so we actually took, um, so traditionally Nutanix has been, um, we have our own hypervisor layer.
Mm-hmm. And then, so we have the hardware, the hypervisor, and then we, uh, acquired a company to do Kubernetes. So we have a Nutanix Kubernetes platform, which was the cloud native approach.
So the first AI solution approach was, Hey, we'll give you all the pieces that you can run on Kubernetes, but you would do it yourself, right? So essentially, and, and that's where it's like, it was PyTorch base. So you could, you could get all of these open source pieces, you could put them together.
And then, uh, we kinda waited for almost a year and gathered feedback because we were like, let's just hear from customers what, what's wrong, what's missing with the solutions approach? And that kind of actually led us down this path where we were like, okay, so we're hearing these clear signals that people don't want a bag of paths. They don't wanna deal with upgrade issues, they don't wanna deal with security.
There is a huge learning curve involved for somebody who's starting out. And, um, that's essentially why when we introduced this product. So what of nutanix's core capabilities were you able to leverage to put this together?
The idea of the Yeah. All in one appliance. Was that the sort of what design principle, or where did you start?
So, so I think the design principle was around, uh, one is, this is cloud native, so it is basically leveraging our Kubernetes platform. Mm-hmm. But we are a storage company, or we did start off as a storage company.
Um, so that's still a very strong pillar of support for us. So everything that we do in Nutanix Enterprise AI uses Nutanix files and objects, so for the model storage piece, and then, uh, for, for all the container storage, all of that. And so it's kind of built on the premise that we leverage the rest of the platform.
Okay. So, um, so, so just in terms of like the full stack story, Nutanix offers compute, storage, networking, so this is like a layer built on top of it. So we're leveraging, That's very helpful to connect the dots of how you got there and why, why you know what you're doing.
So thanks. Thank you so much.