VMware Private AI Foundation Capabilities and Features Update from Broadcom
Presented by Justin Murray, Product Marketing Engineer, VCF Division at Broadcom. Recorded live in San Jose, California on January 29, 2025 as part of AI Field Day 6. Watch the entire presentation at https://TechFieldDay.com/appearance/vmware-by-broadcom-presents-at-ai-field-day-6/ or visit https://TechFieldDay.com/event/aifd6/ or https://vmware.com/privateai for more information.
Transcript
Good morning, good afternoon. Thank you for the opportunity to talk to you again. Some faces are familiar to me here.
You've seen me in action before. I'm Justin Murray. I'm a product marketing engineer at Broadcom.
And, uh, my next 25 minutes or so is to get us into the product. So, uh, those who've seen me do this before, please bear with me for questions. Uh, for anybody who's not seen this before, this will be an introduction at a technical level into the VMware private AI foundation with Nvidia product.
Quite a long name. We call it VPay of, and internally, VMware Private AI Foundation with Nvidia is a product that's based on VMware Cloud Foundation. You have to have Cloud Foundation to use this product today.
So I'll talk a little bit about the architecture at a fairly high level, and then we'll go deeper. Um, my, my, uh, successor speaker, Alex nus, is going to get into the journey of adopting this. I'll give you the toolkit.
I'll try to bring the tools together so as you can see how they're used, and Alex will complete that story in the next talk. Um, so I will also answer why VCF for private ai, but I'm going to give you an answer from the context of the data scientists. Why would they be pleased with the outcomes that we can offer with private AI foundation with Nvidia?
And, uh, I'll speak about the deployment options. This is largely a combination of a deep learning vm. A deep learning VM is just a VM that has GPUs attached to it.
It could have virtual, virtual GPUs attached to it, and virtual GPUs are key to vMotion, which you know about in VMware is certainly part of DRS, and really is the magic of being able to move a VM from one host to another. We can do that with GPUs. Um, and finally, to complete the story on, uh, started by, uh, my predecessor Speaker Tasha, on utilization and under utilization.
How do you see that? What are the tools that enable you to see the GPU occupied? The GPU is not occupied.
We're going to do a huge amount of work on, uh, on that in upcoming releases. 1, the currently shipping product in this, in this talk. So the architecture you're possibly familiar with already.
1 is the currently shipping version at this time in January, 2025. Uh, on top of that, we layer a blue layer, which is VMware's, uh, intellectual property. Um, and some of these names have changed since you saw me talk about this last, and I'll explain that, uh, that that really has to do with infrastructure provisioning.
Uh, it's not data science, it's infrastructure provisioning. The green layer has more to do with, uh, data science or of interest to a data scientist. Like what's my model going to run in?
It's going to run in an an, an NVIDIA inference microservice, which runs in the container. So that top green layer, think of it as a set of rather large containers, actually, uh, containers that run inference servers, inference servers from Nvidia or inference servers like VLLM from Berkeley, where open source inference servers are done. And we want to open this up to you to use other inference servers besides the ones we started out with.
So if you were to ask me what is the progression from last year when you presented this to us? First, I would say there are a lot of thing, there are a lot of new things in the set that weren't there at the minimum viable product that we started off with and made, uh, 20, 24 new things, like more understanding of model governance. Model governance is a new model appears on the market.
I'm interested in using it. What's the safest place to do that? Uh, and the safest place to do that in VMware private AI is in a deep learning VM that's in a safe zone that's possibly disconnected from the internet.
And once I've tested that model and run all kinds of, um, uh, test suites against it to guarantee that it's not harmful, where do, where does it go from there? Part of model governance is promoting that model from a simple, straightforward, deep learning VM into a more complex environment that you mentioned, which is Kubernetes, promoting the model to a level where developers can now use it in, in Kubernetes. So that's what we mean essentially by model governance.
When You're, when you say the, you know, in a private place or, or, or set apart. I, I'm hearing almost microsegmentation language when you speak Of governance. Very, very much so.
Thank you for that prompt. Um, the, the VMware Cloud Foundation, as people possibly know, is, uh, a collection of infrastructure tools and platforms. An important part of that is the vSphere hypervisor that everybody knows and everybody loves.
Um, NSX is another part of that, which is our virtualized networking and overlay network on top of the physical network. Uh, NSX gives us distributed firewall, which allows us to take one community of users and isolate them from another community of users. And you may ask, what, what is your story there about sharing GPUs?
Uh, we can do that even with different communities of users. We can share a single physical UGPU across, across those communities, and we can prevent them from seeing each other or interfering with each other's work as well. Just as a, Just as a quick geek, whatever happened to like the O-V-N-O-V-S stuff, did that ever make it into the VCF?
The OVOV, Open V switch, open v the, the whole open source project that was sorter when you bought Naira. Yes. You Bought basically open V switch.
Effectively, we, we productized huge amounts of the Naira acquisition into NSX, Right? Yeah. And then there was, there's the, there was the blip, there was the NSX dash V.
Yeah, there was the NSX dash, I believe T dash T. Now you're delving into our, our history here. We've used the word delve twice now.
So Hey, AM not ai, right? We we're data, we're hip data science people. So, so NS XT is the currently shipping implementation.
Yeah. 1 looked at the manifest as it applies to the NSX component would be the NSX dash T component. SX just nsx.
It's just NI don't have to say dash tt. There is no more B, it's just S people can't hear you. Alright, repeat.
So for those of you listening at home, I'm on mic and one of the things was clarified is we need not say NSX dash, we just say NSX now. NSX. Okay.
NSX Yes. Internally. NS xt, but NSX is the product name.
Yes. You differentiate model gov model change control Model Change control has a, is a big subset of model governance, but model testing, model proving All done on a deep learning VM and isolation All done in a safe place. Because we don't know what this model is exactly.
We don't know how high quality it is. We don't know how big of A GPU it uses. Right.
What's the latest model on the market today? Um, what kind of GPU does it use? My first question and answer to to that is, what is the number of billion of parameters here?
That tells me a lot about the size. So there's a whole bunch of data science work to be done here and testing work to be done before anybody uses this model. Yeah.
It, going back to that same manifesto bill of materials, there'd be like the ray, you know, environment for Ray for him doing his particular model, and then there'd be maybe the J model where it's, I'm doing a similar but a different change to that model. I Understand all that, but the question now becomes, when you create J two or J three, how does that go through process? Now we have two problems, as they say.
Yeah, Yeah, yeah. There's, there's versions of models to be handled. There's brokens of data to be handled.
There's versions of the experiment that you're doing. So how deep are you going in this? How, how much work is the word governance doing in this chart?
I mean, it's a, it's a block. It's a, it's a block. But there, there are going to be a series of things under underneath that.
Okay. Uh, for example, and, and VMs facilitate that because how do I repeat the same experiment over again with different models? I do it in VMs that I've saved off somewhere.
How easy is that to do in bare metal? Almost ridiculously hard. Um, so there are a lot of things that virtualization brings you here to do with re reproducibility of the test environment, re reproducibility of the tools that you tested with the data sets that you tested with, you could do very, very interesting version control as, as in the question, just using VMs.
When you say VMs, you mean a cluster of like containers and stuff like that, right? Are you, yeah. Just to clarify that, Think, think of a VM as our runtime for containers.
Yeah, That's what I wanted to make sure that we were, yeah. Virtual infrastructure within the VMware Cloud Foundation Where in, in physical Kubernetes, you're familiar with machines being nodes. Yeah.
Those nodes are now virtual. That's what I wanted. So when we clear, when you called vm, you were, we were talking about a KU cluster When we we'll provision the Kubernetes cluster later, if I get to my demo, when we provision the Kubernetes cluster, it's a, it's a set of VMs that are capable of writing pods collection.
Just Wanted to be clear. Yeah, Yeah, Yeah. Good, good.
I'm glad because we haven't Seen, like, we haven't seen other acronyms like TKE or other things like that yet, uh, or anything related to like a tan Kubernetes. So I dunno if we're gonna go into that level of detail, but It's, it's effectively that it's effectively a, I don't a work a workload cluster. It, it, it's included in this, uh, VPay dash n it's Included in VPay dash N and the, is it VPay?
Am I saying that correctly? VPay, um, should I not say VMware? VMware private AI foundation with I should expand the whole thing.
Okay. Be totally correct. Yeah.
Yeah. But, uh, we'll accept that The foundation's important. 'cause otherwise it's vpa, right?
Yeah. VMware private AI foundation with Nvidia infrastructure. With Nvidia.
Okay. So yes. Vector database.
Our favorite one is Postgres with PG Vector. There was a lot of discussion about vector databases. We are seeing other vector databases out there.
We're open to using those. Um, Does that, does a ve So just like that similar conversation, we, you know, the version control, uh, does that extend to things like Vector database or is Vector database that would be coming soon where you could do version control related To that Vector database is shipping from us today shipping, whether you can version control your data in it, that's a feature's a whole Nother. Got it.
And, and you can plug and play your own rag or vector solution in there. Or it just, I mean, is there, is there ability to put in, We'll give you, you'll see from me later a curated rag design using a vector database. But I can put in like a MongoDB or a You can, you can put your own vector, Not, I don't have to go specific what people We're not gonna stop you from using, but It'll work in the architecture seamlessly.
Yeah. Okay. The same, the same architecture, the same attitude to models, new models or D models.
Okay. Yeah. As to every other component.
Okay. Good. Uh, self-service automation is us doing all the work for the data scientist and setting up the Kubernetes cluster with all of the tooling such that the data scientist is ready to go in their Jupyter notebook.
We call it. Sometimes we call this GPU as a service. Uh, we want to remove the infrastructure from the data scientist as a concern.
Uh, and then GPU management, I'll get into, it's about seeing what the GPUs are doing, and if you can predict what workload you're going to be doing in a month's time, what model you'll be doing in a month's time, we can reserve that GPU for you ahead of time to guarantee that it's there for you. So that's the high level. Getting into a a little more of the rationale here, what we see happening among data scientists and AI developers, and they're sometimes the same person, sometimes not, um, is everybody's in model selection.
We have a prime example over the weekend of, there's a new model selection to be done now. So this is a almost a, a full-time job for a data scientist, uh, testing new things. Is this model going to work for us in a, in a, in an orderly way?
Uh, that's just part of it. Then there's the setup under underlying that. What's the infrastructure that this model needs?
The question we get at least two times a day is, I'm going to do data science. What GPU should I purchase? Well, tell me what model you're going to use it, what size it is, and how often a request is going to be made to the application using this as you saw in the charts.
So set up, uh, rag, I'm glad you mentioned it, retrieve augmented generation, which you can think of simply as the combination of a database plus a model to get you a truthful answer. So the database has a source of truth, it's supplying some of the answers to the model. The model is massaging those answers into readable text.
And finally they have, we have the application that you have to build on all this. So this is happening across the data science community, but it's not happening as a single thing. It's happening on multiple models, firstly on multiple data sets in multiple departments with different people doing it.
The answer here is VMware Cloud Foundation and VMware private AI foundation with nvidia, where all of these people are operating in their isolated deep learning VMs, their isolated Kubernetes clusters, they're operating independently of each other. Otherwise, this is going to be a management nightmare, right? So we, what we are doing, if you, if you gave me a word for it, is management of the infrastructure underneath AI management is important here.
And this reminds me a lot of what VMware was doing for conventional data center resources back in the day. You know, abso absolute aggregating them, managing them, assigning them to specific tasks as needed is just a lot more frequent and a lot more flexible with AI applications because it's a shorter term commitment, shorter term use case. Absolutely.
But it's still all about that fundamental question of optimizing the use of resources. Yes. How, how much can these people share that are doing different work in different machines?
Uh, how much can they share A GPU? How much can they share the storage? We can provide answers to this.
And is there a self-service story that's differentiated now? So to the earlier point, Steven made, you know, it VMware very familiar to the IT realm, but getting a developer or a scientist, you know, type interface or interact with a VMware product of, of probably a different conversation. Absolute, absolutely.
We want to do that. And that's the setup phase and the rag phase. I'm going to deploy enough Kubernetes to run RAG in my demo at the end here, and I'm not going to ask a data scientist other, any other question other than how much GPU power would you like?
That's essentially the conversation. Tell me what GPU power, they're gonna say four GPUs and maybe use one, but I, I can, I can provision a certain range of GPUs for them automatically. And So the service catalog is consistent with what's available and, uh, you can do your chargebacks if you need.
I think we talked about that earlier. So, Right. So this is, this is, um, what used to be called VRAV realize automation, we call it VCF automation.
Um, it has the capability of imposing charges on people. Um, for the, for the just A note infrastructure is 90% of the challenge, not half the challenge. Thank you.
Thank you. So the left hand side is infrastructure. The right hand side is data science.
We want to take care of the left hand side and also supply the tooling that NVIDIA gives us. And actually we're building some of our own, some, uh, we want to supply the tooling that the data scientist likes on the right hand side. So the data scientist says, I need a Jupyter Notebook with LAMA three 7 billion, and I think I need one GPU for that, but give me the appropriate amount of GPU, Mr.
DevOps person, Mr. IT person. Uh, within 15 minutes they have what they want.
That's the idea. So cloud-like, but really on premises. So, uh, you can deploy one of these manual, uh, manually.
A deep learning vm, as I said, is a regular VM with A GPU attached to it, or a part of a GPU attached to it. A part of A GPU we, we refer to as a virtual G-P-U-A-V-G-P-U. Uh, for those who are familiar with, with VMware's past a virtual CPU allowed many people to share the same CPUA virtual GPU does exactly that as well.
So that's what differentiates the deep learning VM for many other, and we put some tooling in there as well. PyTorch, uh, Jupyter Notebook, we put that in there automatically for you. So deep, deep learning VMs are the starting point in modern gov model governance.
Kubernetes is the sort of endpoint, and the thing that links the two together is harbor the repository for models. I'm sorry, but you do, um, red Hat is doing some really interesting stuff within Instruct Lab, we're helping you sort of create abstractions for PyTorch. So you don't have to be sort of an ML genius.
Are you doing anything like that? You're, you're, you're including PyTorch, but that, that's an interesting narrative going on where they're sort of trying to help it make it easy to do training for the non AI expert. Is there anything going on like that?
Yeah. When, when you get a deep learning from us and at boot time, everything is placed in there, con is placed in there, which is an internal virtualization environment within the VM itself. So there, there are two layers of virtualization happening.
There's the VM level and then there's the con level, and those two are completely compatible with each other. There isn't, uh, yeah, I was just going about like the, the, if you, I dunno if you've seen the struc lab, I Haven't, uh, Touched you should look because they're, they're taking another step where they're literally making pie touch and training commoditized where you don't have to build, you know, all the details of back application. Interesting.
We'll look into that afterwards. Yeah. Thank you for pointing that out.
Um, I Think you alluded earlier to, um, the motion. Do you have that, that use where you could move workload from one to another one? Absolutely.
Absolutely. And that is, that is going like a rocket ship in releases that we're coming up with soon. I I've seen vMotion of a 48 gigabyte GPU attached to a VM in about a second or two.
And can you use that for snapshoting? You know, like You can certainly snapshot a vgpu aware vm. Yes, yes.
So our intention four or five years ago when I started working on GPUs was everything that you can do in regular v VMware, you can do the GPUs, snapshots, vMotion, DRS, everything. So that's a powerful story for, for a data center that runs on VCF today, or even vSphere moving to VCF. So, uh, we provision these manual deep, deep learning VMs for you.
You don't need to clone them from anything. We'll clone it for you. We'll put the tools in there and later on when you want a Kubernetes cluster, we'll, we'll create that from templates and, and clones and put all the tooling into it that you would need, what that looks like.
Uh, and typically you'll do DLVM for prototyping, for testing, Kubernetes for, for production or pre-production staging. Um, so in, oops, in, uh, VCF, there are two categories of things. There's a management domain and a workload domain.
I'm just gonna make one remark here. We'd like you to start off, uh, a clean AI environment with your own workload domain that isolates you from other people being managed by the same VCF and things. So that's, that's all that's being said here.
And on the right hand side, now you're going to see a slide version of the provisioning tool, uh, which is I've decided to deploy a deep learning vm. I'm going to put it in, in a namespace of its own. This was referred to previously as a, a resource boundary, a place where we can isolate certain people from other people.
Um, and I can do that with my Kubernetes cluster as well, and they can coexist. Um, there's no difference to us between the VMs that run a deep learning VM and the VMs that run a Kubernetes cluster. They're, they're both regular VMs.
Some happen to have GPUs assigned to them. Uh, who Cleans up namespace Is that you're also responsible for consolidating namespace, integrating that? Yes.
Any, anything that's new that when You said the vm, like I think VM sprawl, I would also think maybe namespace sprawl. Yeah, na namespace is like, like a resource pool in traditional visa. It, it's a resource pool plus plus, but it's got all kinds of access in it.
But, um, so we, we do look after namespace. We create them in the first place. We destroy them.
When you destroy one of these, uh, deployments as we call them here, a deployment can be a full blown Kubernetes cluster with many VMs in it. Some workers in the, in that node, in that cluster have vgpu, some don't. They could even have pass through GPUs if you are really extreme about it.
Um, and they can have multiple GPUs assigned to one thing if you really need it, or shares in A GPU. It looks like a small TKG there that you've got in the black. Yeah, I shouldn't be using that acronym.
I should be using VKS VMware Kubernetes service, but it's exactly the same idea. Yeah. Thank you for pointing that out.
And we're not just stopped. There. We're also provisioning data services and Provisioning Harbor as a repository for both containers and a repository for models.
So all of this is installed for you at the, uh, VMware private AI foundation installation time. And so That feels like when you did your bit NAMI acquisition, that became the catalog of all the things that you could kind of load as VMs, you know, on demand of, uh, particular flavors of open source though. Okay.
There is a catalog here too. It's, it's a VCF automation catalog. And these things that you're seeing on the left hand side, the AI workstation, the Kubernetes cluster are catalog items within that catalog.
But the, the, the, the great catalog in the sky is harbor to hold all of the, the containers. And that's particularly useful for offline systems, air gap systems, and we'll put models in there as well. We have OCI ways of doing models now.
Mm-hmm. Mm-hmm. Um, so the workstation I showed you last time, I won't demo that again.
Uh, a, a VM comes up, it's got some base things in the image K and PyTorch, and then we insert into it a bunch of containers from Nvidia GPU Cloud. Now we can get containers from elsewhere. We can get them from local copies on Harbor.
It doesn't have to be NGC based containers, it could be others. But the starting point last year was Nvidia, GPU Cloud authenticated containers tested by Nvidia and tested by US containers like the NIM to run a model. And we'll mount a model into that, um, a retriever, which which accesses data in a ver a vector database and presents it to the completion model at the end of rag.
Just, I don't want Con and Python, you my deployable. So is there a difference between the two different, like in other words, you talked about You can customize this image that we give you to take out things if you don't like. Okay.
But you have to do that. So you, you get the base image. So if you just deploy it by default, you're gonna get con in Yes.
In a per deployable image. Yes. Which is not a really good Thing, but there's no reason you can't take our image out of the content library that you've put it into.
Okay. Yeah. Instantiate, I think customers wind up doing that and then you'll have all these Yeah.
Security we, we'd like that. We like, we'd like customers to customize it Yeah. For their own needs product.
And in your classic VCF days, you had the kind of, the, the duality was you had the virtual infrastructure on one side and you had your VDI kind of experience on our other side. Is that a similar motif or with the, the AI workstation, the, the, the primary benefactor of that, the user, the use case is this data analyst. They wanna get something done.
Uh, to your point, like whatever their magic menu is for their project, there is the customization much in the same way. There was golden images or image management with VDI that we're just kind of taking that same approach and, and personalize this, this kind of new customer internal to it. There's a great analogy there.
I hadn't thought of that before. Thank you that It wasn't ever just BDI, it was just basically a catalog of any VMs that you possibly want. BDIs were the ones that you would typically mask deploy.
Yeah. But they, it was always just the library. I think more in terms of the personas of those VDI images, like you're a knowledge worker versus I'm a data worker versus et cetera.
And just interestingly, the whole VGPU movement started in VDI because you being able to share gp Yeah, it started there, but we've moved on from there. This is talking about types of GPUs that now don't do graphical work anymore. They, they, uh, they're pure compute Oriented.
So John, these, these were always just things that were you custom add to your library and you can pull 'em out if you Yeah, I get that. I just, you know, customers will tend to deploy what this is DevOps 1 0 1, right? Like people will build and not customize it for deployment, and then therein starts all the vulnerability opportunities sharing GPUs happens.
I think this is a key aspect. Yeah. This, that's the secret right now.
This is a big topic. I'll give you two minute, two minutes on it. The question we were, I was asked here was tell me how this GPU sharing happens.
Uh, this is, this is, uh, a feature of two NVIDIA drivers, a host level driver, and a guest level driver that we, I install for you. They're, they're installed into the host, uh, by our lifecycle management tooling. They're installed into the guest s of a, of a VM using, um, our own automation tools.
What they present to you is the capability to say, for any one vm, how much of a share of a physical GPU do I need? Let's say you need half of A GPU, uh, that would be called a virtual GPU profile, and that's identified in the NVIDIA driver, and you assign that to your vm. You say, I'm gonna take this virtual GPU profile representing half of a physical H 100 and assign it to my vm.
Uh, there are a bunch of decisions that you can make about tuning that. Should it be time sharing with other sharers of that GPU or should it be completely isolated from the other sharers of the GPU and that you say Completely isolated. So you could have like three workloads running on a physical GPU, but they're physically or at least virtually separate in that environment rather than trying to time slice.
They, if you choose, if you choose a, an option for virtual GPU called meg, they are physically separated on, on the GPU. They're getting different slices of the compute think. So, uh, a quick question.
Does, uh, does this private AI foundation include the bit fusion technology? It does not. We've, um, we've sunsetted that technology.
Okay. Do you have any way of sharing BCPS across the network? Um, there are technologies around that could do that, but today our sharing technology means the VM that's using the GPU is on the same host as the GPU itself.
Okay. Yeah. Okay.
We decided for various reasons not to continue with that remoting. Okay. So we were talking about virtual GPUs and, and Kubernetes, the way in which it works is those two drivers.
Um, and you choose whether you're going to use time sliced or mig, and then you choose your sizing. And all of that is done through a user interface called virtual GPU profile set up. Once, once that's set up.
And let's say you chose mig, you and I are completely isolated from each other on the GPU hardware. We're isolated from each other in separate VMs. We don't see each other, we don't know each other exists.
Uh, time slice is a bit more lenient than that. We could have a noisy neighbor problem, but, uh, MIG is common in inference workloads. Okay.
So that was a quick introduction there. Um, so we're going to load up your Kubernetes cluster with the appropriate containers that you identify. You identify those by telling us what package you want in, in your Nvidia, uh, from your NVIDIA set.
And to do that, you need an NVIDIA license and you purchase that from Nvidia as part of their Nvidia AI enterprise. The green layer of the, of the suite. Okay.
Data services manager is a tool that comes as part of our portfolio for provisioning databases in general, but here it provisions vector databases, it provisions Postgres with PG Vector on it as our first step into vector databases. Last topic, there was a lot of talk earlier about GPUs and their utilization. How do I see it?
Can I ask One Question? Yeah, go ahead. And so with all of these different packages, how are they kept up to date?
Like if a updated is released, how does, is that on for the users to do? Or is there a push down from the platform To Do it? Yeah, um, there, there, there is version control across two different infrastructures here.
There's version control of the NIM layer. Nvidia takes responsibility for that, and you control what version you're using because you're downloading those containers into a, a Harbor repository. Uh, we'll allow you to patch as we always have, vSphere, VCF, all of the tools that we'll supply.
We'll keep those up to date if you, if you do auto life cycling with us. Yep. So there, there is a bit of a combination going on there between the components that you're using from NVIDIA and the components that you're using from us.
And, and so your Postgres, uh, PG vector implementation is a container that in Harbor, as long as you manage harbor, It can be, it can be containerized, you can also run it just in a straight VM as a, a set of processes in the vm. Okay. Alright.
Yeah. So, um, there are at least two flavors of it today, but I've, I've seen it run as a regular container. Yeah, yeah, it would make more sense.
But yeah, So this is a sort of beautiful picture of the console that you see in VCF operations telling you the temperature of each vm. And, and if the temperature is bad, we're gonna fill it up as red. Uh, telling you the core consumption on the GPU and telling you the memory consumption on the GPU memory consumption is probably going to be the most important because that's the premium with these large language models, but not Power.
Uh, we are going to show power as well. We we're, we, we have that in plan. Uh, uh, now what this doesn't tell you is which virtual GPUs are occupying the physical GPUs and that's coming in the vCenter, uh, U user interface as well.
So we're doing a lot in this space to give you more and more detail about what your GPUs are doing. And that's vs I can't see on the wrong cloud. So that's v uh, this, this is, this is, uh, VCF operations.
Okay. Uh, but there will be also screens like this in v in the V vSphere ui. So does that, excuse me, does that mean now we're, we're back to multiple consoles depending on who you are and what you're looking at and Not, not every user is, uh, confident or, or allowed to use the vCenter ui.
So this is, this is for this specialist DevOps person whose job is to look after the data scientist and, uh, make sure they're happy. But yes, you could, you could be using the vSphere ui, the client and vROps together, and they will feed each other data so that you don't have to go to. Right.
But so it's, it's, it's still two separate products. It's not one product with different windows depending on your role or your responsibility. You're right, you're right.
Now they, they are going, they are going to exchange data. So you what you could see in, in one, you can see in the other. But if you want to, you can use two consoles.
Yeah. So Those basic blocks and, and any of those, I thought those VMs that are using those GPUs, that how I read that The, the green blocks are actually GP representations of GPUs, Separate GPUs. I, I see.
Yeah. So on that left hand side, there are, uh, two GPUs on, on the host. Uh, and the, the host is represented as the enclosing panel around the green blocks.
Okay. Or the red blocks. Um, now we're going to make this a lot prettier.
This is a, this is the first view of It. I guess my question was, and maybe you answered it already, what, do you know what VMs are actually using those? I do, I do because I can see the full topology in the pan immediately beneath the colorful pan.
Here, you can see host, it's on which vm it's on which cluster it's in, um, and which VC is managing it. You'll, you'll see the entire topology Data being saved. So that could layer, for example, do a time series analysis on it.
Yes. This data is in a database that's accessible programmatically such that you could do your own analysis on this. Yes.
There's a repository behind this. So I, I'll stop here. Um, So a persona, this is for, from a ops perspective, like a platform engineer.
Yes. Plat, there are many names for this person. Platform engineer, DevOps engineer, Kubernetes manager, uh, data engineer.
Who's serving, serving the data scientist. Any of these titles, uh, or the traditional IT admin that we, we've always helped. So a chat bot, an NVIDIA chat bot, you get this for free with the package.
Um, this is not a production chat bot. This is a very simple one. We submit a question to it and it gives us various hyper parameters like temperature, the amount of randomness the model is expressing.
And it says, I apologize, I can't give the answer to your question. So the benefit of RAG is I can attach, uh, a database to this and switch that database on the database here contains some, some data that I've, uh, previously loaded up, and I'm going to ask the same question again, but this time with the Vector database active behind the scenes. And how does the question look here?
Well, the answer is very directed, very precise because it's, the model has been fed with data from the Vector database that is the source of truth. This is a very tiny example, but this exhibits how, why RAG is a good idea, why RAG is the starting application that many people are choosing here. So what's behind the scenes of that?
We provisioned an AI Kubernetes clusters. This is the catalog that the gentleman here asked about. Uh, we can have any variety of catalog item that we choose to, this is customizable.
We're going to provision a Kubernetes cluster here. We're going to be asked, what's the name of that Kubernetes cluster? What's the GPU power?
That's the key question. Uh, what's your license for using this? GPU, uh, which we fill in the, the blanks here.
Here's my license. Hit submit. And about five to 10 minutes later, depending on how fast your system is, a Kubernetes cluster is available to the end user that's running that application that we just saw, that RAG application that we, we just saw.
So I, I'll speed through this here. Okay. Uh, something like 55 steps were taken in the background to, to provision that, that, so a lot of work going on that you don't have to worry about anymore.
Those are all the steps. Uh, Is that customizable? Like is that sort of a terraform Or Yes, it has, it has a scripting mechanism that you see on the right panel here.
Sorry you didn't see it. But there, there is a whole scripting language behind this that like terraform, like Ansible, okay. That you can configure for yourself.
The end result here is a URL that you see under the application here. If I were to click that URL, I would see exactly the application that I used at the beginning of this. So, um, what we saw here is VMware private AI foundation within Nvidia Deep learning VMs, Kubernetes clusters, virtual GPUs for sharing and management tools for seeing the utilization of your GPUs Maturity question.
Enterprise IT administrators, the data analyst. You're within the enterprise. What if you're a service provider, a managed service provider, a cloud service provider within like the larger VMware community, any of this, uh, fit for that purpose?
Or is that coming soon to a theater near you? It's in, in progress. We have cloud service providers as we term them, very interested in providing AI as a service to their clients.
Exactly. Yeah. Um, people who serve the government, for example.
Mm-hmm. Um, people, people who are by definition making their money out of cloud provide providing, they're very interested in using this and are some are testing it right now, Roger.