Three Reasons Customers Choose VMware Private AI from Broadcom
Summarize this presentation by VMware by Broadcom at AI Field Day 6 based on the following Abstract and Transcript. Write 3 paragraphs with no bullets or headings. Begin the summary with the Abstract text.
Title:
Speaker:
Abstract:
Transcript:
Overview of why customers are choosing VMware Private AI and popular enterprise use cases we are seeing in the field
Presented by Tasha Drew, Director of Product Engineering, VCF Division at Broadcom. Recorded live in San Jose, California on January 29, 2025 as part of AI Field Day 6. Watch the entire presentation at https://TechFieldDay.com/appearance/vmware-by-broadcom-presents-at-ai-field-day-6/ or visit https://TechFieldDay.com/event/aifd6/ or https://vmware.com/privateai for more information.
Transcript
I am Tasha Drew. I lead the engineering team for the AI and advanced services engineering group in the VCF division of Broadcom. So focused on bringing, uh, private AI services and capabilities to the VMware private AI foundation within Nvidia, uh, product suite.
Uh, super excited to kind of take you all through, uh, kinda what I'm here to talk to you about today, which is reasons folks are choosing to use VMware private ai. Um, my approach to this presentation and most of my presentations is, uh, not to scare, uh, our PR team, uh, but I really like to show slides and data from outside of VMware to kind of show you, hey, like this isn't just our data. This is popping up around the industry for many different research labs and other, other industry groups and customers.
Um, so kind of expect that from me today. Um, after my presentation, uh, Justin's gonna take you on a deep dive through how to use VMware private AI foundation and a lot of our features and capabilities. I'm really talking to you about the, the bigger question, which is why are people excited about this, and what are we hearing from customers and seeing in the market?
Um, so I know that this is intended to be really interactive. Uh, feel free to ask questions. Uh, and you know, I'm, I don't need to get through all my slides.
I have a bunch of slides, but I'm super happy to kind of take this in whatever direction folks are finding interesting. Um, so yeah, uh, kicking off. Uh, so the three reasons that I'm gonna dig into today about why we see folks really getting super interested in private AI to begin with is, and the first one's a little spicy, like the first one.
Uh, I, you know, I, I don't know that we're hearing a ton about this in the market, but we are actually seeing it in our data centers under utilization of GPUs. Um, why is that a little spicy? Because right now it's really hard to get GPUs.
So what do you mean? I have teams in my organization who are waiting for a long time to get these very powerful GPUs, and then you're telling me they're underutilized. Uh, you know, why aren't I getting the top, uh, ROI the top return on investment in this infrastructure investment?
What are you seeing and what do you mean by that? Um, you know, today, I think kind of the most common, uh, spot is the GPUs are getting bigger and bigger. The models are getting bigger and bigger.
And what I'm hearing is that as the model gets bigger, I need to get the big new fancy GPU, and then I can run one model on that GPU. And if I need to horizontally scale, go get more, right? Um, so I have some interesting data I'm gonna dig through to kind of show you what we're seeing, uh, about.
Maybe that's a myth. Um, maybe people are under utilizing their GPUs. Uh, the second driver of private AI that we're really seeing right now is cost of AI in the cloud.
Uh, cloud, as you know, kind of we've been hearing about for years and we've been experiencing for years. Cloud is amazing to get started fast. I don't have to order the GPUs right now.
I don't have to have all the hardware racked and stacked in my data center. There are higher level services that I can lock into and start using. And so with AI especially, we're seeing a ton of sandboxing experimentation.
People starting to figure out what is this technology and how might I solve business problems with, with it in the cloud? It's a great way to get started, but then as people approach production, those cloud costs really take off. And then folks start, and then maybe your CFO's on the phone with you and they're saying, Hey, like, why did our cloud spend just astronomically increase?
I would like to understand this. And so cost is a big factor that you can start to intelligently manage by intelligent investment in your private AI cloud, uh, private cloud. So that's something I'll be digging into a little.
And then model governance. Um, model governance. Governance is a huge topic right now as, uh, enterprises are starting to kind of start to understand how to think about models, how to think about generative ai.
They're realizing as soon as I expose and train a model on proprietary ip, that model now needs to be treated with a sensitivity. Um, it is now as sensitive as the data that it's been trained or fine tuned on. So how can I start to have nice regular enterprise grade practices for how to manage the sensitivities of these models?
That sort of case one that we're seeing a lot. And then other cases are just how do I programmatically pull new models in, make sure that I'm doing all of the security scanning, make sure that I'm also aware of all of the ways that model behavior itself needs to be guaranteed. Um, make sure the model weights haven't changed, make sure my model behavior hasn't drifted, um, make sure there's no data poisoning.
So there's just this huge model governance question now, um, as enterprises are starting to understand kind of the risks inherent, um, with this new artifact, uh, that is a big driver of why folks are starting to also be very interested in the private AI solution. Cool. Uh, so under utilization of GPU, I'll just kind of dig into this and super happy to take questions as we go.
So, uh, I wanna start off with all of my slides are from the uc, Berkeley Sky Computing Lab. Um, why, uh, the Prism team at Sky Computing Lab did a, has done a really awesome job of sort of distilling this problem that we're also seeing in the market, and frankly in our own data centers. Uh, and I just really love pulling in, uh, data from other sources, uh, so that you can kind of see like, Hey, this isn't just a VMware storyline.
This is kind of around, uh, academia, around the enterprise, around the market. Folks are starting to realize, Hey, suddenly we need to serve all of these LLMs. Yes, that's important.
It's also expensive. Let's dig into the data and kind of see what we're seeing to support my sort of, uh, my sort of straw man to all of you that GPUs are being underutilized. Uh, so kind of starting off, uh, applications are starting to diversify.
Folks are serving large language models. Uh, often you might start with a more generic large language model and then fine tune it or start to have many different smaller models that might be custom, uh, trained for specific use cases and have much better, much better outcomes for those specific use cases, but be it smaller size, and we need to run these across high-end GPUs, which are expensive. Um, so kind of just setting it up, um, when you dedicate GPUs to individual models, this starts to cause under utilization.
So today's common deployment practice, um, again, we see this in the industry, but this is also just Sky Labs like setting it out. Um, everyone's dedicating a GPU per model, and this leads to significant under utilization because of inference workload patterns. So what are those inference workload patterns?
Let's take a look. So this is a data center that the Prism team, um, was monitoring to kind of prove their, uh, hypothesis that under utilization of A GPU is a problem. And here this data center is running five different models.
You can see that model number one is really popular, it's getting a lot of hits, and then the other models are have, uh, less popularity. So they're not consistently, uh, driving up their c their number of, uh, requests as often. Um, so this kind of is showing what you might see in a standard LLM as a service offering in the backend.
Basically. I have some models that are super popular. They're getting tons of requests all the time.
I have other models that are less popular that are utilizing their infrastructure less. But a big thing to take away from this is even the most popular model has a significantly spiky load. So when we start to look at this, uh, I, I know that I'm coming at this from a certain perspective, but what I see is there are cycles that we can be reclaiming from ev all of the GPUs who are running these loads.
So in this model of dedicating A GPU for per model, we are under utilizing our hardware. So is it fair to say then that the VMware software is acting as a orchestrator for GPU workloads? Yes.
So you, you actually look at the GPUs, what they're doing, and then assigning specific workloads so that the GPUs are fully loaded? Yeah, we're doing the scheduling. Um, Justin's gonna dig into some of the nice ways that we're basically making sure that you are not, uh, programmatically under utilizing the GPUs by being able to say, Hey, like this, if if I have GPUs that are, have memory that I should be saving for a larger model, don't take it up by, by scheduling a smaller model in it.
So smarter scheduling of the model to the underlying hardware, um, a bunch of different ways. The other thing, uh, that I really see kind of from a more VMware perspective is, uh, and whenever I bring this up, I do a lot of executive customer briefings. Um, I call it GPU hoarding, but basically because GPUs are hard to get, um, when teams at an enterprise get their GPUs, they don't wanna share them with anybody else, especially if they're running on bare metal.
They just kind of wanna hide their GPUs in a corner so they can do the job that they got the GPUs for originally. And that means that the, when you have traffic patterns like this, you are by definition under utilizing that expensive hardware because they're afraid to share it. What?
So when you have a platform approach, like the private AI foundation approach, you can basically go to those teams as an IT team and say, look, bring these in under management, and you are going to get first priority. All of your workloads get first priority on your GPUs, but when your GPUs aren't being used, they're then made available to lower priority workloads. And those could potentially even be from their own team.
Maybe their own team isn't sharing well with, for example, these are production workloads. Our data scientists don't have enough GPU to do experimentation. We could open up experimentation, uh, workloads by having more cycle time.
Uh, but the, but really, like, it's basically bringing the teams together so that you can have first crack at the GPUs that you acquired and that your BU paid for, but not have to do that by hiding them in a corner from everybody else. Ashley, you see this more of a problem for inferencing than, than training. I mean, 'cause training is a different game altogether, right?
Totally, Totally. Yeah. So this is really looking at inferencing running the model.
Um, yes, absolutely. A question. Is the control plane for this platform based, like, for example, Kubernetes, or is it sort of an underlay of that?
So the data that we're looking at here, it's Control plane for scheduling the, the hoarding, or that's Yeah. Oh, so what I'm talking about specifically our control plane is going to be, uh, Kubernetes. Yeah.
Okay. That's good. That's what I wanna know.
That's awesome. Yeah. When I look at the first three things that you've, you've carried forth, um, it was not, uh, it was almost deja vu a little bit.
Yeah. So if we went back, say 15 years, and we were having the same conversation around, well, you know, turns out we've been, um, we've been buying too much X 86 server, we are, uh, we're under, we're under utilizing the processing capabilities of those servers. They're, they're, they're not as busy as they should be.
Mm-hmm. Uh, they're server huggers that are, you know, not letting their server talk to the other department. And I think that we have some, some policy inconsistencies around how we're actually utilizing that hardware most effectively.
And if you did the search and replace and then put like cloud and AI and GPU respectively in their places, I feel like we're almost repeating a similar pattern. Yes. Um, is, is that consistent with what you're hearing from customers that you are also sensing some dejavu because it's literally a new technology, but it's falling into the same, uh, pattern?
Yeah. Okay. Yeah, I absolutely am.
Um, I've even had really interesting conversations with, uh, C-level executives who, uh, basically started looking at this, and then they're, they're like, oh, so cloud repatriation like, got it. Right? Like it's, it's all the same bullet points.
Mm-hmm. Um, I think the other day someone said to me, this really just starts to speak to the VMware bread and butter. Mm-hmm.
And it does, right? Like, it's like, look, let's just get, let's make sure we are helping our customers utilize their infrastructure as efficiently as possible. Well, I'm, I'm glad you used the R word.
Uh, there's many to choose from, but you said repatriation. Yeah, because a month ago at Gartner, I-T-I-O-C-S in Las Vegas, uh, one of the distinguished analysts, uh, spoke on stage and said, and I quote, yeah, repatriation is theists effectively. Mm-hmm.
Like, it, it repatriation doesn't exist. Uhhuh. It's, it's, it's, it's, uh, it's an illusion.
Mm-hmm. So there's, there's that side of what people when they, when they hear repatriation, I think that's the broader cloud conversation. Yeah.
I think, if I'm hearing correctly, you're, you did mention there's cloud costs. Yes. But this is really about the AI specific utilization use case of that cloud AI specific utilization of maybe GPUs in a data center as opposed to a general cloud consumption pattern.
Yes. I think that there's all of the patterns, it looks like, uh, I dropped off of here for some reason. All of the patterns that you kind of see traditionally around just where people are running their workloads, um, continue to be true.
Private AI is interesting because when you kind of dig into like the private AI premise, it's AI is really powerful. We have a ton of data that could do amazing things. Uh, when you start to use AI capabilities with it, some of that data is inherently, uh, very intrinsic to your business.
It's ip, I really need to protect it's data that I don't want to have to completely move into some sort of cloud service to start being able to take advantage of some cloud capabilities, right? Like there's data locality, there's, uh, different restrictions of Move is a four letter word you had move, so, yeah. Yeah.
So there's very in, there's like intrinsic reasons why I wanna be able to take advantage of AI and cloud doesn't make sense for me, and how do I just run AI incredibly well in my private cloud that this entire solution really seeks to automate and deliver. Um, but then I also think that there's just, uh, you know, this is just a, was such a rapidly growing space that we all kind of have, uh, been communicated some truisms, like, oh, bigger, bigger model, bigger GPU, no way to optimize from an inference perspective. And now we're starting to get all of this runtime data, and we're like, that's not actually true.
Right? Like, there's a lot of ways that we can start to manage the costs of ai, um, more, more effectively. Yeah.
Now you've, you've, you know, acquired many things over the years Cloud health Dean, one of them. Um, and so I, I don't wanna sight read your, your slides, but at some point, I'm assuming we're gonna talk about this notion of the port of manto of finance. You mentioned finance, but also the operational, so fin ops.
Mm-hmm. Um, you also maybe go into the technology business management. So you have the domain of the traditional IT spend, you know, category.
Then there's this cloud thing. Yes. Then there's the labor component of how many data analysts are tied to this, not just the person that knew they were hugging the server and the data analyst team.
Yeah. But what are the other constituents that are tied labor wise to that IT asset or that cloud expenditure? So is, are you gonna cover any of that today?
'cause I saw it was a second bullet, but I didn't know how in depth we're gonna go. So we do have, uh, some, uh, capabilities included in the private AI solution to help folks start to manage, uh, predict and optimize the costs. Okay.
Um, I'm gonna leave all of that to Justin, uh, to kind of what, what features he's gonna be highlighting today. But those that we're like, basically like managing your GPU spend and being able to make sure that, because it turns into an operational question, even at the hardware level, right? Like we're saying, Hey, there's like cycles I can capture from A GPU here, but if you run that GPU too hot, you're gonna melt it, and then you don't have a GPU anymore.
So it's like there's this sweet spot of performance that you really have to aim for, but then as far as your meta question is like, how am I predicting labor and, uh, getting and, uh, all of the acquisition pieces of like having the infrastructure? Yeah. The totality of the cost I think is really the, like if I was a CIO, I have that perspective and I'm A CFO.
Yeah. It's a more holistic than just what the IT team is up to. Exactly.
Okay. So there's a number of tools and capabilities that we've been making available to folks to use that. Um, and then, uh, I think, I'm not sure if we're digging into that today, but yes.
Okay. Thank you. Yeah.
Cool. Tasha, if, um, was you and I who were speaking the other day, I'll tell you, you know, if, if Broadcom VM were not thinking about this Yeah. I wouldn't be shocked.
That's right. Yeah. Just going back to the, you know, the, the, the early days of why you and virtualization and such, I mean, it's, it is this question of utilization in, in the largest sphere, you know, wow, ai, this, this new paradigm, let's just forget 70 years of computing history, right?
And everything we learned before. So now that we're kind of catching our breath, we can bring all these things forward, right? And, and, and the best practices and the learnings, and apply them now to how we actually get these, these jobs done.
So thanks for, thanks for hauling these out so explicitly. Awesome. So yeah, so I think this one, everyone kind of, you know, immediately, uh, you know, kind of looks at, so here's another reason.
GPUs are significantly underutilized. So we have a long tail, as you saw on the previous slide of low requested models, which means if I'm dedicating a lesser requested model to an entire GPU, there's just gonna be a lot of workload I can recapture there. Uh, and so low ba, low batch size is not gonna, j uh, is not gonna saturate your GPU.
And so here's just kind of an example of the percentage of cold models and the percentage of requests received by the top 20% versus the top, the bottom 5%, uh, on the right hand side. And so you can just see like there's a significant gap in the utilization of the GPUs that we can start to recapture here. Um, and then there's also a long idle period between request bursts.
Um, and so this is another kind of thing that you could see in that original very spiky graph is like sometimes your GPU might be getting hammered, and other times it may be sitting fairly idle and there may be long idle periods between those bursts. So when and models with no requests cause this GPU to just be sitting there idly. So again, like, uh, you know, this is all data from, uh, this research team over at uc, Berkeley.
We have similar data from our own data centers, uh, that we used when we were starting to, uh, basically prioritize a number of the features we've been building, uh, just even two years ago. Uh, but I really like just kinda showing external data because it starts to show there's this pattern across the industry. And when you're serving models and really paying attention to GPU utilization, you can, Hey, Natasha, can you go back to the prior slide?
I'm trying to understand this Yeah. The second chart here. Mm-hmm.
So the percentage on the vertical graph is percentage utilization, percentage of models, percentage of what? Yeah. So if we take this one back to over here, like what we basically see is the, Some models are busy and some aren't.
Exactly. I understand that. That's pretty straightforward.
Yeah. And then this one, they're digging into the different, uh, arena battle arena direct and hyperbolic. And so they have a percentage of cold models that are receiving less than 5% of the total requests.
And those are not arena battle direct and hyperbolic. They're something else. Uh, Yeah.
Uh, I think arena battle, uh, arena chat and hyperbolic are showing are different ways of testing the model. And so they are, so this is just kind of showing the difference in, um, when I have these two, uh, workloads running, um, we're going to have a difference in the percentage of what we're actually utilizing on the GPU. And when I have a low batch size, Was that saying the cold models are using the CPUs better?
GPU is better than the, the No, no, it's active models. No, I'm sorry. Uh, so the, we have, uh, the fraction of the model, uh, that's actually being served, and then we're comparing it to the low late low tail of low requested models.
I probably should not have thrown this one in because I didn't dig into this enough. I apologize. Okay, fine.
Sorry. Um, yeah, no, it's okay. Uh, I can actually send you the YouTube where the research team goes through all of these in depth.
That would be great. Yeah. Um, and so then we have this one, which is just so showing the models with no request causing the GPU idle.
And this is, again, going back to those five models that are running with the different paradigms of running them and showing, uh, the median request interval, interval distribution for all of the models. And sometimes we even see, uh, that we have over 30 minutes, um, with nothing being used. Okay.
Cool. So I'm gonna dig into cost of AI in the cloud. Uh, I don't think any of this data is gonna be a surprise to everybody, but please, uh, stop me, uh, if you wanna dig into any of it.
And, uh, my third section, I actually have an entire session at the end of this four. So if we don't dig into model governance, I have 30 minutes, uh, later today that we can start digging into that with. So, uh, from, I, I wanted to just pull in data from, um, a bunch of different places.
Obviously, I have a lot of the kind of colloquial data around our customers talking to us about how AI is significantly increasing their spend. Um, we also see that the AI infrastructure spend, uh, in enterprises is causing folks to shift budget around. So a lot of projects, uh, were getting canceled to free up budget to handle the increased cost of infrastructure and services to serve, uh, AI workloads.
Um, so with this kind of coming together, uh, we have reports out of big data wire showing us the average company is spending 30% more on cloud compared to last year. And they're identifying AI and generative AI as the big drivers, uh, of that spent increase. So, um, a lot of folks are saying, you know, well, can you explain this to me?
Because it looks like it and cloud costs, cloud costs are going up, but it budget as a whole is staying relatively stable. And the reason we see that happening is because folks are shifting budget and canceling other projects that might otherwise be going on. Does the data indicate if it's a compute or storage increase, or Probably both.
But yeah. Is there one more dominant than the other? Um, not from that report.
That was like a pretty high level, high level number. Um, so I'll just kind of, uh, uh, Dr. Pull you through all of these.
Um, and again, this is me just trying to give kind of a survey of the industry as for as to like support my claim, uh, that AI in the cloud is fairly expensive and that folks are gonna need to look at how to reclaim, uh, or start managing that expense. Uh, so AI development costs, uh, so we see that we have small to medium projects, and this is from Tech Magic that are costing between 50,000 and 500,000 of, uh, in development costs. And then large scale projects are costing over $5 million.
I think these numbers probably aren't surprising to anyone just because of when we start to look at the cost of the infrastructure involved and then bundle in different services and capabilities. Um, so this is a report, I'm not sure if everyone's already seen this before. This was an IIDC white paper that, uh, our team worked with IDC on creating, and it was surveying a large number of enterprises about their own perception of ai, um, in private cloud versus AI in the cloud.
And here we have, uh, from that, uh, large scale report, which is available for folks to download that 60% of enterprises are seeing AI on-prem as less expensive or equal to public cloud. Uh, so, you know, you have to kind of read the text pretty carefully here to like, what are, what are these different colors? But basically, does your organization perceive the cost of developing or deploying AI models in the public cloud as more expensive or less expensive than developing or deploying AI models on prem?
And so most folks are saying that they see in the cloud as more expensive or about the same, um, and then about 40% are perceiving it currently to be less expensive. And there's a good chance that that really has to do with where in the production journey they are as far as like how they're perceiving it, because again, it's very easy to get started with sandboxing and experimentation and then have those costs creep up over time. And then this is from the same report, uh, from IDC, that's basically showing reasons for the perception that public cloud AI models are more expensive than on-prem.
So here we have, why does your organization perceive developing or deploying AI models in the public cloud as more expensive than on-prem? So if they said they perceived it as being more expensive, why? And a big part is it skillset and training, um, which is actually something I was speaking with some folks about, uh, prior to this session.
Uh, how do I get folks to understand how to deploy, manage, um, and experiment with all of these new tools? And then platform costs are, are really high up there. Um, we also see data security costs.
This to your earlier question, storage infrastructure costs, um, compute infrastructure costs, scalability. So again, that kind of question of it's cheap to get started, and then, oh my gosh, what is this bill? Um, AI skillset and training, and then network costs including egress and ingress, this one becomes really big because if I had data, um, that was, uh, local to my data center, and now my cloud provider's telling me I have to move all of that into their data lake in order to be able to use their native cloud services for ai, that becomes a huge expense as those data services bills start landing.
So Why is the IT skillset and training more expensive in the cloud than on-prem? I, I'm trying to understand that is because you're using cloud services rather than on-prem services. I mean, Yeah, I would assume it's upskilling everybody to understand how to use the tools.
Um, and I'm not sure if the report digs more deeply into that. Yeah. Okay.
Cool. Uh, so just kind of digging into the solution that we're providing. I know folks have already asked me questions about this.
I kind of kept this here. Um, but why, why do we think that PCF is providing an exceptional capability for private ai? So the AI space is moving really quickly.
Um, and VMware private AI foundation is designed as an AI platform that helps you deploy the infrastructure required to run your generative AI infrastructure. Uh, so you can come in, if you already have VCF running, that's amazing. You can have a workload domain, uh, that and apply our service to that workload domain.
And then your, all of the standing up the operationalization of your AI infrastructure is ready to go. Um, and so this kind of really ties into the turnkey solution. So we're offering this turnkey solution that allows customers to deploy their AI workloads successfully.
And then most importantly, our platform is designed to be extensible so that you can move and adapt to the trends of the AI landscape. Uh, when new models come in, when new technology comes in, when the latest, uh, chip set comes in, how can I quickly make that available to my internal customers to start using that on the private cloud? All of that is operationalized with our platform, uh, so that you can just lean on our automation and then, and, and then give that to your consumers to then be able to build their applications and services on top of.
So to kind of sum this up, why do we hear from customers that they're choosing VMware private ai? Um, policy and control is huge. Being able to set policy over who gets access to what expensive infrastructure, but then even more so having control over who has access to what model, having unified RAC across all of these, uh, new, uh, artifacts and being able to, as an enterprise be able to control, here are the models that we're bringing in, here's our security processes around these.
Here are the teams that are enabled to use these, and here's their self-service model, um, where appropriate, uh, all of that's baked into the platform. Um, secondly, resource sharing. Uh, you know, this has kind of come up throughout this presentation, but this idea of we will get the best ROI out of our infrastructure, um, when we are making sure that it's being utilized, uh, to the maximum appropriate capacity.
And so by being able to share things, by being able to have pools of resources that folks know that they have guaranteed access to at a certain quality of service level that's baked into the platform, um, this ties into our lower to total cost of ownership. Again, if you're getting the most out of your infrastructure investment, if you're able to successfully share things between teams easily, um, and successfully, then you're going to drive down, uh, your, the, the required money investment, uh, in AI from an infrastructure and services perspective. And yes, Just, uh, yeah.
Um, what seems to be an arms race right now is being able to run like an RAC inside embedding and vectors. Yeah. And there's a lot of vendors coming up with all sorts of crazy solutions.
Is that work you're trying, stuff you're working on? Right. Because that would be very interested instead of you having to explore all these vendor booth Right.
Nonsense things of way there, because that's a hard problem. That's a really hard problem. It's a very hard problem.
Yeah. So, so to kind of peel the onion of your question. Ah, uh, so the way that I think about it and tell me if this is not aligned with what you were thinking about is when I store the data, generate the embeddings and store the data in my Vector db, how can I make sure that only people who should have ACC had access to that data in the first place?
Will, Or PII data That might be, you can see it, but I can't see it, right? Salary is, yeah. Yeah, Yeah.
So yeah, the salary example is really good. I am creating an HR right, uh, agent. I'm going to embed all this HR data into the Vector db, but only someone who should have access to see that salary information in the first place should see any result with that return to me via any LLM or other, The non determinist nature of how it gets produced.
But Yeah. Yeah. Yeah.
So there's a couple different ways that I would look at it. Um, one way that I'm looking at it, which is a much longer term way is, uh, are back in the, in the vector database, right? Which doesn't exist yet.
Um, so really what we have to do until that exists is a couple different ways. One way, one is how are we architecting the vector database as part of our solution? Um, and so right now if you're saying, you know, this is the HR team and their data's all gonna be here, uh, the VMware private AI foundation, you would just set up different, uh, VMware, uh, namespace.
Yeah. And within that, the, the HR team gets access to this and you wouldn't add anyone to the namespace who shouldn't have access to the data in the namespace. So that's kinda like a, like a very, And not to go on, but the Yeah, you need optionality too.
'cause you've got allowed people not to use your embedding a vector space, like for example, a rag. Yeah. Right.
You need a, a customer needs be able, Whatever. Yes. So you need, Yeah.
Something. It's, it's just an, I'm glad that you're sort of working on it. 'cause it's a, it's a hard problem and it needs to be solved.
Yeah. Yeah. So I think that there's things that we're seeing invested in the database itself that may come, you know, just be able to basically check your access and be like, Nope.
But right now that's not there. So right now you've gotta control access to the database. Um, and then another more removed thing that you can do that's also just a best practice is to have guardrails on the output of the LLL egress.
Yeah. That's basically checking it at that point and saying, okay, are you, this, this data that's coming back to me was a member of this group. You're not.
So I'm going to, basically, That's how most of the vendors are trying to solve it aren't any best. Yeah. But thank you.
Thank you. Yeah. So two different levels.
Yep. Okay. I see I got kicked off the zoom again, but I'm gonna keep going until I reconnect.
Okay. So, uh, so just kind of wrapping this up and then setting you all up. For Justin's really awesome presentation where he's gonna dig into all of the concrete capabilities of, uh, VMware private AI foundation, uh, what we're providing is a platform to help our customers be able to operationalize delivering AI on the private cloud.
Um, we think there's a number of reasons that folks are gonna want to do that. We think that they, that folks need to look at really getting the maximum utilization out of their expensive GPU uh, investments. Folks are going to start to realize that the cost of the clouds is very expensive and that they can realize significant cost savings through running AI within their data centers.
Um, and then finally, model governance and the ability to have that control over your data. Um, the ability to protect your ip, um, protect your, uh, data locality and be able to, uh, be in line with all of applicable rules and restrictions as far as data transportation is going. Folks are, that's gonna become more and more front of mind for folks is they realize the value of AI and they want to be able to use it with really sensitive data, um, and still remain, uh, operationally fast and able to develop.
Um, so with those kind of three big value props that we see in the market, uh, we are starting to see really awesome uptick, uh, in excitement around private ai and.