51. Production AI Applications with VMware Private AI on VCF – Tech Field Day Podcast Spotlight
Learn More about Broadcom and VMware
See More Field Day Presentations from Broadcom
You already have the people and the platform to run production AI applications in your on-premises data center. This episode of the Tech Field Day Podcast, presented by Broadcom, features Tasha Drew, Gina Rosenthal, Jay Cuthrell, and Alastair Cooke. The public cloud is a great place to innovate and test new technologies or for bursty workloads where on-demand access to near-limitless resources is essential. Predictable and steady-state production workloads are often more cost-effective on-premises, and AI applications are no different. Your existing on-premises compute platform, based on VMware Cloud Foundation, is a great place to run production AI applications with more direct cost control while keeping your data on-premises. Running your AI applications on your existing platform capitalizes on your investment in software, hardware, and your staff, who won’t need to learn a new paradigm.
Host:
Alastair Cooke, Tech Field Day Event Lead
Tech Field Day: https://techfieldday.com/people/alastair-cooke/
LinkedIn: https://nz.linkedin.com/in/alastaircooke
X/Twitter: https://x.com/DemitasseNZ
Broadcom Representative:
Tasha Drew, Director of Product Engineering, AI and Advanced Services
Connect with Tasha on LinkedIn
Watch the Events Live or On Demand: Website | LinkedIn | YouTube
Follow Tech Field Day: X/Twitter | Bluesky | Mastodon
Follow the Tech Field Day Podcast: X/Twitter | Bluesky
Subscribe on YouTube or find it on Podcast Services
Tech Field Day is part of The Futurum Group.
Transcript
AI applications are everywhere in the enterprise. Maybe these AI applications should belong on premises running alongside all of your other applications with all of your own data, and maybe they should be on a platform that you already understand and that you already staffed up for. This episode is supported by Broadcom and the VMware private ai.
Join us for all of the details. Welcome to the Tech Field Day podcast, where we bring together a group of IT technical experts to discuss a single idea about key concepts in the industry. This podcast features a variety of perspectives from members of the Tech Field Day community, and is often recorded in association with our, our events.
Tech Field Day is part of the Futurum Group, and this podcast is published on our sister company site, techron tv on this episode presented by Broadcom. Uh, the idea we'll be discussing is that you already have the platform and the people to build on premises ai, but before we start the discussion, let's meet who's on the panel today. Um, Gina, it is so nice to have you with us again, and, uh, it's nice to have old friends.
Hi, I'm Gina Rosenthal, and I am a product marketing, um, engineer in Austin, Texas. Good to have new friends too. Jay, welcome to the show.
Great to be here. Alistair, how are you? This fine, fine.
Day or evening, wherever you are in the world. Well, here in New Zealand it is, it is a beautiful, clear morning in, uh, early autumn for us. And Tasha, you are joining us from Broadcom.
Uh, whereabouts in the world are you and what's your role? Yeah. Hey, everybody.
I'm Tasha Drew. I run product engineering for the AI and Advanced Services Group in the VCF division of Broadcom. Uh, and I am joining this call from, I think, kind of mildly cloudy, Palo Alto, California.
And of course, I'm Alistair Cook. I'm an event lead here at Tick Field Day, and Joshua was one of the presenters for Broadcom at AI Field Day six, talking about VM VMware's private AI and how it fits within the VMware Cloud Foundation product suite. And I thought it would be interesting for us to talk about this move towards having AI on premises or sometimes a desire to have AI on premises, because although the public cloud is a wonderful place to run things that you're not sure you want to use, uh, like they're saying it's a great place to fail.
It's a terrible place to succeed, or it can be a terrible place to succeed. Uh, building large scale AI applications on top of cloud platforms is a great theoretical idea, but ends up being quite expensive at times and sometimes ends up being rather inflexible And being a longtime VMware person, having taught VMware training courses for 10 odd years and been heavily involved in the VMware community, it's how I made, uh, Gina as a friend. Um, it's, it seems to me that there is again, a need for this pooling and sharing of resources that allows us to run a, a variety of different workloads on some sort of common platform.
Maybe a cloud platform that runs on premises. It seems like a sensible thing to do with our modern applications. Um, Tasha, what are you seeing customers really doing?
Because I, I think what I liked about your presentation at AI Field Day was that you brought in market survey information, information from, you know, real actual customer use cases rather than entirely relying on things that happened in the lab. Yeah, yeah. Uh, so what we're seeing right now, uh, just kind of across the industry is a couple things.
I feel like you gave me such a meaty topic that there's a bunch of different things we could kind of dig into there. Uh, one thing that we're seeing, uh, is that, uh, enterprise's cloud costs have gone up 30% year over year, driven by adopting ai, uh, and really starting to use AI in the cloud. And that is, uh, and a lot of times use ca.
Uh, when folks are starting to adopt ai, they really start doing their sandbox and testing in cloud environments. Um, it's easy to get started. There's GPUs available.
I just swipe a credit card and I can start to really rapidly experiment and prototype, but then they get, start getting the bill, they start taking use cases into production. They start having more reliable, predictable workloads. Suddenly the CFO's on the phone and he wants to know, Hey, why is our bill so high?
And they start looking at other solutions Now that we've prototyped, now that we've experimented, is this a workload that we could bring in on-prem and have under cost control? So there's an element of just kind of predictable cloud repatriation, definitely. Um, there's also just the long lead times to acquire GPUs that may lead to cloud, first on-prem later.
Um, but then there's also just rapidly experimenting, prototyping be able to, being able to try out all these tools and services. And then once you realize what works for you, let's replicate and scale that in a place where we have stronger cost controls. Jay, your perspective on sort of business management, cost control, data governance was one of the ones that also kind of swung into this, this, uh, running AI on premises.
Sometimes there's a, a necessity from a data governance point of view. There's a couple of things inside of that. First, if you scroll all the way back to the beginning where we decided that it costs should be contained, it, it is a reasonable thing to do.
Uh, the next thing that occurred that the, the cloud provided us was, uh, all these operational expenses. So, hey, we don't have to go buy a bunch of, you know, capital intensive equipment and sweat that asset and appreciate it and advertise it. It's lovely, you know, swipe the prove credit card.
But of course, the controls, uh, the governance, the guardrails, uh, didn't necessarily ship with that experience immediately. A lot has changed since then. And I think there's, there's methodologies, there's the technology business management, uh, methodologies, TBM, uh, that has kind of vaulted into the world of a three-legged stool of there's the, it cost traditional, then there's the clouder, uh, sassy type of, of cost.
And then there's the, the labor, uh, elements that, of course, with, with all of this, uh, amazing technology, we still need humans in the loop, uh, or just humans to take, you know, and care and feed for all the data pipelines that maybe haven't been fully, fully, you know, evolved into a managed service. When we do all these things and we see the totality, it's A-C-I-O-C-F-O view at that point, I think then you get to make these business trade-off discussions and compliance is arguably more important now, uh, than it's probably ever been. Uh, what, what if, if cloud is the ability to make mistakes at scale, um, what mistake are we making that, uh, you know, appears in print?
Um, and yes, you can weigh PR by the pound, but I think being associated with the wrong, uh, you know, uh, press coverage is very important, uh, for boards, uh, C-level executives that, uh, do want to embrace this ai, uh, this, this promise of ai. But I think the control element does still favor even if you're doing, to your point, experimentation dev test, QA environments in a public cloud context. Um, and there's probably some screening deals now, you know, GTC is, I think, taking place as we're recording this.
And so you can anticipate there will be provocative pricing offered by the large cloud service providers to kind of entice you, uh, bring it over here. But I think, uh, for long running enduring workloads, it is still quite proven that a, uh, traditional AI on-premises or a machine learning on-premises architecture as part of a greater hybrid by design approach is probably, uh, a more advisable, uh, story than just going quote unquote binary all in on cloud or quote unquote binary all in on just private ai. So that's my thinking.
So your thinking is, it depends. It does depend. It does depend As, as predicted.
But I, I think that's a, a really important element wherever we're getting into production, deployment and, and large scale adoption of technology, uh, use case by use case, the, there's right places to put applications for a variety of reasons. And some of them are financials, some of them are short term operational. If I can't buy the latest and greatest GPUs, then I can't run the workloads that require those latest and greatest GPUs in my own infrastructure.
But of course, we don't always need the latest and greatest GPUs because that was a, an angle that, um, Gina brought up as we were getting ready for this, is that not everything needs the latest Grace Hopper or, uh, the, the Fineman architecture that's coming in a few years time. Um, lots of things run just fine without the latest and greatest, including deep seek. Yeah, I definitely, I mean, people have been running AI workloads for a really long time, right?
So I think a lot of the times when we think about, um, the Grace Hoppers and whatever else will be announced this year by the different powers that be, um, we think about generative ai or we think about, you know, some really intense deep learning with lots of, uh, training and lots of retraining and, and lots of, um, just lots of CPU power needed to do that part of the, of, uh, of the AI pipeline. Um, but you don't need that for some simple machine burning things, right? So if you've just got some predictive ai, you're trying to figure out, well, if everybody loses, uh, if everybody lives in Austin and this part of Austin and there's so many crashes, then what really, what do we need?
How many crashes would there be? And how high does the insurance have to be? That's kind of my situation right now.
So, but that, that already happens now, and it doesn't require tons and tons of, of ma manpower like, uh, like the, like some of the generative AI things would need. And, um, so, so I think just understanding what it, what's the business problem that's being solved by ai, then, you know what type of workload it is, then you know how to architect for it. And when you think about, I think about performance, not just as the, the clusters of machines that we have or the platforms, it's also by the people.
So the performance of the business is you're trying to get to a point where you can, you can see predictably what's going to happen, how the insurance rate's gonna go up, how the farming's gonna go, all of these types of things. You may not need to go to a bigger model. And honestly, when you start looking at the labor, like Jay said, if you've already got a platform with people that understand how to architect and manage and, um, and control that platform, this is just a little stretch to help them understand another workload and understand how to do all the management and all the data hygiene things like protecting the data and security and all that kind of thing as well.
And you might not have that. Um, the care and feeding for that typically doesn't go with the development teams. The more the let's make this happen and let's get it done, and then it's up to the operations teams to stick it to the ground and, and make everything work in a, in a performant way from a organizational standpoint.
So I think, yeah, we can, we don't have to do everything at once. And I, I think you brought up some really, um, important topics and you, you and Jay, I think are in, in a solid agreement that having people that know how to operate your environment and are familiar with that environment is one of the ways of significantly reducing cost and also enabling agility. It's really hard to move fast with something that you're unfamiliar with the first time you start using a cloud platform to deploy your applications.
It feels very different to on premises. Tasha, as you are seeing customers deploying the VMware private AI solution, are you seeing them starting with large language model predictive ai, or is there, does there continue to be a lot more of the more machine learning gen, um, predictive rather than the generative models being used? Yeah, so looking at our customers, uh, we have, most of our customers have been been doing machine learning, uh, predictive AI for a long time, right?
Like they already have that infrastructure, they're often already running it on VMware. Uh, the real difference is starting to look at private ai, uh, as our solution that is really, folks are coming to it from a generative AI perspective. They're very interested in creating, uh, agents.
They want to create rag applications. They want to open this up, um, and be able to serve, uh, these capabilities from a platform perspective. Uh, how, uh, today without a platform, you might end up with every team at your company running their own copy of LAMA three, right?
Like, you might end up with all of these different copies of models that bunch of different teams are supporting, needing to procure GPUs for not sharing with each other. But with, with a platform approach, you can source the GPUs, you can run the models as a service, you can have an API gateway in front of that. You can horizontally scale the models behind that based on demand, but you can run things in a much more efficient way if you have a baked in platform approach.
Um, and so really from customers, what we're seeing is a lot of really interesting use cases around scaling out generative AI assistance for their various industries, um, starting to think about how they could have assistance work together, um, to have, um, a true agentic workflow. Um, folks are really interested in RAG and just the ability to intelligently, uh, get data out of sort of the massive amount of documentation that every enterprise creates. Um, but then they need to have strong privacy, security are back around that to make sure that the right person is getting access to the right sets of data.
Um, so definitely interesting challenges from a privacy perspective. Um, but having a nice baked in platform approach, um, that to, uh, the earlier points, you already have a team that knows how to manage, deploy scale and is now just stretching to a new use case, um, is definitely very attractive for folks. I really like also that, um, API interface, you know, the, the API gateway interface to it, because one of the things that I see is a lot of developers who would like to have AI or being told they have to sprinkle AI all over their application, they're not actually data scientists.
They're not people who are going to be able to write the prompts and, and understand how to tokenize their existing data to build that rag solution. Um, being able to segregate that sets of skills and duty is gonna be really useful. And I think, uh, another element for sort of the integration ju were raising was, was around some sort of universal interface for talking to agents.
Yeah, I think right now we're probably in what I would call the earlier days, but, um, you know, uh, uh, code which people use the company behind that philanthropic has been, uh, releasing, um, uh, certain things to the, to, to the world. One of those is the notion of a model context protocol. And you could, again, think of this as, uh, a, uh, uh, a a paradigm of secure two-way communication between different agents, uh, that one might have.
And so if I wanted to have one agent talk to another agent, uh, security, obviously, you know, to Tasha's point, top of mind, um, defaults folly of defaults is that, hey, it worked, ship it. Um, and, and then we, and then we learned that, um, uh, one of the examples that was, uh, shared yesterday at All Things Open, uh, for AI in, uh, Raleigh was, uh, you know, uh, in a Canadian court of law, I believe I'm not a lawyer, uh, someone was, uh, uh, they they tried to cancel their, uh, ticket and the LLM that, uh, was supposed to pull a rag and then tell them what policy was for, you know, getting a refund for their ticket, said, yep, not a problem. You were refunded.
And so through the, the glory of the court system, um, this company that tried to refute, like, no, we don't, we don't wanna refund you, that actually is against our policy. Yes, but you're a rag and whatever system, whatever it connected to, and whatever it communicated to actually did say, in fact to that person, yeah, you're covered. No worry about it.
We'll, we'll refund your money. And so I think we'll all learn some of these valuable lessons through, again, the news, um, because some of these, uh, some of these spectacularly wonderful failures, and we must have failures. That's the only way we're gonna get better.
But some of those are gonna be very instructive. And I think, um, having something like, um, MCP, this model context protocol will allow for a, a greater de-risking, if you will, in the type of two-way communication. We all want to see, we want to see agents talking to agents doing things better on our behalf than we ever saw before.
But to Tasha's point, uh, much like b and w, it's not how fast you go, it's how well you go fast. I think the another really significant element is, as we see production adoption is, is around governance and understanding the model history and making sure that we're not getting a lot of drift from in, in our models as they're being prompted. We've seen those, again, as Jay says, the the disastrous things that happen with, uh, unfettered access for humans to these, these models where we can destroy things really well.
Um, actually governance is, is central to the, uh, VMware private ai. What sort of functionalities for governance are, are being delivered? Yeah.
Uh, so when I talk to customers, uh, I, a lot of times when I'm talking to them about model governance, uh, it really resonates, especially if they're in a regulated industry. Um, especially if they have started experimenting with building their own models, fine tuning or training them against sensitive data, um, they become very aware of how critical model governance is to them. But it's even important to folks who are just trying to pull a model, like a really well-trained model off the shelf, so to speak, from a model repository like NGC or hugging face, starting to have enterprise workflows for how to ensure that that model is properly vetted, security scanned, tested for appropriate model behavior, starting to have, um, an enter a stance in your enterprise about what is our risk tolerance around the model, having problems answering this kind of question or this kind of a attempted, um, escalation, right?
Like, and starting to kind of measure all of that, um, and, and continuously test your models for how they're behaving against a series of prompts. All of that really feeds into this model, model governance story. And so what we have is a set of basic automatable capabilities that are part of the private AI solution.
We have a model gallery. Um, we've, uh, basically opened up Harbor, which is one of, uh, V C's most popular open source, uh, donations, uh, to the Lennox Foundation. Uh, and we've added the ability to store models in Harbor.
Um, the OCI specifications been extended, uh, so you can package and deploy models, um, as OCI artifacts store them in Harbor Harbor within VCF already has the ability to do regular security scans. And then we also started to show our customers how to automate when they're pulling in models, but also as part of A-C-I-C-D process, testing the behavior of the models itself using tools like gift guard. So now you can basically red team models programmatically and have one LLM try to get the other LLM to say the wrong thing or say something outside the boundaries of what you as an organization are comfortable with.
And then you get a really nice automated report saying how your LLMs are performing against that perform. So you can use that to validate the models that you're importing from outside. You can use it to continuously validate the models to ensure that model behavior drift isn't happening.
Um, and then this is all baked into the platform as automatable against your own software release processes. Um, once you then store the models in Harbor, um, we have a nice automated, uh, process that's exposed via CLI, where once you wanna import or store a model, you can use our CLI, we'll bundle it up as that OCI friendly artifact. We can attach tagging, you can attach tagging, and then you store it in a model gallery that's only shared with the appropriate people.
So if you have a model that's just for a small group of data scientists to test and experiment with, that would go in their private gallery. But then if you have standard models that you're very comfortable with the behavior of, and you want to provide to large teams of people, you can then have public galleries. Um, so just kind of adding the basic capabilities in the platform to make sharing, deploying, validating, and testing models part of your regular practice.
I really like that Tasha. Um, that is really cool to have that model governance built in. Not only that, but have it as part of Harbor and have it part of the CICD pipelines, all the rest of it.
So, you know, when you get back to talking about the performance of the teams that are using it, you know, already you, you know, your op your ops team is good to go. 'cause they, they stretch a little bit, they understand the workload, but this helps bring in the other teams that are responsible for the work to, you know, to, for using the model and training the model and, and all the other parts of that to make sure that, you know, if you had to pass an ethics test, uh, my podcast partner is that's what she does. She does ethical AI and does those kind of reportings for IEE.
Um, but yeah, if you had to pass a test on, okay, you're doing AI for medical or you're doing AI for the government, let's make sure like these processes haven't harmed anybody or that they're actually doing what you said. And this seems like you said there's reporting like that just seems really good and you, you go back to the people performance piece of it, you're not learning a whole bunch of new stuff, you're just learning how to automate it, which is, which is great. We could spend hours discussing this topic and, uh, preferably o over a meal or a drink somewhere, and we will carry this conversation on.
But thank you all for joining us today at the Tick Field Day podcast. And before we go, where can people connect with each of you in order to carry this conversation on social media websites? Where can people find more about you?
Gina, Best place for me is g uh, is LinkedIn. Um, it's LinkedIn gmx is what you can find me under on the URL or Gina Rosenthal. Just search for me And Jay, whereabouts can people connect with you and, and have some interesting conversations?
org or like Gina, you can find me on LinkedIn as well. And Tasha, where can people find more about you, what you believe in, uh, how the private AI is working in Cloud Foundation as a whole? Yeah.
Uh, so we have a great, uh, VCF blog series, uh, that's really digging into private AI and how to automate and create all of these capabilities. Um, so definitely recommend checking out the VCF and private AI blogs on the VCF blog. Uh, and then I'm also on LinkedIn, uh, and it's just slash tahi, so Tasha, but with a y instead of an a.
Yeah, Nice. You're getting a nice short, uh, handle that often requires that sort of manipulation too, doesn't it? Yeah, and of course I'm Alistair Cook and you can find me online as Alistair Cook on, uh, LinkedIn.
You can find me as DMI tess nz as my, uh, social media handles. co nz for my own site. Uh, you'll find me also at the Futurum, um, blog as well, or the analysts, uh, elements on there, and periodically some of the places over on Techstrong tv.
So thank you for listening to this episode of the Tech Field Day podcast. If you enjoyed the discussion, please subscribe on YouTube or your favorite podcast application so you don't miss a single episode. Do consider us giving us a rating and a really nice review.
Uh, this podcast was brought to you by Tech Field Day, the home of IT experts across the enterprise, and a part of the Futurum Group. com/podcast or view us on text on tv. Thanks for listening and we will see you next week.