Unlock AI Cloud Potential with the Rafay Platform
Haseeb Budhani, CEO of Rafay Systems, discusses how the Rafay platform can be used to address AI use cases. The platform provides a white-label ready portal that allows end users to self-service provision various compute resources and AI/ML platform services. This enables cloud providers and enterprises to offer services like Kubernetes, bare metal, GPU as a service, and NVIDIA NIM with a simple and standardized experience.
The Rafay platform leverages standardization, infrastructure-as-code (IaC) concepts, and GitOps pipelines to drive consumption for a large number of enterprises. Built on a Git engine for configuration management and capable of handling complex multi-tenancy requirements with integration to various identity providers, the platform allows customers to offer different services, compute functions, and form factors to their end customers through configurable, white-labeled catalogs. Additionally, the platform features a serverless layer for deploying custom code on Kubernetes or VM environments, enabling partners and customers to deliver a wide range of applications and services, from DataRobot to Jupyter notebooks, as part of their offerings.
Rafay addresses security concerns through SOC 2 Type 2 compliance for its SaaS product, providing pentest reports and agent reports for customer assurance. For larger customers, particularly cloud providers, an air-gapped product is offered, allowing them to deploy and manage the Rafay controller within their own secure environments. Furthermore, the platform’s unique Software Defined Perimeter (SDP) architecture enables it to manage Kubernetes clusters remotely, even on edge devices with limited connectivity, by establishing an inside-out connection and a proxy service for secure communication.
Presented by Haseeb Budhani, CEO, Rafay Systems. Recorded live on September 10, 2025, at AI Infrastructure Field Day 3 in Santa Clara, California. Watch the entire presentation https://techfieldday.com/appearance/rafay-presents-at-ai-infrastructure-field-day-3/ or visit https://rafay.co/ or https://techfieldday.com/event/aiifd3/ for more information.
Transcript
Hi, my name is Hase. I'm with Rafa. Uh, I am about to talk about the Rafa platform as it relates to the AI use cases.
So portal where somebody can log in as an end user. This is white label ready. So our customers will change it to whatever they want it to be.
Um, and they will decide to sell whatever they wanna sell on the left, because each of these is a template on our platform. But for their nd customer who may belong to a team that belongs to an enterprise, and they may have a thousand enterprise customers each, they all get a different experience potentially because of the standardization work we've done in the platform because of the IAC concepts we have thought about because of the GI ops pipelines that are built into the platform. It's not like there's Argo, of course, all very cool things, but when you start thinking about a platform that will drive consumption for a thousand enterprises, GI ops cannot be a tool that we use.
GI ops is actually part of the system. So our platform is actually built on a gate engine. Otherwise, we can't do configuration management across so many people.
Our platform needs to understand what identity means at a very different level of multi-tenancy, because in our world, yeah, my customer could have a thousand enterprise customers each with their own, uh, Azure AD account. Mm-hmm. By the way.
And they will log in, God knows, with what u uh, URL, and it should work. But eventually they come here and they can come and ask for a bare metal server. This SSI, I have no policy for that.
I, I hope I have something for VMs. Where am I logged in? I logged in in the wrong place.
Clearly I did not tell my colleagues to, okay, no wrong. No problem. See, this is how we know it's a real environment.
Right? Number seven, if you're counting Brian, So I'm gonna log in actually to a, a test environment for a customer. Uh, because clearly my environment is not what it, uh, needs to be right now, but that actually makes it more interesting.
This customer, for example, has decided to offer three kinds of services, compute platform, and industry ai. So, by the way, this is all configurable white label. They can call it whatever they like.
And to their customers. They offer three compute functions, form factors, Kubernetes, bare metal, GP as a service, bare metal VMs. There's GPU as a service.
It is the thing. Here's the thing, 19%. And the idea is that if all you want is Kubernetes, which why would you need it?
But okay, you gotta sell ku. You gotta sell compute, right? Something that you must do.
It should be so simple that you just tell me what version you want. Everything else is magic. And, but if I were to run it, it'll take 15, 15, 20 minutes.
It's gonna gimme essentially a service, a service account. Here you go, connect to it. Done or cube config, if that's all what you want.
Or you may sell platform services. Now, in this case, they're offering NVIDIA services, uh, n video, for example. It's a tool that you may have heard of, but the experience is very simple.
I just want to consume a model. I don't need to understand anything else or agentic applications. I just want an agent library.
So our work collided with this AI space because of two reasons. One was the core platform, which does this sort of the self-service consumption using standardization. Of course, that turned out to be a really important thing because otherwise it's really hard to run a cloud.
See, as a cloud company, I can't behave like an IT organization where I talk to somebody and there's a ticket. None of that has to happen. It has to just work.
And the second thing was we had built a platform in our company that we built because we thought it was really cool. We didn't really know what to do with it till we ran into ai. So we built, uh, essentially a serverless layer where you can bring any code run into top of a Kubernetes or VM environment, and you can run n pieces of code at different points based on what the output of the previous piece of code is and deliver an experience.
So we were just playing around with it, and we took Q Flow and we built that as an experience. Mm-hmm. And we showed it to people and they said, this is pretty cool, we'd buy it.
So we sold it, And that became an AI product. But we started with something entirely different. We were trying to solve a very simple problem.
We were trying to solve a platform engineering problem. Mm-hmm. You think we've done a pretty good job of that.
Uh, but, but we were a small company in that space. Mm-hmm. And then using this serverless layer that we built, where you can bring your own code and deploy it, which we have.
And in fact, we've had partners do that as well. Partners like DataRobot, we can deliver DataRobot as a service because we own the infrastructure for the, for the provider. And we wrote the serverless layer where DataRobot came and wrote their own code.
And we can sell that as a card on our, on our, in our platform. Mm-hmm. We can sell anything.
We can sell Nvidia applications, which we do. NIM is an application from Nvidia, or Tool a DataRobot is a full fledged app. Dataiku, qre, many, many, many things.
And when we kind of recognize what we had in our hands, what we said, gosh, man, we can, we can take this to market. Mm-hmm. So then we look My, all these other services are customer generated solutions that they are on top Of this.
These are the purple ones are partner generated, the partner. These are, these were written by a partner who's supporting this specific telco, this partner, if you recognize the logo is number eight. Yep.
Who is, yeah. But that's their logo. But, but yeah.
I mean, Ray, these could be, this could be, uh, homegrown. It could be, it could be Homegrown people Could be a lot of this. So when I'm looking at, like the code converter, for example.
Yeah. Right. Um, think of it, I, I put in, in my mind ahead of that word compliant.
Like, you know, some kind. There's, there, there are some, you know, every shop has their own set of rules and, and, and customs and like, whatever. Right?
So, uh, rather than like, they, they're training some, some sort of model or tuning some sort of model to get you halfway there. Yeah. Uh, using conversions they, that they've already done.
And then now you can expose it in the, in, in, in a catalog. So here's another example of something similar. So I finally remember the, the user just took me forever to remember the user.
Uh, but okay, this is another environment we have. We keep playing with these things. So, uh, but, but basically same idea.
Um, uh, I just want a notebook. I Like a Jupyter notebook. Yeah.
Just a notebook. Uh oh, I know why, because I don't have a catalog. There you go.
Uh, I just want a notebook. Yes. I want a notebook.
And I got it right. So, mm-hmm. Now these are, these are test environments.
You'll see a bunch of variations of the same thing. But the idea is that our customers will pick one of these and, and, and provide them to their customers. But I just want to know what profile I need.
And maybe optionally, I have these GPUs available and it should be completely an experience like you get on Amazon. This could be on-prem, this could be in the GP cloud. 'cause both are my customers.
But I just wanna make sure we tie this back to the original, previous conversation. The reason why we're able to do this is not because this is an IDP. How do I know there are GPUs available?
Because we have an inventory management system. Where does the Jupyter on a Kubernetes cluster? Where'd that come from?
We built one, how do I know this? Jupiter Kubernetes clusters compliant to the policies of this, this company. We have the blueprints for that.
How do I give this guy access? We have identity management. How do I do this across many, many, many companies.
We have multi-tenancy. All of the rules from before. Mm-hmm.
Must be true. Otherwise, you cannot run a platform engineering function. You cannot run a cloud.
Just to be clear, the end point of this is that the Jupiter notebook appears at my, uh, mail stop or whatever, and I could go, it's an IP address provisioned. And, or, or, or is it a virtual Jupiter notebook that I'm Getting? Would you like one in your mail?
Uh, I'm just asking. Yeah, It's a URL. Yeah, it's A-U-R-L-L.
Yeah. It's a URL. Yeah.
You'll get a URL. Yeah. You just connect to it on internet.
Okay. And it's a token. And then you play a cloud.
It's cloud experience. Now the customer may decide, I don't wanna do this on the internet. 'cause that would've been an interesting trick, is, uh, you know, Send, send a mail Physical notebook as a service.
PNAS. Yeah. If I build it for you, physical notebook, that's a Tricky acronym.
There. H there's A certain connection though, that I think you gotta, I gotta, you gotta explain that. Pull back.
Pull back. Well, so like, new role. Are you doing something fancy around like the network connectivity to things like vSphere?
Because I mean, I'll be honest, I was able to go and log in and create myself an account on your website here. Yes. And get into it for a, not a place for me to go add these here.
Yeah. Because somebody has to, somebody has to allow you to do that. But yes, I, I, I, once we take a break, I will do that for you.
I will, I will give you access and You can, oh, no, I mean, that's what I'm saying is does this need head to have, is there, like, is this designed for, I can deploy my own vSphere and then give myself access to it with this? Or is it a, I mean, is there like A-G-R-P-C fund, like tunnel thing that gives access to it? Or?
So for vSphere, what's gonna happen is you're gonna deploy, uh, uh, what we call a, an agent. Okay. Uh, uh, in your environment.
What a, what a fancy name get out agent. Right. And that agent is gonna talk to your vSphere environment, uh, uh, where's my cluster infrastructure cluster?
And basically we'll create, uh, no, where is it? Cloud credential. So we'll create a cloud credential for quote unquote vSphere.
Yep. Okay. Okay.
And that agent is gonna talk to your vSphere, and that agent talks back to our cloud command and control. Now we can do anything to your vSphere environment. Gotcha.
Cool. Okay. Poor choice of words.
We will not do anything to your cloud cloud, but that's how it works. That just Bring up a good question, which is, where are all the security? Um, yeah, let's talk about that.
So, so the, so our enterprise customers use our SaaS product. It's a SOC two type two compliant platform. Uh, we Very, very frequently share our, you know, for, uh, uh, pen test reports, et cetera, with our cus with our customers.
Um, seems pretty good. Um, so, you know, the, the nice thing about SOC two is, uh, that you understand that our operation controls are pretty good. So not nobody can just log in and kind of do whatever they want with it.
So if we don't for sure, we also provide reports on our agents, ma'am, so that you are comfortable that anything deploying in your network reporting network is sound. But the larger customers don't do that. So the larger customers do not use the super large customers or the cloud providers.
They are not using our SaaS product. What they are doing is they are using our product. So the Rafa SaaS controller, The sell, we can deploy it anywhere.
We can help you remotely manage it, or you can work with a partner. Many of our partners are trained to do this. Some enterprises can do this themselves, some cannot.
The ones who cannot, they can work with a partner who will manage it for them on-prem. If you want this to be truly agat, which many of our customers do, but every GPU cloud provider is doing this in an AGA fashion makes sense. How can they rely on a third party status control plane for their command control?
So that's how Gina, we address the security questions. Thank you, ma'am. I, I looked down and I talked to Gina because there's a screen at the bottom that, that I'm looking at.
He's looking down on you. Gina. I can sit on the ground.
It's all good. So pilot. Oh gosh.
Alright. So, so the idea is that, you know, different customers will want to sell different things, right? So they'll want some compute functions that they sell.
And you see a bunch of them on the left. This, this is a test environment. So these, there's a bunch of silly examples.
Uh, or they want, uh, some AI ML platform. Okay? Now, you can call them other things.
You can call them, you don't have to call it on slot. You can call it, uh, rag because you, maybe you don't want to tell people what, what, what it's built on. But the idea is that you can sell anything.
And we've now seen some very creative things happen on our platform where a, they build these use cases on our, our platform. So we bring a library of them, and you can sell them. We don't, you know, it's fine.
Uh, b they bring, so they bring them, and then they put a price point on them. If I can find an example of that, I'll show you. But basically, you can tie prices to these things, uh, maybe here, I don't know.
And the intent of that is that we have a billing engine, or I should say a price, uh, a billing metrics engine in our platform Mm-hmm. That you can tie to Salesforce or NetSuite or whatever tool you like, monitor 360, and you can start generating bills for your customers, or do this internally. You can do chargebacks if that's all you care about.
Or you can actually generate bills for your customers. So in our platform, you can make a decision. Uh, Brian's gonna pay two bucks an hour for something, but guy's gonna pay three bucks an hour, and the platform will track it.
It's supposed to be at the cloud. You should be able to do an agreement for the customer. If they pay for six months or a year, they should get a better deal.
Versus the on-demand customer, you can do that too, trying to help you build a cloud. Now, even that capability we built for enterprises, thinking enterprises needed, and they do. Mm-hmm.
But it turns out that the cloud customers need exactly the same. Right. Every large enterprise has been trying to become a cloud provider without owning the infrastructure.
Mm-hmm. And I wish what I just said was clear to me three years ago, it wasn't, we were solving for platform engineering. The problem is much bigger.
The problem is, how do I operate a cloud? Mm-hmm. I may, I may own some of this infrastructure.
I may rent some of this infrastructure, or I may lease some of this infrastructure, but it's my cloud sitting on top of somebody else's infrastructure. That provider could be Amazon or Core ViiV, my cloud, my policies. And that's the opportunity.
Going back to the original question, what is the opportunity? And when it comes to that big problem, nobody does it better than other Regina's question about who do you compete with? Ma'am, we compete with people who build this themselves Because there's nothing, there's no one thing out there who can do this.
Now this diagram, as I shared earlier, is on NVIDIA's website. So of course it talks about NVIDIA's applications and tools. But of course, we've looked at a whole library of things that can be delivered.
And of course we do. The number one use case right now is model as a service 100% of the time. Everybody wants that.
And the number two use case is single G-P-O-V-M. That's the, like across the world. Yep.
Those are the two use cases. Everything else comes second this and third this and fourth. Mm-hmm.
I mean, do you support, uh, MIG and MSP? Yes, sir. Absolutely.
Virtual DPU virtualization. Absolutely. Oh, a hundred percent.
So, uh, uh, in fact, I thought my colleagues just, uh, wrote a blog about this, uh, or ds, or they write a lot of blogs. When do they do work? Mm-hmm.
Maybe it's easier to do, Meg. Uh, yeah. Yeah.
Yeah. So there's a bunch of strategies we recommend that you use depending on the situation. Just to clarify something on that previous slide that you were just showing, Which one sir?
The, the grid. Was it? Oh, yeah.
Nvidia slide stacks the Stack. The Nvidia. Yes.
There you go. So when you say data center down there in, on your environment, vening, then that's your data center? Correct?
We don't own anything. We are just a software company. It's a customer's data Customer.
Well, okay. No. If, if, but when you, you, I mean, using your product, it would be my data center.
Yes, sir. Okay. Yes, sir.
Yeah. But it could also be my Azure environment. Right?
Exactly. That's right. It could also be my, my Edge implementation.
Yeah. Okay. When we raised money for this company, nobody even thought that, you know, six, seven years later, people would be building data centers.
Again, at the time it was all cloud. Right? But of course, now the world is very different.
It's the dessert topping and the floor shine floor wax. How does this run on? Hopefully not in the same spoonful.
How does it run on the edge systems with, you know, three servers and, you know, it's, it's, doesn't seem like it's applicable in that environment. Yeah. Okay.
Let's talk about that. So, uh, we have a customer Who Runs Kubernetes on cruise ships. Oh, man, this website is very different than would I remember it to be, uh, zero trust.
How do I click on it? Docs. Okay.
Oh, there. Okay. So this customer runs Kubernetes on, on cruise ships.
The cruise ships have tit links. There's no VPN. Okay?
So think about this for a minute, huh? So the controller, the Raffi controller in this case happens to be SaaS. It's running out there somewhere on the internet.
The cluster is on a cruise ship. Mm-hmm. How do I get to it to manage it?
Gina, this is a security issue that we thought about a lot, and we applied a very, uh, very, very unique way to think about solve the problem. So, so we wanted to solve this problem. So our thinking was, what if we run a SaaS service and there are a hundred thousand clusters that we need to manage now for our customers?
How are we gonna open up ports on their firewalls to get to them? That's not gonna happen. Nobody's gonna let us open up all these ports.
Mm-hmm. Ah, but what if the cluster could reach out and talk to us? There's nothing special about that.
Right? That's, we all understand that right. Inside out connectivity.
Yep. But the problem with that is, uh, how do I know to whom I give access to what? So in the security space, there's a concept called software defined perimeter, SDP.
So we basically built an SDP for Kubernetes. So we essentially run a proxy service on the internet where the yellow box, the yellow or orange color boxes, right? Mm-hmm.
So when you do a cube cuddle to a cluster, you think you're talking to the cluster. Actually, you're not. You're talking to an endpoint on the internet, and it authenticates you, authorizes you, and it knows who you are, where you're trying to get to.
And you talk to the proxy, and the cluster talks to the proxy. CDN, it's a CDN. Mm-hmm.
For, for q we're the only company who does that in the market. But because of this, we can now manage a tiny cluster on a cruise ship. And on top of that, when it comes to Kubernetes, we can run, uh, a single node remotely managed Kubernetes cluster, single node, and the overhead for Rafa bits, and all the stuff that we do in our, in these clusters, is less than two cores virtual core.
So we can actually run a pretty small box with Kubernetes at the edge, and we can manage a thousand of them. Nobody can do this, but, So I, I apologize, but I'm still a little confused. I'm trying to reconcile your SOC two compliance with we just a software company.
So are you providing only the software stack, or are you actually providing services as well? We don't sell, we run a SaaS service, and we also optionally provide that same software as subscription to our customers. And they deploy our SaaS service in their data center, basically.
Right. So they don't a version of ve we don't sell any services. We have a very strong support team that will help you all day long.
But we don't provide professional services or Anything. No, no, I don't. I meant our service.
Yes, sir. We do run a SaaS service. The reason why I keep switching back and forth, and my apologies, I should have clarified, thank you for asking the question.
We started this company thinking, all we will do is SaaS did, we got punched in the face by a large company. We said, no, I like it, but I'm gonna run it here. So then we had to think about a different way to architect this, where we created this concept of cells like truly cells.
There are, so these are three known high, high lable systems, right? That you can deploy any of them. Mm-hmm.
And in theory, we can, I say in theory, because we don't, we don't do them all the time. We can manage them all centrally if the customer chooses for us to do so. Mm-hmm.
Or a partner can do it, or the customer can do it themselves. So it's a very elegant architecture so that we can have a multi control plane, control plane, if you will. That's what we have built.
In reality, what happens is we run a SaaS. We had, we done multiple SaaS instances on the internet, because different regions will have different needs, right? Mm-hmm.
But then many of our larger customers, they just run a Rafa in-house. So if I'm a a, an enterprise, a smaller enterprise, or smaller MSC, I might use the SaaS version. But if I'm a larger, or I want to do my own thing, I could deploy in my site, whether it's a cloud or my on-prem data center or wherever the RA tool set and the whole stack Yes, sir.
Myself. Yes, sir. It's a pretty significantly larger, uh, that you're gonna download to get that going.
But yes. Okay. You can do that.
And, and the, the pro, the, I was gonna say the problem, but this is, this is not a problem per se, but, uh, a lot of our business right now over the last year has been GPU clouds globally. Mm-hmm. But they're all, they're not gonna use our SaaS service, right?
'cause they want to be sovereign. Right? But then we end up deploying our Rafa instance in their environment.
So that's become, uh, sort of our go-to right now. And that would be the AI use case. Everyone, AI use case is consistently sovereign.
They want a complete sovereign environment, but they want the power of your tools to manage and deploy the infrastructure rather than 20 different building their own stack of 20 different open source tools and Whatever. You can, you can launch something in two years, or can launch something in two months. What is better, right?
And right. Two months is better. My how much, much you lost our product, me or the product.
Well, They seem to be the same thing. That's, I guess it's all the, it's like, you know, pretty much economous I guess at this point. Uh, the look, the, so here's how we do math on the, on the AI side, right?
So, uh, we, at the beginning of this conversation, we talked about per GP per hour pricing, right? Um, the, the premise is that, hey, Mr. Customer, if you're cost to build infrastructure, power cooling, everything is, let's call it a buck 50, as we discussed earlier in this conversation.
And you could sell it for a buck 80, or you could sell it for three, or depends right? On the use case. If you sell model as a service, you can actually make more money than that, depending on the model.
If you make three bucks, four bucks, whatever, would you pay me 20 cents an hour? Yeah. If you have consumption.
Yeah. So the answer is yes. Mm-hmm.
Will you pay me 30 cents an hour for the right, for the right use case? And that's a conversation we're having. So we're basically talking to them about consumption based pricing on the upside that they will experience now that they have the right software stack.
And so's a very simple conversation, Ashley, right? We, we say, tell them, Hey, look, pay us something upfront, right? We wanna get some minimum commit to a hundred GPUs to hundred something, something, so we can, you know, where our beaks, and after that, everything is consumption based.
The good news is when people have software, they seem to have consumption. Everybody wins and they're happy to pay. Now, in the enterprise world, obviously that's not a concept, right?
There's no concept of consumption. Is there infrastructure in the enterprise space? We're going also something similar.
So we're saying, Hey, let's look at your infrastructure. You may have a, I don't know, whatever, a thousand clusters, but let's not do that. Let's go with your top team.
The, the somebody in, in the most pain. Let's start small. Always start small.
Get, you know, once they have your, once they, once they kind of feel good about the investment. Once you have their trust, it all grows. But start small.
And that's, look, as a small company, we can get away with these things, right? Uh, large companies sometimes have too much overhead. We are still a small company.
We're 150 people, right? So we can get away with those kinds of strategies. But the, on the AI side, with GPO clouds, they love the consumption model.
So back to the models of consumption. So you've got either, um, you provide as a SaaS, or you can give it to people they can install locally. So when it, when they get into the SaaS, what's it, what does it sit on top of?
So our SaaS service runs, uh, in the us we run it in Amazon, in other regions, we do run it in OCI, it depends on the, on the region. Uh, because we, we end up running multiple endpoints on the internet, uh, looking for something cheap, right? Basically.
Right. O Cs, they, they've been very kind to us. They've given us great deals globally and AWS we've been with forever.
So we just kind of keep using it. Uh, um, if it's on-prem, then you know, you decide that'll be VMware, bare metal, whatever's easy for you. Right?
And I would definitely recommend, you know, you don't need bare, you know, bare metal systems, don't waste the money, right? Just spend a VMs good to go. Are there any advantages they get with Amazon, further integration to things that they might already have?
Is that a value Product? Ex Thank you. Excellent question.
So, uh, sort of the Terraform provider you've written to set up our controller case might as well, right? Might as well write a Terraform provider for that. Mm-hmm.
It assumes that this is AWS or Azure or GCP or, or, or vSphere. So it understands that if it's in Amazon, it's gonna spin up the services that it knows. It needs to spin up the, you know, in Amazon.
Mm-hmm. But then similarly, it understands that if it's in an air gap environment where there's no access to anything, it's gonna spin up, for example, uh, uh, uh, my sql because it knows that I need, I need my sql, I'm just gonna spin up the pods. And that just a, you know, that's a health chart thing, right?
So we, we know exactly what's going on. Thanks. Yeah.
Absolutely. Such, such interesting questions. Should I go back to the deck?
I got like 70 more slides. Why not Take one? Okay.
Uh, slide there. My, oh, there. New slide every 15 seconds.
Okay. So I'm gonna skip a bunch of these things, but, but I think we've already discussed this, right? The whole point of standardization, et cetera, was, Hey, look, if you do all these things, cell service becomes very easy, uh, because then you actually can reliably let people press buttons, because the buttons will always mean the same thing.
So then you would be comfortable giving themselves service access. And that's when people ask us, Hey, can I do this for ai? The Q flow story that I was telling you earlier?
Mm-hmm. Sure. And that became a whole business for us, and that led us to our, our friends at nvidia.
You know, I mentioned earlier that, uh, uh mm-hmm. You know, there's this content on the website about us. We, we publish a reference architecture with them that talks about this notion of a GPU platform as a service.
GPU has, uh, at present, we are the only company for quite some time now that has a reference architecture at that layer with Nvidia. So that gives us a lot of advantage because in the AI space, RAs are, you know, they're the gold standard. If you have one, you're good.
If you don't have one, go fight it out. Right? If we have one.
Mm-hmm. And definitely the blog is, is, is very good. I mean, it actually, let's go back to the blog for a second.
Uh, so the blog, the reason why I like it a lot is because it focuses on self-service. There's clear alignment that this is the real problem, doesn't matter. G-P-U-C-P, of course, GPU, there's very few of them in any one environment.
So you wanna do a better job of that, right? You don't wanna waste time. But this is the problem to the prior point that, you know, we were talking about, right?
This is the problem. Everything steps for this. If you solve for this.
And the, of course, the supporting cast of other technologies that need to be in place to deliver self-service consumption. Everything is fun. This is it.
Why do you think this problem exists? I think that for the longest time, people have gotten obsessed over the tech, right? And we've forgotten we've forgotten the, the purpose of it all.
The pivotal, I think Pivotal is amazing. Pivotal. Pivotal, pivotal is ama.
So, you know, I, earlier I showed you an example of somebody who deploys Java Code MySQL, you know, who, you know, who does that? Pivotal does that They've done it forever. Don't exist anymore, unfortunately.
Yeah. Right. So our perspective was look pivotal, you know, because they existed before Docker and Kubernetes and all these, so they had to invent all these things in a very, in a very proprietary way.
Mm-hmm. What if we could reimagine pivotal in for the new world? I mean, in theory, I mean, not that we, we, that was not the, the, the sort of the problem statement in our minds at the time.
But that's, that's a very good way to think about the problem, right? They did it right. Yeah.
I have a different theory, which is that so much of the AI marketing and discussion and strategic focus has been around training foundational models and delivering them at scale. When I would think that, you know, 50%, 60% of the activity right now is experimentation and small models and messing around with stuff that does not lend itself to The training guide. Well, the, well, when you're training foundational models, you can, you, you know, you, you can start building out the service and this team and all that sort of stuff for three months before you start revving the next model and all that other sort of stuff.
You don't need the self-service. And This, you don't need any of this because you have the right people. So, But I don't, that's just, that's, that's so what your, that's like a hobby horse of mine is, But that's a, that's a very sound point, right?
Because most people, a year, a year and a half ago, I mean, look, people would talk about this, right? People would say, every company's gonna build a model. Why the hell are you gonna, are you gonna build a model?
I don't understand. But that didn't work out. Yeah.
Build your own model instead of tune it or build an application on something that exists. Uh, add rag or something. I, I did a one of those flash presentations here about how there's actually three different avenues of AI development.
Foundational model creation is only one of them. So now people say inference is going to be the, is the going to be the 10 x use case? And we will see, uh, but definitely right now, the biggest use case is experimentation.
Mm-hmm. Where people start small and cell service is key. I think that, um, look, I mean, I tie this back to the last five years of sort of my life, right?
As I've been outselling product, right? People get caught up with the, with the tech man there. People are enamored with the tech.
Um, January of 2024, I did the presentation for our board. The first slide in the, in the presentation said, Kubernetes is now furniture. Nobody cares anymore.
There was an expert in there as well that I just removed, right? And the point of that slide was, um, I think that the world's gonna change very fast. This is 2020.
So basically a year and a half ago. I think that, um, we need to think beyond just Kubernetes management. We need to think about what are the use cases, and we need to do this fast.
And we did, which is why we are here. So we basically reinvented ourselves like four or five years into the company. Like, or we could have not, and we could, we would still be where we were.
But because we took a pause and we thought about, okay, look what is the real problem here? We started this company thinking about self-service. And all we talk about these days in our company is Kubernetes management.
And the better c and i, and this is not it, this is not why we started this company. Let's just take a pause. Let's take a step back.
What is the real issue? That's when we started building that service layer I talked about earlier. And because of that, we are here now because we remembered the original purpose.
The original purpose was consumption, self service consumption. You do that well, everything is fine. And this block talks about that in a very elegant way.
I definitely recommend you read it. Uh, it has, it has that diagram we looked at earlier for quite a while. There's a, somewhere here, it's a thought.
There was, oh yeah. There's a customer reference, right? So it's, it's pretty cool, right?
So we use this a lot, uh, as we talk to customers, right? Having, Hey, did you see that blog about RA on side? We say that a lot on a daily basis.
Um, but, but, but they, they do this, uh, because they understand the problem. Like, I really do believe that, like, of course they, they have the biggest market, uh, uh, uh, sort of, uh, percentage, right? So they are the kings of this market.
They see it all day long. They understand that consumption is key. Mm-hmm.
Uh, they do s et cetera. And that's where we're seeing our business. You solve for self service consumption.
Everything is great. Sorry, go ahead Gina. I'm sorry.
Um, where are you getting your customers from? Like what, um, verticals or, or what, what kind of, um, what kind of, uh, things are they producing? Right?
Right. What customers are coming to you? So the, the GPU cloud market, as it's called right now, ma'am, seems to be very active.
So people who have actually, a bunch of them happen to be sort of crypto monitoring companies who have now become GP providers. Uh, they are looking for software, uh, to deliver high value services to enterprise customers. Mm-hmm.
Uh, there are data center companies who have also invested in this globally, uh, who are doing this. There are now government funded, uh, sort of organizations globally who are deploying GPUs for their local regions. So that's where we are seeing a lot of uptick, a lot of uptick on the enterprise side.
We see two, uh, sort of sectors where there's more activity than others. And this could be anecdotal. This could, this may not be represented at the market.
Uh, financial services and pharma, pharma, healthcare, life science mm-hmm. Bucket. We see a lot of AI activity in those two segments.
Yeah. Finserv is like, it's like everybody's doing something and they all have the money to invest. So, uh, we whiteboarded this with, uh, some friends of ours on the dev rail side at Nvidia, where in the early days when I was trying to explain to them, 'cause they were asking similar questions, but what the hell do you do?
Right? I thought I already had Kubernetes solved for, and that company looks cooler and I like cross plane. And I've had this conversation many times with them, and I was trying to explain to them the problem is not the environment when anymore.
The problem is the governance layer. You solve for that. You solve everything.
So we've also talked about this, and I'm happy that all those slides that are gonna come up, we've already talked about in some fashion, but these are just representations of the problem. How do you run mm-hmm. A cloud, be it an enterprise, or be it a company that is gonna sell to multiple enterprises.
Mm-hmm. You really have to think about tenancy in a different way. And at each level, you better be able to set policies, quotas.
When you, when you set up an account on Amazon, by default you have a policy, you will not have more than this much compute. If you are, if you're a free account holder, you can only do T ones, for example. Right.
It's a policy and then somebody goes and changes the policy. Mm-hmm. You have to file a ticket for that, by the way.
Right. We should be able to do the same. We, we can't.
We were doing this for large enterprises anyway, again, because we built this for this crazy large enterprise use case. It just so happens that, you know, in the GPU space, they sort of look the same. So, you know, very blessed and lucky, uh, to have been pulled into this market with a product that meets this need.
And the alternatives out there don't seem to have thought about multi-tenancy in this fashion. The only company in our journal space who understands this notion of an MSP is actually VMware vCloud Director. Great product doesn't do this.
Mm-hmm. Not that, I don't mean the tenancy part. I mean the AI use cases is one or not.
You, You mean AI cloud director? Yeah. Yeah.
AI cloud director. AI cloud director. Don't give them ideas, man.
They'll launch it tomorrow. No, they, they've already gone. They have one.
They are, yeah. They already call it that. Oh My God.
Okay. All right. All of us should have seen that coming.
Right. We talked about this slide as well. So when you're box six network segmentation, there are you segmenting the front and the back networks?
When you say back, what do you mean? GPU networks. So the GPU networks, I have a, Oh, And this, this is better.
This is a better distinction. Right. So, so this is a very, very dumb diagram.
This was Whiteboarded. Uh, I happen to be in India with a, with a specific, uh, uh, solution architect from Nvidia who I, who will know who he is when he, when he hears it, hope he does. And he whiteboarded this in their office in, in Bombay.
Uh, and then I made a slide out of it. I should at one point make it better. But the point of the slide is that there are many things an orchestration system has to do.
Right. Two of them are the north, south, and east west network. Mm-hmm.
So what you call the back backend, which is east west, the GPUs need to talk to each other. Correct. It's, it's the east west that nobody sees.
So the UFM layer at the bottom of the screen mm-hmm. Is that, so UFM is the endpoint in the, in NVIDIA land mm-hmm. For InfiniBand configuration.
So I, if I need a new P key for InfiniBand, I talk to UFM. So we have, we have a solution for that. We talk to UFM and it, it is gonna, it is gonna create a new peaky so that your, your, uh, GPUs can sort of interface with each other mm-hmm.
In a mesh. But similarly, I need to have a, a north south network. Mm-hmm.
Right? So I need to VLAN for that, and then I need to make sure the servers are belong to the right vlan so they can all network with each other. And That north south feels like a solved problem.
So I, I have a good sense of how that's done. Same problem actually. Right?
They both have the same problem. The issue is the automation. Right?
Right. So we've built now a nice integration with our friends at Nvidia. We do this for the other vendors, at least in the infin space.
We only, only Nvidia on the ethernet side. You know, we have a solution for Dell switches, Fortinet switches, et cetera, and we keep adding to our library. Mm-hmm.
So that over time, as more and more different options come our way, um, the next guy is gonna get it for free. The first guy, if you gotta build something together, right? Next guy's gonna get it for free.
Okay. Uh, and similarly, if you look at this picture, um, so we talk about switching just now, B-C-M-B-C-M is NVIDIA's, uh, BCM. So this is Bright Cluster Manager.
So if you are an NVIDIA certified, uh, uh, NCP is what they call it. So Nvidia Cloud partner, uh, which means you're GPO Cloud that follows their reference architecture for GPO clouds, uh, then usually you will end up using BCM for bare metal provisioning and server for server management, if you will. Yep.
But we support that as well if you don't want to use it, which is fine. There are multiple options for bare metal provisioning such as, uh, bmas or Ironic. So our team seems to like ironic a lot, but Mass is a great option as well.
But these are some examples of automation, uh, uh, sort of touch points, right. That one has to think about when one is building a network mm-hmm. Including the billing ing on the top left.
We must have some way to help people. It costs this much money. Mm-hmm.
So this is the sort of the final outcome, right? So this is the stack. You've seen variations of this already today.
So that Nvidia picture is, is a, is is a variation for sure in some way. Uh, but this is how our product management team thinks about the problem. Um, so they think in terms of sort of cloud management and cloud services, the cloud management is the network, the servers, you know, the VPC, all of these things have to be thought through, uh, somehow.
And then the cloud services could be everything from the compute services, VMs, Kubernetes, et cetera, to the tools themselves. There's a gray box on there. It's gray because we don't have it today to sell, hopefully.
So, so the team believes Q4. We have a, a sort of a very nice fine tuning service built into the platform as well. Today we have an inferencing endpoint.
Mm-hmm. Very good. Um, don't have a fine tuning product.
And you had a couple of bugs in the earlier slide under infrastructure. One was edge and I don't remember what the other one was. Data center Edge remote cruise ship.
No, it was, yeah. Should I, But, but, but this was the other for them, other Picture is, is this picture? Yeah.
Yeah. It was just remote. Yeah, remote.
Okay. Yeah. All right.
I don't know that those are included in the one you had there. You just, It just, yeah. The, what Is that, that OCI or something?
The red cloud? The red one is Oracle. Yeah.
Yeah. OCI. Yeah.
Just, uh, one quick question. The one thing I don't see provisioning for is memory Provisioning for memory. So memory in the AI cases is going to be sort of part and parcel.
Uh, so there'll be some memory, you know, sort of on your memory or memory on the side. So in the VM case, there is a configuration when you, when you, when you, you can have a pre-packaged solution. So you can say, I'm going to sell two virtual cores and 30 gig, two gigs of ram.
And we will create that on demand. Or you can say, let my customer choose. But we provide actually both options.
Okay, great. Yeah, we must, because, uh, most people actually don't want to sell a mix and match because of been packing problems. So what they do is actually they'll create small, medium, large configurations or whatever, explorers, et cetera.
And they'll have clusters dedicated to only small cluster for only medium cluster for only large bin packing. You don't wanna waste any capacity. So if you may let people pick and choose, then you'll have fragmentation issues.
Okay.