Bridging the gap from GPU-as-a-Service to AI Cloud with Rafay
Rafay CEO Haseeb Budhani argues that to truly be considered a cloud provider, organizations must offer self-service consumption, applications (or tools), and multi-tenancy. He contends that many GPU clouds currently rely on manual processes like spreadsheets and bare metal servers, which don’t qualify as true cloud solutions. Budhani emphasizes that users should be able to access a portal, create an account, and consume services on demand, without requiring backend intervention for tasks like VLAN setup or IP address management.
Budhani elaborates on his definition of multi-tenancy, outlining the technical requirements for supporting diverse customer needs. This includes secure VMs, operating system images with pre-installed tools, public IP addresses, firewall rules, and VPCs. He highlights the difference between customers needing a single GPU versus those requiring 64 GPUs and emphasizes that all necessary networking and security configurations must be automated to provide a true self-service experience.
Ultimately, Budhani argues that the goal is self-service consumption of applications or tools, not just GPUs. He believes the industry is moving beyond the “GPU as a service” concept, with users now focused on consuming models and endpoints rather than managing the underlying GPU infrastructure. He suggests that his company, Rafay, addresses many of the complexities in this space, offering solutions that enable the delivery of applications and tools in a self-service, multi-tenant environment.
Presented by Haseeb Budhani, CEO, Rafay Systems. Recorded live on September 10, 2025, at AI Infrastructure Field Day 3 in Santa Clara, California. Watch the entire presentation https://techfieldday.com/appearance/rafay-presents-at-ai-infrastructure-field-day-3/ or visit https://rafay.co/ or https://techfieldday.com/event/aiifd3/ for more information.
Transcript
So I'm gonna shift to the, the gap. Now, we've talked about it at length. I'll just kinda structure this with slides so our friends online can also, uh, uh, sort of consume the same information by the, I'll repeat this year.
My name is Hasi Bani. I'm the CEO Raffe. Thank you for joining us.
Uh, and let's talk about the gap. So, reating the same point, right? What is the purpose?
The purpose is, you know, be like CSPs. Be like, Mike, if you can do that, you can make money. So, okay, this is what I promised the other day when we spoke, right?
So what does it mean to be a CSP? So this, this slide, I will tell you, I've had two hour conversations on with people. This is a, this is a contentious slide for many people.
Uh, because basically what I'm claiming is that if you cannot deliver true cell service consumption, you're not a cloud. If you, I cannot go to your portal. Mm-hmm.
Set up an account, get access, and you can stop, you know, control my access because you need to do some back channel. By the way, we could support that too, right? So there's a concept called KYC, know your customer.
Mm-hmm. Right? Where you have to validate that this is not some, you know, unknown or Chinese entity.
Um, yeah. You can go do that process in the back till then. We'll hold them.
They can look around, but they can't consume anything. And then once you say in the back system, good, we will let them in. Mm-hmm.
But it should be self service. If you cannot do true self service, I am arguing that you're not a cloud. Agreed.
Good. I would, I would add that's even happening in the rack, rack scale, compute space as well, right? Yep.
The next generation is happening. There's not go wrong. Fabrics and infrastructure right?
At rack. GP clouds have spreadsheets that they maintain. Doesn't matter what their websites say, it doesn't matter what their blocks there.
Spreadsheets. Yeah. But I'll, I'll make, I'll take it one better.
Right? So Beal service is one thing, right? If you don't sell me actual applications, like an experience, not a cloud, right?
And this is where people fight. People fight for this now because all my customer wants are bare metal service. Okay?
Alright. That's a great business. Making money.
I don't my my definition, sorry, you're not a cloud. You gotta sell an app, a tool, a platform, whatever words you wanna use, you gotta sell one of those things. A developer should be able to consume it, or a use case they have Bare metal is not it.
'cause then somebody in ops or dev, whatever, somebody's gotta take it from here and make it ready for a developer. No. Now that may be a use case and you should service that customer as well, but you must be able to service a developer finding words.
And the third one, multitenancy. Multitenancy means many things in a, in, in two slides. I'll, I'll, I'll explain my definition of multi-tenancy at a high level.
So we all agree with these things. This is just to set up, uh, particularly for as, as we look at the, the product demo. Well, I, I quibble With things in the backend, like multi-tenancy.
Yes, sir. Um, I think it's, uh, it's really about the, the self-service on demand experience more than anything else. I mean, uh, uh, for Azure really came around for many years.
M six, M 365 was not multi-tenant. It was, And it worked just Fine. Yeah.
And it worked fine on demand. So, so I think that Single tenants has at that time. Yeah.
Yeah. Which is completely okay. So, uh, so long as you can hide it, I guess is the answer, right?
Question is, did they hide it? They had it, they hit it very well, right? Because I log in, I do whatever the hell I want to do.
Mm-hmm. Right? Mm-hmm.
It's not my problem. See how it delivers. Well, it's intentionally o it's intentionally opaque.
Yeah. Because it was better for them at the time. Mm-hmm.
And then over time they made it multi-tenant because multi-tenant over time is cheaper, actually That Right? That was a benefit to them that they could pass on. Yes, sir.
So that's a margins optimization. That that's, that that's their problem. But from an experience perspective, I'll just switch technical slides.
I'll go back to them later to make the point that we're discussing right now. Right. So here's my definition of multi-tenancy.
So there's a lot of things that happen under the covers, and let's talk about them here. Somebody comes and says, and I, I used this example before, uh, I want 64 GPUs or whatever, you know, slinky and Kubernetes, uh, or somebody else comes and says, I need a single GPO man. I just started my journey.
I don't even know what the hell this is. Just gimme a milk flow. Mm-hmm.
And I'll take it from here. At least I'll read some blogs. Right?
Okay. So now what needs to happen for both of these customers to be happy in a self-service way? Now yes, we can create two different accounts for them.
Absolutely true. Right? So both companies have, or other, each company has its own account and environment.
I'm wasting money, but it's my problem as a provider, not your problem. You get a great experience. Right?
But now let's look at the second level, right? So for me to give somebody, let's look at the left side, then we look at the right side, right? So for me to service a customer who wants a single GPU, what needs to be true?
So let's start at the bottom. You need to create A secure VM that, Ah, my, my friend's already there. How do I get there?
Right? So I got a server. Let's say it's got eight GPUs.
You want one, not eight one. Mm-hmm. I don't wanna give you one and then waste my seven.
That's bad business. Or maybe it's, maybe it's again, my problem. Maybe I, I'm okay with this seven.
Yeah. Okay. I should have some vm some conclave.
Right? So I have some, some hypervisor layer on top of it. We can deploy V vm.
Let's argue it's KVM. Mm-hmm. Okay.
Then I need an OS deployed. Uh, the OSS should have the image so that the tools that I need should be deployed. Mm-hmm.
Thank you. Right. Otherwise, what is the point of this?
Right? Right. Uh, okay, then I need to access it.
I'm not in your data center. I'm at home. Mm-hmm.
I need a public ip. Where's that gonna come from? Who's gonna manage my I iPad for public ips?
And then that public IP hits a firewall. Mm-hmm. And then I need rules so that my traffic can go to my VPC, where the VPC come from.
Who is gonna do this for now, all of this is inclusive in the definition of multitenancy. If you don't do all of this, this is not a network. This is not a cloud.
If I gotta ask somebody, Hey, can you go set up a vlan? It's not a cloud. Again, my definition, my friend, we made this No, no.
Nolio no, I I I get where you're coming from. And that actually refers to the original definition of cloud. Mm-hmm.
Before all of this Agree Salesforce and Gartner and everybody started stepping in. But, uh, I think the, the, the focus on the, on the tenant experience is the, is the important part. Absolutely.
Upscale, upscale down. I'm, I'm in control. I'm defining the environments I need.
Self-service. Self-service. Yeah.
Self service, multi-dimensional self-service. All of this is to serve one goal, self-service. Mm-hmm.
The goal is self-service. Mm-hmm. But to get there, these other things must be true.
Right. And because self-service, what I need to do something, I need an app, or I call it apps, by the way. Sometimes my friends at NVIDIA say, nah, not everything is an app.
It's a tool. NIM is not an app. NIM is a tool.
It's an agent. It's an, oh yeah. Some things are agents.
Oh, don't say that word. So in this room, Gentech, I like that even better. I was about to tell you to serve at Accenture, but, And They're Gentech framework.
I, but, but I learned to that. But see, the same problem applies on the other side now, right? So if somebody says, I want 64 GPUs, which usually means eight servers, these eight servers somehow have to exist.
They need to be networked together. Yep. Yeah.
So I need to assume for a minute, this is an infinite environment, right? So I need a new, new PQ for this tenant. I need to tie the ERES for all these together so I can mesh the GPUs.
Mm-hmm. Then I need a new vlan so I can do east, uh, north star traffic across these servers, because otherwise it's an on network. And then all of them have to be, uh, uh, you know, in their own quote unquote, uh, VPC, but VRF, right?
And then I need to get access to these servers. Again, same problem. I need IP addresses that are public so I can hit them and blah, blah, blah.
All of this must happen, you know, but the hyperscalers do this today. We don't even think about this anymore, right? I just get an EIP just magic.
Right? Right. But now our new friends building GPU environments, I have to take a second to think about, what do I call it?
A cloud or not. GP environments must do all of this work. If they do not, it's not a cloud.
And While you're at it, please make sure nickel or rle works and the collective infrastructure and libraries, I gotta show a demo another day. I gotta show you demo another day. Okay.
So my, my, my colleagues have done such an amazing job with some of these problems. Uh, of course we don't solve every problem in the stack, but we solve a very large percentage of this. Uh, and I'll talk about that for sure, uh, before we, we part paste today, right?
But this is, this is the, the definition of multi-tenancy. To go back to this, but I think this was a good, uh, question asked at the right time, because now that we agree on this, we go back to the primary thing, right? The goal is self-service, right?
Self service consumption, but self service consumption of what? Some app, I wanna go to a portal. I want an endpoint.
I just wanna consume a model. Mm-hmm. I, all I wanna do is try that.
I just gimme some, gimme some curl example. I just wanna hit this end point and ask a dumb question. I think you hit it right in the title it's apps or tools.
It's an endpoint, it's a model. Yeah. Jupyter Notebook use case, right?
Let's call it a use case, right? Depending on your need. Juniper notebook, I'll show you a demo of that before the end of the day, right?
So this is what we believe is the right answer, and we believe this is the opportunity. Now I'm using our product screenshots to show you what our product does. Mm-hmm.
So clearly the point here is we do these things, and of course we'll go into this in a minute. Say that if, if, if the, if the goal is to actually have, uh, apps and tools available, and you're not really caring about GPUs, quote unquote, then we need to throw out GPU as a service. So I would agree, ma'am.
Done? Yep. So I would say that, look, so when, so there's a phrase CP as a service.
Have you guys heard of it? Mm-hmm. You have.
I just made s**t up. Right? It's not, nobody says that.
And nobody says CP as a service. Why don't they say CP as a service? Nobody ever said that I can get a m server that nobody cares with that.
Right? I, I can get a vm, EC2 happens to have X 86 in it. Oh.
But There's everything as a service There, so, sure. Right. But the point being that we created a, a new concept called GPS service to your pan map, which is, uh, which was the look, which was the need for the time because the market was so early.
Mm-hmm. Right? But the market is actually maturing pretty fast, actually.
It's quite amazing to see the things people are building, uh, in enterprises as quite fascinating microservices. Now, consistently, we see people have some basic ragging happening in internal applications, actually pretty amazing. Mm-hmm.
Um, okay. But none of them are thinking about the GP itself. In fact, I just care about the model that I heard about, uh, the fact that, uh, you know, the, the full fledged version of DC R one requires eight GPUs is not my problem.
But the, but the, what do they call it? The, uh, what's the, what's the, what's the short, the reduced version of DC called, what's the word? Distilled.
Distilled. I was thinking rated, which is obviously wrong. Distilled.
Yes. The distilled version runs on a single GPU. Yep.
That's not my problem. I just used it. I used the model right.