Accelerating AI Infrastructure Adoption for GPU Providers and Enterprises with Rafay
Haseeb Budhani, CEO of Rafay Systems, begins by highlighting the confusion surrounding Rafay’s classification, noting that people variously describe it as a platform as a service (PaaS), orchestration, or middleware, and he welcomes feedback on which term best fits. He then pivots to discussing the current market dynamics in AI infrastructure, particularly the discrepancy between the cost of renting GPUs from providers like Amazon versus acquiring them independently. He illustrates this with an example of using DeepSeek R1, highlighting that while Amazon charges significantly more for consuming the model via Bedrock, renting the underlying H100 GPU directly is much cheaper.
Budhani argues that many companies renting out GPUs are not true “clouds” and may struggle in the long term because they are not selling services on top of the GPUs. He references an Accenture report suggesting that GPU as a Service (GPaaS) will diminish as the market matures, with more value being derived from services. He emphasizes that hyperscalers like Amazon have understood this for a long time, generating most of their revenue from services rather than infrastructure as a service (IaaS). This presents an opportunity for Rafay to help GPU providers and enterprises deliver these higher-level services, enabling them to compete more effectively with hyperscalers and unlock significant cost savings, citing an example of a telco in Thailand that could save millions by deploying its own AI infrastructure with Rafay’s software.
The speaker concludes by emphasizing the increasing importance of sovereign clouds, especially in regions like Europe and the Middle East. Telcos, which previously lost business to public clouds, now have a renewed opportunity to provide AI infrastructure locally due to sovereignty requirements. He states that Rafay aims to provide these telcos and other regional providers with the necessary software stack to deliver these services, thereby addressing a common problem across various geographic locations. He highlights a telco in Indonesia, Indosat, as an early example of a customer using Rafay to deliver a sovereign AI cloud, underscoring the growing demand for such solutions globally.
Presented by Haseeb Budhani, CEO, Rafay Systems. Recorded live on September 10, 2025, at AI Infrastructure Field Day 3 in Santa Clara, California. Watch the entire presentation https://techfieldday.com/appearance/rafay-presents-at-ai-infrastructure-field-day-3/ or visit https://rafay.co/ or https://techfieldday.com/event/aiifd3/ for more information.
Transcript
Hey guys, my name is a nice to meet you all. Thank you, Alistair for a minute. He said, this is not an ai.
And Friday, I was very worried. I don't know the difference. I didn't even know you had so many.
Um, I presented at a tech field day. The first time I did that was 20 12, 13, 12. A long time ago.
In fact, when, uh, Ethan, uh, and Greg used to have their podcast way, way backpack pushers, uh, I did episode 61, maybe. And I remember talking to Ethan about, you guys should, you know, there should be some charge for this, right? 'cause I, I represented a company and we had a call and, and we agreed that we'd pay five bucks for the podcast.
And I think you guys charge a little bit more than that now, but yeah, a little bit more. So it's been a, uh, it's been a while since I've been a, a follower, uh, at least the various podcasts now. There used to be one.
Now there's many, many of them. Uh, so I do listen to them, to the, to the best of my ability. And I still find that some of the things that, um, I used to work on when I was a real engineer.
I get to learn from you guys when I listen to you guys. So thank you very much, uh, for the, for the community that, uh, continues here. So today I'm gonna talk about, uh, Raffe.
Uh, I'm the CEO here. Rafa is actually, I'll tell you the three words people use and my request for the computer is in our closed session. I'd love for you to tell me what you think is the best one.
So depending on who we talk to, maybe four. Some people call us platform as a service. Uh, uh, in fact, we have a reference architecture with Nvidia that calls us a GPU platform as a service.
Uh, many of our customers, once they understand the platform, they say it's orchestration. I guess at that level, a lot of things are orchestration. And the third phase people use for us is middleware.
Um, we don't know which one is the best one. So depending on who we're talking to, we use different words. Uh, but it would be very nice to get your feedback.
I think this is kind of something that is top of mind for me and I'd love to get your thoughts on this as we proceed today. So we'll talk about a few things. I wanna start with, uh, just how we see the market, uh, the specific area that we focus on.
Um, we talk about each one of them. Um, uh, we'll talk about sort of in this market, the, the gap. And as we talk about the, the market dynamics, the, I think, I hope the gap will be, uh, obvious to you.
And we can spend time when we talk about a bit about Rafa, uh, about what Raffi does for set gap. And then we'll go under the hood to the point you had made the other day. And I'll try to get there as quickly as possible.
Um, and then we can talk, right? Um, I get very distracted if I have slides to present, like really distracted, you'll find me d joining on. So please, uh, let's talk.
Yeah, let's, let's please let me not present slides all through, but I'll start now with the market dynamics and to, and to tell the story of the problem that we see. I spent two weeks thinking about a problem, and then I had a five minute conversation with chat GPT. And actually it was amazing.
I should do this all the time. Um, so what I was thinking about was, let's say I'm building an app. I'm building a travel agency app.
People will come and they'll ask the question, is my flight delayed? Or, I need a new flight. Simple stuff.
Very simple stuff. So HR GPT, I wanna build this and I think I need a model for this. I'm hearing about the gen gen AI thing.
What should I use? I said, well, you have a bunch of options. You could use deeps seek.
Mm-hmm. Okay, what's deep seek? Well deeps seek, there are many options.
Deeps seek R one. There's all these different things you can use, but you know, we recommend you use deep R one. Uh, and, uh, you can use it on Amazon.
This is Charge Beauty ra. Mm-hmm. Okay.
Alright. And then, uh, ask you, okay, so what is it gonna cost me? Like, well, it depends on how many interactions you have per day.
I know, let's call it a hundred a minute. That seems like a reasonably healthy website. So it said, okay, based on what you just told me, if you use Deeps seek in Amazon, the cost of consuming deep seeq is, you can see it on the screen, right?
61. That's the cost. If you use Amazon to, to consume this model, redrock gonna cost you.
And let's, let's pick the number in the middle. $6 and 80 cents, just medium, whatever that is. Okay?
So I gotta pay Amazon seven bucks an hour to consume this model. Then I'm gonna use a hundred times a minute. Okay?
I'm using an H 100 for this from, uh, Nvidia. So next question I've asked, it was, what does it cost for me to just get an H 100? I don't wanna use Amazon.
What if I could just get the GPU? Mm-hmm. So deep Seek said, well, you can, it pulled up G list AI essentially, and said, if you go to gps ai, you can rent H one hundreds for like a buck 60.
You gotta buy more than one, obviously, right? You gotta build an app. You're not, obviously nobody's gonna run a business on one H 100.
They're gonna have more than that, right? Okay. So it costs like a buck 60 to rent one for an hour.
And Amazon charging me seven, eight bucks, 10 bucks. Thanks. Why?
And that's the opportunity. That's the opportunity, right? All these people building data centers, all this power stuff.
And we talk about cooling, and I've now seen a few of these, these, uh, these liquid cool data centers look pretty awesome making this kind of money. H one owners don't need liquid cooling. Obviously those are the gps, but, you know, why are they not making money?
Any thoughts, by the way? Any, any, any, any assumptions on this? Too hard to get started and use too hard to get started.
Yeah, absolutely. That's a, that's really important. Yeah.
But this is where we found a spot for Raffi. Mm-hmm. And I'll talk about that at obviously at length as to what is the issue, some, some more and, and what do we do about it.
I think that the biggest problem right now in the market is there's a bunch of companies that have been started who may call themselves clouds. I would argue they're not clouds in a little bit. I'll define my, I'll give you my definition of cloud.
I told you that the other day. I'll do that today. And I believe that there's a journey that these companies are gonna take from here onwards, where they're gonna recognize that just rent out hardware is not a great business.
So either they will evolve, do the hard work mm-hmm. Or they'll die at some point in the next four or five years. It's too expensive to run this stuff, right?
And the margins they're making are not that attractive right? Now, if your landed cost to deploy this hardware is about a buck 50, a buck 50 cooling power, everything in the us and you're charging a buck 60, it's not a great business to be in. But Amazon, they make money because they're selling bedrock.
They're not selling a GPU, they're selling a service, SageMaker, whatever, right? So many things that Amazon does. And by the way, when I say AWS or Amazon in this context, I don't mean Amazon.
I mean it's like, you know, internet search is now Google, right? You know, tissue paper is Kleenex, right? So I'm using them, you know, sort of across the board.
That also includes the Googles and the Azures of the world. All of them have very incredibly well thought through services. So this does not, this applies to all the hyperscalers.
The question is the hyperscalers figured this out a long time ago. And some of our new friends in this new community have not figured this out. So I love this graph.
Uh, this is the only time I'll, I'll reference Accenture I think in this presentation. Uh, otherwise you'll probably kick me out. But, uh, this is an important point.
Uh, Accenture as a, as a, as a side plug is, is a, uh, is a strong partner for ours. We have a global teaming agreement with them. So any sovereign AI deals that Accenture's working on, they take us along.
So we're very thankful to them for, for the business they bring us. Uh, but this point they make has nothing to do with the fact that they, they, they, they generate revenue for us. This is actually a really important point.
Yeah. GP as a service over time is gonna go away, or it's gonna ask them to, to some percentage, which is now gonna be 50% or 70% can be a much smaller percentage. This seems so obvious, but we're not seeing this happen in the industry.
And I think it's a reason because of what you said. It's just really hard. In fact, this point, very specifically, the hyperscalers have known for a long time.
So this is an IDC report that talks about where the money comes from for hyperscalers. So look, look at, look at the I line number less than 20%. Makes complete sense.
Where's the rest of the money coming from? Not is services. 'cause if people start looking at, I can get an easy two instance and I can deploy this and I can deploy this and I can deploy this in theory, it'll be cheaper.
It's actually not cheap at all. It's really expensive. And Amazon understands that.
And that's why they say, Hey, you want a hundred requests per, per minute on deep seek eight bucks. Turn it on. Right?
Tokens, right? So, and token math, I don't understand. I don't know if you understand, but nobody can figure out how much a token costs all magic.
But this, this is the opportunity. So the question then becomes, how do we help both the providers that are investing in GPUs sort of look like a public cloud, a hyperscaler. If they can deliver similar services, maybe they can make more money.
The answer is yes, they can make money. And similarly for enterprises, 'cause it's not just a a, a hyperscaler or cloud problem. Many large enterprises are also investing in buying in their own infrastructure sovereign reasons at scale.
It's just cheaper. And the bigger problem is, how do I do it? In fact, I'll tell you a great story.
So, um, we are working with a, a customer in Thailand. Uh, it's a very large telco, similar ideas. They wanna build a GP cloud.
So many telcos are building GPU clouds these days. So their chairman asked a very simple question, if everybody's doing ai, we must also be doing AI in our company who find me teams that are working on AI today. Great point.
So they went around and asked the question, who's using Gen AI today? They found many teams. One team was very interesting.
So they found a marketing team that puts together these marketing campaigns, uh, and they, they run these ads globally for people who are traveling into the us I'm sorry, into Thailand so that they can buy phone services, et cetera. These, this team said they're using OpenAI, that G PT basically, right? So they have an enterprise agreement, they pay a bunch of money.
So they're paying apparently over the next three years, five plus million dollars, given the scale that they think they will need 5 million bucks for three years. So we did an analysis, what is the capacity that you actually actually need per minute? We did this with Accenture, sorry, second time I mentioning Accenture, so sorry, two no more.
So we decided to go, uh, get a box from hp. Our friends at HP gave us a box with eight H one hundreds. We deployed Rafi software on it.
I'll talk about the software. Of course we put Quinn on it. Quinn is an open source model, is very good.
And we delivered the exact same experience to the customer as they were getting from Public Cloud. Now, we're not saying don't use Open. Of course Open is a great company, use it all day long.
But in terms of costs, right? So the box physically boxed from hp, if they just bought it, it would cost $240,000 USD one time pay and support, whatever, right? Our software for agpu, $10,000 a year, something pick a number or three years, what is that?
$280,000? Whatcha you gonna pay? OpenAI 5 million bucks.
20, 20 times. 20 times. That's the opportunity.
That's why enterprises, when they look at this math, they go, chip, man, if I do this myself, if I can afford to do the upfront investment, and if I can build it, this is better for me. Right? And if I don't have the data center capacity, because all those people left like 10 years ago, uh, okay, I can go to a local regional GPU cloud and I'll get the same experience.
But if I go to the GPU cloud and I get a bunch of H one hundreds, then I gotta deploy software and get coin working and blah, not gonna happen. I'm gonna go to Bedrock or I'm gonna go to OpenAI. This is the opportunity.
Coursely. Do you agree with this or are we already in violent disagreement about something? No, A hundred percent.
It's, I live this. Yeah. Picking It up from ground and running it Is an order of magnitude harder than people can tolerate.
Yeah. Well, I hope we continue to agree for the rest of this conversation, but look, everybody wants this, right? Um, I just wanna platform whatever platform means, which is why, by the way, PAs, right?
That's why the, the, uh, the RA we have with Nvidia, or rather blessed by Nvidia, uh, calls us a GPU pa. But look, somebody's gonna want to come and say, I want, um, I don't know. I'm gonna train stuff.
So gimme s SLM and 120 GPUs, I'm 20. GPUs is a bunch of servers. It's not one server.
Uh, but now, I don't know, I'm just starting my journey. I don't even know what Gen AI is, but my boss makes me do this stuff now. So, okay.
I guess I just need one GPU. Mm-hmm. Okay.
That's a VM actually, right? I mean, that's, that's not a server. There's no, I mean, you, I guess in theory you could have a server with a single card.
Nobody does that. It's a, it's a portion of a server. So, and somebody's gonna come and say like, last, very, very late last night, we had a conversation with the customers.
Like, my customers want Q Flow, dude, show me one customer wants Q Flow, but okay, all right. We can do that too. Right?
So people will come and we'll ask for many different use cases, and somehow we have to deliver these use case, right? And this is a daunting challenge for us as a company as well, because of course, right? I mean, it's, it's an early market.
We're all learning together. Mm-hmm. Uh, we've done many things already.
We have many, many things to do over time, but the problem we're trying to solve for is, and if I can just make the people, the developers right, who come and use these platforms, this means their life better. Mm-hmm. Right?
Or good, good enough such that they may at least give this internal platform a shot. I could do that. It'll save people a lot of money.
Mm-hmm. Right? And, and maybe they'll pay us money as well.
Right? But this is the, this is what people seem to be looking for in the market. And again, AWSs the, the Googles, the Azures, they already do this.
I mean, there's no discussion. Hyperscalers have nailed it for a long time. Of course, they're hyperscalers, they have, God knows.
I, I don't know. Do you, do you guys know, actually have a sense for how many developers work at Amazon? A WSI don't know.
Like it's, I don't know. I I just, tens of thousands, hundreds, tens of thousands, right? I don't know.
Right? I mean, did somebody say hundreds of thousands? Hundreds.
Hundreds. Yeah. Even we had that many, my friends, I don't know.
Hundreds is not that many people anymore, but no, I mean, I, we hear they have tens of thousands of people. The question is who's gonna build the next hyperscaler? Who's gonna have tens and tens of thousands of people?
It's really, really hard at this point in, in history to do this all over again. And it takes 15 years. It doesn't happen in a year.
Amazon didn't happen in a year, 15 years. Right. So this is the problem, right?
So I'll stop talking about the problem. Oh, by the way, we didn't talk about this. Yes, yes.
Very important. Uh, so Sovereign Cloud. Mm-hmm.
The first deployment we had experience with was not in the us it was in Indonesia. Mm-hmm. Telco Indonesia called Insat Uri Hutchison, they go by Insat.
Uh, they're off customer public information. Um, and they wanted to deliver the first sovereign AI cloud for the Indonesian market. Uh, ah, shoot, I'm gonna say Accenture the third time.
Now they, this is where we met Accenture notes. Now, are you gonna, oh, shoot. Uh, but, uh, that's where we actually met our friends at Accenture, because they were the ones doing the implementation.
And we showed up at, with our friends at Nvidia, they made the introduction and they said, Hey, the problem we're trying to solve, looks like these guys claim they have software that could sort of give you the stack and you could deliver a service. Inad has enterprise customers in Indonesia using their platform, not a hyperscaler. We have more than one of those now.
Right. But this is a pretty significant driver. So in the Middle East, for example, in Western Europe, very, very significant driver, uh, cost, very important.
Sovereignty sometimes is more sacrosanct, uh, even to enterprises who all, all they care about is money. But sometimes, man, I just can't use a hyperscaler anymore. Mm-hmm.
There are RFPs globally that are being rewritten for 26, calling out that I need to use somebody regional. Yeah. Right.
So that is a very significant opportunity. But at the same time, the issue is are there enough vendors who can solve this problem? And sometimes the answer is no.
The question on that, yes, sir. Who do you think has the, the most data right now on the region or country with the most pressure for sore data centers? I don't know about the most, but I'll tell you the countries where this is happening in a big way, right?
So France, for example, right? Multiple, our friends at ze, uh, for example, they're doing amazing work there. If you're not familiar with this company called ze, we're very, very much worth looking at, um, in Europe, like between NEAS and N Scale and, and Carbon Three and seo, many companies are popping up.
And then the EU is funding multiple projects. Now, they call it the EU AI path, uh, uh, Gigafactory Projects, right? Which we are involved with, with our friends at Dell, uh, in Europe.
The, the activity is incredible, right? Right. In Asia, absolutely, uh, the same, but the, but the, the number of countries is, you know, there's just a few more countries in, in, in Europe, um, in the Middle East, same thing, right?
I mean, every, like, like, you know, in Qatar, right? The, the, the, the primary telco there, they're a customer, right? So each of those regional providers are seeing ai actually, for them, this is an opportunity, right?
So, you know, 10, 15 years ago, they used to run these massive MSP businesses and then came to public clouds, and it took all that business away, right? Like, oh man, I have this opportunity to bring it back because sovereignty requires my customer to work with a local provider. And in many cases, it, it requires for you to work with a local company, right?
Telcos, right? So telcos have an amazing opportunity. And for us, for sure, a big focus is telcos.
So we, we look for these telcos in these regional locations who are investing in this infrastructure. And the, the business is good. I mean, there's a lot of 'em right now who have basically the same problem.
And for us, the good news for us is it's exactly the same problem, right? So we can now rinse and repeat again and again, same problem like our friends at in side and others.