From Infrastructure Chaos to Cloud-Like Control with Rafay
Rafay, founded seven years ago, initially focused on Kubernetes but has evolved to address the broader challenge of simplifying compute consumption across various environments. Their solution aims to provide self-service compute to companies across verticals.
Rafay typically engages with companies that already have existing infrastructure, automation, and deployments. The core problem they solve is standardization across diverse environments and users. They help companies build a platform engineering function that enables efficient management of environments, upgrades, and policies. The Rafay platform abstracts the underlying infrastructure, providing an interface for users to request and consume compute resources without needing to understand the complexities of the underlying systems.
Rafay’s platform allows organizations to deliver self-service compute across diverse environments and teams, managing identity, policies, and automation. The goal is to reduce the time developers waste on infrastructure tasks, which, according to Rafay, can be as high as 20% in large enterprises. They offer a comprehensive solution that encompasses inventory management, governance, and control, all while generating the underlying infrastructure as code for versioning and auditability. In summary, Rafay enables companies to move away from custom, in-house solutions to a standardized, automated, and cloud-like compute consumption model.
Presented by Haseeb Budhani, CEO, Rafay Systems. Recorded live on September 10, 2025, at AI Infrastructure Field Day 3 in Santa Clara, California. Watch the entire presentation https://techfieldday.com/appearance/rafay-presents-at-ai-infrastructure-field-day-3/ or visit https://rafay.co/ or https://techfieldday.com/event/aiifd3/ for more information.
Transcript
So now I'll talk about what the hell does Rafa do? I'm, I'm really happy that you guys allowed me to talk about the, the problem at this, uh, at this length in depth. So thank you very much.
Again, my name is Hasi Bani. I'm the CEO of Rafa. Uh, I, I saw Alistair's look in the back saying, dude, you didn't do that.
So that's me. And we'll talk about Rafa Now again, without referring, oh, My name is as Bani. I'm the CEO of Rafa.
Now we're gonna talk about Rafa. Did I get that right? Lister?
Now do it again without asking. So, we've been around a while. Uh, we've been around for seven years now.
Um, if you go back and you, you will not do this, but were you to go back and read my blogs about the things we've been working on. The word changed because markets are different. And, you know, two, three years ago, we only talked about Kubernetes because Kubernetes seemed to be the only thing that people cared about.
Mm-hmm. Um, uh, but the one thing we focused on since T zero in this company, in fact, before the company started, after I'd left my previous job, I wrote a blog called Introducing the Programmable Edge. But there's a phrase, programmable edge, uh, that, uh, is used in the market.
Uh, I own their trademark for it. Mm-hmm. Um, so at some point it'd be really fun to go, just, just, just for fun, just sue a bunch of companies who use the phrase programmable edge in their, their content just, just for fun.
We're not, not engine money. Uh, and even in that blog, which is a long, long time ago now, uh, the point was how do you make it easy for developers to deploy a bunch of apps in a bunch of locations? So that was the point of programmable edges.
Now, I used to work at Akamai, so Edge was big for me because everything is about Edge at mm-hmm. Companies like Akamai at CDMs. Uh, but the problem was not just about Edge.
The problem was anywhere. How do you make it easy for people to consume, compute? So we sell our product, uh, for the self-service consumption of compute to companies in those verticals.
Um, uh, and we have been doing it successful for some time. And if you Google, you'll find a bunch of companies like Guardant Health and MoneyGram and others who use our platform today. Mm-hmm.
When people engage with us, broadly stepping back from the GPU piece for a minute, right? Like, what is the, what do we do for them? Right?
So, usually we meet companies who already have infrastructure. Of course, we usually don't talk to companies who have nothing. They have infrastructure.
They have some automation, and they've deployed something. Kubernetes, we have something. And, uh, when they have one team, couple of couple of clusters, maybe.
Everything is fine. Basic Terraform works fine. And the man, the number of times I've been in conversation, when people look at a product and they go, but I can do all this with Terraform, right?
And I consistently ask them the question, tell me how long you've been working on this in this company. And they will always tell me less than one year. Mm-hmm.
I'll call you in six months. No Calling will be different. That conversation will be different.
And usually those people have already left the company company. Mm-hmm. And it's somebody else's problem now, right?
But we, we engage with companies when they have built something, and the problem we help them solve is standardization. You have all kinds of environments, all kinds of users. Um, how do you know the right version of Promet is running?
What is a dumb question, right? Mm-hmm. It's, it's not that easy actually.
Yeah. Okay. There's a CVE for whatever eng, uh, engine X.
How are you gonna go across your Kubernetes clusters, across multiple environments, across multiple accounts? Are you gonna update it? Mm-hmm.
If you have the right standardization and if you have the right thinking around automation. These are non, these are non-issue. Our customers can do this in minutes, not three months.
Uh, and then we end up essentially helping our customers using software, build a platform engineering function. This is what we've been doing. Before AI came about, we had never thought about ai.
Mm-hmm. All we did was we said to our customers, Hey, the platform that you will end up building using all this cross plane and Terraform and all these nice things, is going to look like Rafa. What's Rafa?
Rafa is the company I work at. No, But what's the, why is it that, why is it not? Is it Radian a, does it stand for something?
Is it someone's name? What is, I gotta tell you another story about the example. Okay.
Now, So one of the A's is for, That's why we have two of them. Two of them, them, no. So Rafa actually, it, the word means, uh, uh, he who elevates, it's, uh, it's a, it's one of the names of God in Arabic, and it's my son's name, hence the name of the company.
Got it. Okay. But the abridged version of the story is that in the previous company I worked at, when we started it, uh, we named it something that was, I thought was awesome.
Everybody hated it. Uh, and then we had to change it for a very specific reason. io, SOHA short, very short, very simple.
And our head of marketing thought, that's pretty cool. We, we will use it. So was my daughter's name.
Uh, so that company was by the Grace. I got quite successful. And, and then we worked on the next one, and then there was an expectation.
There you go. It worked. Well Done, dad.
I just need to be home. Right? This is really interesting.
What, the way I'm interpreting this, uh, graphic is that Rafa is, whatever it is, platform something as a service, whatever it is. I am, I'm, I'm, I have to say, parenthetically, I have IDPs in mind here. So I, I really wanna see like where you're going with this, but you are presenting it as a journey, and these are all outcomes.
I'm not seeing what it is yet. I assume you're getting to that. But this is a journey where whereby, uh, you, you, there's a, there's a, there's a, a staging to it.
Stage one, standardizing environments, creating the blueprints, applying workflows to them. I, that seems like the first instantiation of self-service, although for Ops people, but Yes. Right, exactly.
The ticketed self-service that we, that was so famous for so long, yes, they'll have to put in a ticket, but now the ops person is more efficient. Um, is this what your customers experience? Yes.
Or is this imposed by Rafa? Or is it both? This Is our product, our product.
On day one, you can go all the way to the core, to the, to the right. Nobody does. At, at least the traditional customer, they don't do that.
So usually they take steps because they already have something, right? So the first thing we come and help them do is, uh, like, well, let's make it very specific, right? So let's say you run an environment, which AWS plus on-prem, you run, I don't know, Kubernetes on in the cloud and VMs on in your data center.
So we provide, firstly, because you said IDP and interface. Now you can use backstage for it. Or if you, you can use ours, we don't care.
We have a nice backstage plugin as well. The interface is not important. The important part is the engine that you use to create the configuration and apply it on an ongoing basis so that anybody in the company can come and say, I need an environment and all the things that need to be true for that environment to happen.
So let's think about that, right? I need to go first. Make sure that you're authenticated.
So I need to talk to your IDP. So, uh, actual, I dp, I don't mean the ip, you mentioned the identity provider, right? And make sure you're in the right group.
Okay. Now I know. Then I need to figure out, okay, based on your identity and the group you belong into, when you say, I need a cluster, actually, I can't give you a full cluster because that's not the policy.
I'm gonna give you a virtual cluster on a shared environment, because that's what the agreement is. You don't know. That's my problem.
System knows this sort. No, when I say my mean, the system. Yeah.
Now I create an environment for you. I create a service account for you, and I hand it over to you. But before I do that, you may have decided as a company that you want to have a ServiceNow ticket that goes to the boss, and the boss says, I agree.
And then do everything. Yeah. And all of this automation is done on the platform.
So Rafi platform understands the notion of Kubernetes. It understands what a VM is. So that when you say to me, I want to upgrade my thousand clusters and make sure they go from one point 33 to one point 34, our platform actually knows what that means.
Mm-hmm. It's gonna go to API deprecation checks beforehand. And if it finds that one other clusters is going to not work out, because your app's gonna break, it's gonna raise a ticket, do something about it.
It can actually go write a ticket in Jira. You go figure it out. I'm gonna keep going on the rest.
Or you can say, no, stop. Mm-hmm. And you can set a policy for that or, and or the environment that I just gave you, a virtual cluster because you're a developer and we know you waste a lot of, uh, uh, you know, storage and whatnot.
I'm gonna take it away in three days, because that's a policy. Every enterprise wants this on top of Amazon, on top of on-prem, on top. We sell that.
It's a very good product, actually. It just looks way more attractive in the AI space because AI is so expensive and the hardware is very, very expensive. But the core notion of the platform is you can deliver a true self-service consumption model for compute across environments, across teams.
Like we have a large, uh, global conserving company that is not Accenture, six time now who runs 192 views across the world on our platform. Nobody knows about each other. All of them have different policies.
Everything just works. Mm-hmm. And when the cluster in a specific country breaks, the only people who can provide support for it are the people in that region.
Because we also manage the identity for them. Not that IDP, the other IDP. Everybody wants this.
Everybody's building one of these. Every enterprise as a platform engineering team that is building this right now. And our recommendation to them is, I mean, you know, if we're all smart people, we can all build stuff just really expensive to build.
It takes forever. Don't build that, but use it. Deliver the service that your developers are looking for.
Don't waste their time. I ask the following question a lot. In a lot of, in a large companies I talk to, if you had to take, pick a number, red finger in the air, what percentage of time do your developers based on infrastructure stuff, understanding terraform, you know, subnet to subnet connectivity developers?
I don't mean ops guys or developers. You know what, most people tell me the numbers. You guys wanna take a guess.
20%. 20%. Very nice.
20%. Now, if you have a thousand people in a company, which is not a big number in most large enterprise, 200 people are doing what? Nothing.
Well, it's worse than that. Sorry. It's worse than that.
Yeah. It's probably worse than that. No, no, I don't mean that it's higher than 20%.
Yeah, I, I'm sure I'm saying that they're spending 20%, but for, uh, the, the, the, um, uh, for a good ultimate user experience, they really should spend 40% based on the platform that they have. It's just they don't want to, so they spend 20% and they have a more, a bug year, a less reliable platform. So that's why by it's worse than that is what you're saying, right?
Yeah. Yeah. At 40%, if they don't, developers don't want to deal with it, and they deal with it as much as or as little as they can.
And that comes maybe to 25. The first time asked, I asked the question. This question was directed to, um, a gentleman.
His title was head of developer, experienced at a very large F 50 bank. Uh, and, uh, he said, eh, 2020 5%. Mm-hmm.
So that bank has 6,000 developers. Yeah. So 25% of 6,000 is a big number.
And on average, those people fully loaded Right? Benefits and whatnot, they're probably making two 50 fully loaded, but two for three times, that number is a lot of money. You could pay me like 2 million bucks.
And we call it a day, man, let's go. Right? So, so this is the conversation.
We have a lot, right? Hey, you're gonna build this anyway, we all agree on the premise, we're gonna build this anyway. Mm-hmm.
Well, what I'm, I'm still trying to work out, I get that it's a product, but looking at your customer journey here, how much of this is actual professional services? We don't do NPS. So we provide it with All, it's like automated standardization of the environment.
Let's go to the product. Okay? Maybe this is a good time to actually log in and talk about it.
Uh, okay, So, so this is a, this is my environment. So there's like literally nothing on right now. So if you see a bunch of things that say red or whatever, it's because we're cheap.
Um, but the idea is that, okay, first and foremost, you're gonna help need clusters. Our platform will help you set up en environments. So the platform basically will tie into your Bsphere environment or your environment, just, you know, IM rolls whatnot, right?
We actually help you question. Now I'm using the, the UI for this because it's visual. Please do know that yes, we do have a Terraform provider for everything that you can do on our platform.
It's public. You can look at it, or there's a very nice swagger. Docs if you like, uh, APIs, you can do that as well.
Uh, it's very extensive. It takes a while to just to, to scroll through it. Uh, and there's a nice CLI as well.
So if you're using, I don't know, Jenkins, right? So you can actually tie the automation into your Jenkins pipelines if you like. Everything should be automated and everything is backed in git.
In fact, not only do we back in git in our platform, there's an actual pipeline concept. So everything that we will do for you, including starting database account setting up, like laying down the foundation, which really means the VPC with two subnets, nat gateway, everything set up, it's all backed in Git, so that you can do this again and again and again. And based on the team member who joined, they may say, I need an environment in Amazon.
But what does that mean for you? It could mean an AW WS account. For, for you it could mean something entirely different.
And the platform understands it based on, based on your identity. So all of these configurations, sir. So, so I remember last time I did this presentation, people used to have their name cards in front of them.
Uh, so I could actually, instead of saying, sir, ma'am, I would tell you, I will call you Who I am. Brian. Brian, yes.
I'm easy. I'm guy. I said, who I am.
Like, am I Supposed to? I'm who I am. I'm guy, very easy guy.
Not that guy. You can call me that guy. Guy, Okay, this Guy, he can call me that guy.
Yeah. Alright, Excellent. But, but that's the idea, right?
So, and we, and we can walk through that. I'm sure we will a little bit more as we talk more, but the experience end of there should be all of the things we've discussed, the automation. So somewhere on the left, you'll find things that the platform does that will solve the problem we're trying to solve for.
But that's what we're trying to do. We're trying to build these, abstractions is the wrong word, these automation components that come together to solve a problem. And then it becomes repeatable, reusable, so that you don't need 50 experts, but Strictly above the infrastructure automation, Strictly above the infrastructure.
You, you, you're leaving the infrastructure Aware or something. Yes, sir. So, so, so now in the platform in, we can help you bring up a bare metal server as well.
That is not a big use case for us. Well, no, I just mean that, that, that whatever management scheme you're using, I-A-C-I-A-C, ai, like there's all kinds of this newfangled stuff. Your container management platform, whatever it may be, that's below the line here.
Ra Rafa is, is, is above the line. I mean, it's almost telling a platform engineering teams to stick what they wanna do, which is create this incredibly resilient platform and pretend that the application doesn't exist. 'cause that's my take on what they want to do.
Meanwhile, the developers want to pretend the platform doesn't exist, and they just want to code a whole bunch And they meet in the middle. And Yeah. And this is the dividing line between the two.
That's a great way to think about it. Absolutely. Yes.
In fact, to that end, right? So, uh, right to add this example, I have a very cynical view of everybody involved in this whole enterprise That, that nihilism is, is what it is. My friend, right?
Yeah. Yeah. He's referred, yeah, I picked it up from Andy over here in the room.
So he's nodding. So Here's an example. So this is my favorite example.
Uh, it took us a long time to get here, by the way. You're looking at a portal, which is a RA portal. We have sort of this notion of three personas as a, as an aside, you see infrastructure, past studio and developer hub.
So infrastructure is, you know, people like us, right? So like IT guys who kind of set stuff up, right? The past studio is actually a product manager because, and I use my face loosely, they may decide what to sell to whom, right?
And the price and whatnot. So we actually have a layer for that. And then the developer is a developer, right?
So as a developer, I may have access to, uh, I get the zoom out of the way, uh, a catalog where depending on my needs, my catalog will say different things. So this is all very, like, there's a rights management concept in our platform. So I will decide what brand gets versus this guy gets what the die guy gets, right?
Uh, that joke is not gonna end today. Uh, but as a developer, if let's say in my entire purpose in life is to write Java code, and my company has a framework be right apps here, spring React, MySQL or whatever combination, you should be able to package an environment where I can write my code and I can test it. I just need a test bed.
What does it take for me to get a test bed a lot in many companies? But what I should be able to do is I should be able to tell the system my code is here And everything else is magic. Mm-hmm.
Including at the end of this, it's gonna give you an endpoint, which is gonna be a routable endpoint. It's gonna say, here's the dco test it, everything else will work. Now to get there, what did, what happened, right?
So there probably is a shared infrastructure somewhere where we created a new namespace for this guy with the right limits on the namespace we created. Probably we cluster because this guy probably has an operator deploy on top of which we set up, okay, okay, I need a database because he's gonna do some testing, load it up some, some test data, because otherwise, what is the point? Then I set up an ingress controller.
'cause I can, you know, I want to test this app, or I need a DS entry because otherwise, how will I get to it? And I an endpoint magic. So y'all must have a lot of competitors.
'cause this sounds like what a lot of people do already. So we find that our biggest competitor is people building this in-house, ma'am. And I know your name is Gina because it says on the screen, uh, but, uh, yeah.
So we find that people are building this, right? So there's a lot of these different, like lots of companies do a lot of things, right? So there's a lot of different people doing components of this, right?
So there's companies who do the front end, right? So the, the IDPs are the, the internal developer platform companies that provide TypeScript based, uh, front ends, which is pretty cool. Uh, people will have pipeline implementation, somebody will have something else, something else, something else, right?
And our perspective was, at the end of the day, people want to deliver an internal platform. Mm-hmm. And, uh, we're gonna take our time and we're gonna build one.
So it took us really long to build it. Yes, but sorry, sorry. So, So you have a, so what you're saying is Rafa is a, um, is an IDP, but it also has the IAC components on the backend.
So IDP is a, is a, I don't know, two, 3% of the problem, right? Because that's just the front end. Yeah.
The fact that the platform understands the tooling, it understands the application, it understands the identity, it understands the clustered cells, it understands the fact that there's inventory here, and I'll talk about that momentarily. It understands what an app means, right? We just walk through an app, right?
It understands what that means. Um, uh, it'll understand, for example, that, uh, I need to go set up a DNS entry and at some point it understands that I need to create a ServiceNow ticket. All of those things happen in a platform engineering organization, right?
And we are helping them build these environments. So the, the phrase that perhaps could be used for this is environments, I guess. So, so question on that.
But, but are you Sorry, sorry, go ahead. Sorry. But are you, you, so you're actually trying to, um, eliminate a lot of those, right?
So you're trying to, we can do your IDP because, and a lot of other companies do the same thing that are building the IDPs. They do understand all of those different requirements. And it seems like, and and you guys were known for Kubernetes for a while as being the kuber ma'am.
Yes, ma'am. So you understand mm-hmm. The back end of it as well.
So you can do those declarative things that obviously this is this, this, uh, platform is doing because Yeah, Absolutely. Right? Yeah, absolutely.
This output. So, so what I'm asking is what are you replacing and what do you do different than all the rest of it's pretty crowded space right now. Pick two companies that come to mind, then we'll talk about them.
Ah, I think there's a, well, yeah, the one I'm thinking of doesn't really do the IDP side, but there's Anybody Platform nine and cast ai. Yeah. Okay.
I'll talk about both of them. Anybody else? System initiative?
I have not heard of that. Morpheus. Okay.
Well, we've, okay. I thought you had Morpheus too. Okay.
We've Already had some discussion just among us as, um, v realize VRealize. Okay. Alright.
So in the Kubernetes management space, I would say that the company that matters the most, so I let, let me walk through each of them, right? So Platform nine has now stopped doing, mostly stopped doing Kubernetes now, right? So they've gone to their root, back to their roots.
OpenStack makes sense. They, they were pretty good at OpenStack. And, uh, there's a need for that in the market right now with GP Clouds and their Kubernetes was, I mean, it took a while.
It's, it's not a, it's not an easy problem to solve. Cast AI is solved on the cost management part of this, right? So they help you reduce the cost of your infrastructure.
It's a slightly different problem system initiative. I don't know. Uh, we realize, or broadly speaking, VMware, right?
So their strategy, of course, vSphere is vSphere the best of the world. Nobody can compete, right? Their tons of strategy has not been great.
I hope I don't get into trouble when I say this, but I've never met a happy tons of customer. And you're laughing because I think you're agreeing, you know, he is nodding. Um, and, and, and I, of course, we all kid.
But it's a hard problem to solve. It just, it is. I I'm telling you from experience making Kubernetes work well is a really, really hard problem to solve.
Okay? And, and actually one of the advantages I can spot right away, overview, realize, is that you actually are generating the code, which is you, you can then, like you put into a Yes, sir, we generate the code. Yeah.
Force code control. In fact, our customers can start with UI and we can actually auto generate the code and you can write it back to gate. So you can tell your, your, your, your, uh, your boss.
Yes, yes. Everything is INC. We absolutely do that as well.
I missed that. So You're generating infrastructures. So you're generating the Terraform Code.
Yes, we can. Some people start, a lot of people start with it, because they already have it. We can consume yours and we can generate it as well.
So how would this compare, uh, with the, like an OpenShift where we diagram. So when I said, tell me a name, I was thinking OpenShift, by the way, you know? So OpenShift is amazing.
It's amazing. It's a great product. Um, I, I think for all the right reasons, they've just, red Hat has done such an amazing job, right?
The problem with OpenShift is, again, I hope I don't get in trouble for this, but very expensive, very complicated, very, very expensive, very complicated. So for, for a bunch of time, and Gina made the point that we were kind of known as a Kubernetes company. So when we still are, I mean, most of, majority of our revenue is Kubernetes.
The vast majority is just Kubernetes management. And those are customers that usually were OpenShift customers at some point. And the reason why they talk to us is because on top of OpenShift, they need many other things to solve the entire problem.
Mm-hmm. But we tell them, Hey, we'll do all these other things. And oh, by the way, the Kubernetes distribution in Rafa is free.
So we don't charge for the Kubernetes distribution. OpenShift is not free. So we tell them we are a management platform.
If you want to continue to use OpenShift as the distribution, no problem. We'll work with that too. We'll work with EKS and Amazon, and yes, EKGS from tan, Zu, whatever.
Why? Why does it matter? Kubernetes is the engine.
Why doesn it matter. Let's focus on the automation on top. So we would sell them the automation, and we tell them, oh, by the way, we also provided distribution.
It actually is pretty good. And it's free. People would replace, Thanks.
So what about rancher? And Rancher is great. They have, they haven't really done great work in a few years.
They worked the best. Were mortis. So Mortis is OpenStack implementation is amazing.
Their ENT platform is very good. It's a great automation solution. Uh, and, uh, we have not yet seen them in the market.
We, that just means me one person. I can only talk to so many people in a day. Um, but yeah, of course, this is my space.
I know every vendor. I can tell you exactly in private also what I, what I think more, thank you. Uh, but no, we, we know them in.
But OpenShift, that's the, the point you made. That's a Red Hat is a, has done an amazing job in this market. So in, in, in Refa, what we say that it's mostly a declarative model underneath, right?
That's what we're, that's the code you're generating. That's the code. And that you're managing and pushing the infrastructure.
Are you already running into the shift from the declarative model into something else beyond? Let's stop writing scripts. Let's stop writing those things to database tables and driving real infrastructure in that way.
So end of the day, everything has talking about Like digital twinning. Yeah, totally. No, all, yeah.
Clear. So end of the day, everything has to be IAC, right? 'cause otherwise, it's really hard to kind of figure things out when things break.
Uh, I, I don't know. I I usually see a diff, right? Uh, We just define yeah, semantics on IAC.
Look, we Much like all of our friends in the community. We also built an agent. Very cool.
We should try it. It'll tell you how to generate configuration. It's pretty cool.
But most of our day-to-day actual hands-on customers, they're pretty good at this. Man, know the real problem we solve. The real problem is not replacing 50 people.
The real problem is most companies, there's one or two people who know everything, and the rest do not, not because they, you know, not because they cannot, it's because there's like 500 things to do in a company. But two or three people just know everything. So with a tool like this, you can make everybody else do that job because those two or three become the, the sort of the configuration masters, if you will, right?
They do all the stuff in our platform. There's always a, you know, a, a a Brian, right? Okay, here's, this is the guy, right?
He's the Rafa admin. We have large financial services companies. We like two or three guys on our platform.
Everybody else is a consumer. The ops guys, when they need to deploy a new database account, they go to a portal, our portal, they press a button, everything is magic. But because the core blueprints were built by one, those two or three guys, and those two or three guys, although the, the CCOE could be 20 people, you know, there's like one guy doing the talking.
That guy. We are making it such that, that guy's thinking and, and, and ideas get proliferated across an enterprise with guardrails. And I just brought up a blog.
This is a blog about Raphael, on, on NVIDIA's website. Uh, uh, and I was gonna talk about this, this, uh, diagram later. Anyway, but this is a great time to talk about it.
This is on NVIDIA's website. This is Rafa's sort of architecture. Architecture.
You see Kubernetes anywhere on the screen? Yes. Little thing on the bottom in the middle is, is this is NVIDIA's website.
Just to be clear, this is NVIDIA's website where I pulled this up because I couldn't find the, a copy of that in my own environment. So I just pulled it from their website. I remember that.
I, I I go to this, uh, link a lot. Um, So the blue layer at the bottom, yeah, that's EKS, OpenShift, whatever we center, et cetera. And, uh, can we help with that?
Yes. We don't charge for it. We're the little guy in this market.
We have to come up with creative ways to find a foothold. And the way we found a foothold is we said, you wanna bring your own blue layer? Do it.
No problem. We do this too. We do it really well.
But two years ago, nobody would believe us. Right? Or three years ago, whatever.
Right? Now they do. Our customers use our stuff.
They don't use these other products. But two, three years ago, we couldn't start there. So we said, okay, what is the problem with this market governance and control?
'cause without this, you cannot run a platform engineering organization. So now I'll use an example from the AI word to make a point, right? So let's say somebody says, I have a hundred GPUs, 104 GPUs.
Nobody's got a, who's using how many? Most people have no idea. They don't.
Once they give somebody access to a server, SSH access, right? Mostly they don't even know how to take it back. What a simple problem.
Okay, we should do inventory management. What a simple problem. We should have a layer where we can do this sort of triaging, Reservation, Reservation system.
Well, we have one of those, essentially, or in effect, right? Because people don't think in terms of reservation, they think in terms of, I need stuff something now. And we tell them, Hey, by the way, sorry, there's nothing available right now versus Kubernetes will tell you in 40 minutes because eventually our deployment will fail.
It's a very bad experience. You have a, you should have an in management system. So we started building a governance layer, but this layer works on top of, like, for the longest time, our primary business was people would go to like Amazon or Azure or GCP and use Kubernetes there.
And then they would have these challenges. So we were getting business from AWS for example. You go down one layer, where would you say your customers start?
Are they starting on bare metal? No. Are they starting on VMware?
Are they starting, They're starting on, are they starting on something really heterogeneous, if not out of out of control, so to speak? Like, so Hybrid is our, is our sweet spot. Mm-hmm.
So definitely like, so it'll be vSphere, OnPrem on Nutanix, a little bit of Nutanix, mostly vSphere, PHE nut. And in the cloud it's gonna be E-K-S-A-K, SGKE for Kubernetes. Yeah.
Right? Or they just want to consume it like a, like a landing zone, for example. Right?
But they want a consistent experience across multiple environments. I Feel like that's almost like Brian, like almost the deal is like, there's some of these, some of these organizations doing a lot of things all over the place, uh, for many of them for a long time. And so the cost of trying to make them make it more, uh, consistent from business unit to business unit or agency to agency or whatever, it's just really high.
It is high. But then we threw GPUs into the mix, and that effed up everything couple years ago. And if we're trying to optimize for GPUs and Don't knock my business, man, This is, this is the meat here, right?
No GP make it more fun because the, well, it's limited capacity, right? Right. So CPUs are effectively in finite, right?
You'll always get some compute in the cloud because you keep pressing EIP button and eventually run out. But, but GPUs are, are definitely very, uh, finite and non fungible in some cases. Yeah.
Because their models are different. And, And in my experience, looking at the top row to the left is easy when it's a single server, small and medium or easy to deliver. The minute you cross servers and have to make sure that the networks are working functional, all the backend, uh, and then the library's on top of that, that's where I find a step function.
And so many people fall down. Let's talk about that. So, sorry.
Completely off script, and I'm loving this. Um, so our docs are pretty good. co.
It's very extensive and extensive enough that I usually forget where things are. Uh, oh, multi times. So here's, here's a, here's an interesting diagram, right?
Because you made the point that doing this inside a server is, is not a big deal, right? Right. Okay.
Let's call it a single node cluster or a cluster mm-hmm. On the, on the right hand side. And that could be the, the green sort of R thing, of course is ra.
Then you see OpenShift to, to the point made earlier and EKS, and, you know, God knows what else could be there for me to deliver a truly shared environment experience says that no one developer can affect another developer. So no lateral escalation allowed. But I want a good experience and I wanna keep my cost low.
I don't wanna give everybody a whole cluster too expensive. Mm-hmm. Most people aren't like a pod.
Why would I give them even a one node? It doesn't make any sense. All of the controls that you see on the screen have to be true.
Otherwise, we don't believe you have a secure environment. Now we do this out of the box, it just works. Mm-hmm.
Right? But, and we use this slide a lot or this diagram a lot because we're making the point that, Hey, Mr. Customer, look whether we work together or not, eh, please make sure you have all of these things, these things in place.
Now, for you to do this Mr. Customer, you have to understand all these things. You know, network segmentation number six, right?
You know, QOVN is actually pretty cool, right? So it's like a VX plan implementation inside a cluster. It's really cool.
Gotta understand this. Now you've done BGP pairing, right? Mr.
Customer, because you gotta do this now. Mm-hmm. Right?
So it's e BGP implementation, right? So now, okay, so all these configurations happen. Now the problem is these technologies add up, right?
Like individually, I can do anything as an IT engineer, I can do anything. But the problem is I gotta do 50 things. Now.
I don't have time, and my team is limited, right? I have three, four guys who can really do this. I gotta move fast.
Mm-hmm. Right? And those are the people, right?
With the, with the clarity that my project, my, my goals are very large and I don't have enough time. They seem to like a product because they see this as, as a, as a multiplier effect of their own brain. But this is an example of what happens inside a cluster, not across, because across you are right.
There's a much bigger problem. Mm-hmm. I completely forgot where I was before this, by the way.
So, governance lawyer. Yes, sir. Thank you.
Yeah. What type of, uh, network policies do you handle inside of this? Uh, when, when you, when you're thinking network policy, are you thinking more so in terms of, uh, EastWest or lateral escalation, sir?
Or what are you thinking? Uh, I was thinking, well, east, west and outbound. East, west, north, south.
Yeah. Okay. So, uh, I'll just go back to the picture for a second.
So, so in this specific implementation, right? So the app running at the top is running at a V cluster. We like it a lot.
We cluster, we think is a pretty cool technology because the issue is if you just give somebody a namespace, even if you do all the security, if you just have, uh, uh, a role binding versus cluster binding, you cannot deploy an operator. Mm-hmm. And most of these new packages that you find now for ai, they all happen to be package as operators, which means you can't deploy it, which means you need your old cluster, which means we, clusters are pretty good solutions.
Now, v clusters running on, uh, an a quota implementation layer that we wrote. Okay? So we enforce this.
So this is not a Kubernetes construct. We, we came up with this idea of enforcing quotas across your multiple environments. Under it underneath.
We essentially are, uh, so this is a, uh, this is a calico plus ISO construct. We sort of kind of chain is ENT with Calico somehow. Okay.
Uh, to do the sort of, this becomes more, not With either one, but go ahead. So, so Calico, most people seem to use it. It's, by the way, it's, it's all IP tables oriented, right?
At some level, right? Okay. Okay.
We all, that's, That's enough to know C, right? It's a CNI, right? Mm-hmm.
Um, and but the more important thing is at the bottom, I said the network implementation happening. So we wanna make sure that a, that a, that a, a tenant inside this cluster when they, their traffic is only being seen by their traffic. So that's a V extent implementation.
That's an overlay. Okay. That was, that was actually the, the question that I was going for.
And that answered it perfectly. Thank you. Yeah.
And of course then there's an egress, which is standard egress. So for that, we, we use, uh, uh, uh, engine X, but with egress egress management. No, I just wanna make sure there was, there actually was a private network.
Yes, sir. With The cluster, sir. Yeah.
Although, yeah, I mean, just, just doing I PIP is not good enough. There has to be an overlay. Otherwise this is not secure, right?
Because the, the, the goal is both ways, right? So yes, traffic should go in and out, but, but east sort of lateral escalation protection is really key. So you should not be able to even figure out the IP addresses, because if you're running on the same network with IP tables, that's not secure enough.
Right? Okay. Right.
So, alright, going back. So I hope this gives you a sense of, look, you know, this is an introductory section. Um, if you guys would like, and we would really appreciate it if, if, if we can do this with you, right?
Individually or in a subgroups, we'd love to show you a, like a two hour demo of what's possible. 'cause in your own sort of day jobs, you may find value in it or you may be able to give us feedback. We'd really appreciate it.
If if you're open to it, it'd be very nice. That's pretty good. But we've, and we've discussed this already, right?
So our experience has been people start in the middle of the screen journeys and companies start in the middle of the screen, not at the left of the screen. Mm-hmm. People just start writing code because terraform enough, mo enough modules are available.
So I've already, uh, you know, there's a bunch of open source code available, uh, in the registry. So lemme start there. And most people don't think in terms of, I have multiple people in my organization, multiple teams and their similarities and differences, and they don't start there.
And they do that later. And usually that's generation two and a company first team is left, right? Uh, and for whatever reason, I don't know, luck of the draw maybe, right?
We meet a lot of those teams, right? Because they have a significant pressure on them to solve a problem. So then, you know, they, they, they hear about us from AWS or Gartner or whoever, and they'll say, you know, if you're working on this level of automation, RAHA can probably help you.
So does, uh, does RAHA Alpha also have the ability to ingest one of these from someplace else? We have built tools to that end. It's always a little bit of a, uh, effort.
But yes, we have built tools where you can show us your terraform and we'll try to do something with it. Or a question was asked about rancher, we take your rancher config and we'll convert it to something non rancher or more generic in our world, or OpenShift for that matter. So we do those things and we've built those tools over time.
I would not call them part of a platform, but definitely these are tools in our, in our, in our, you know, on our tool chains, okay. Uh, to solve the problem for the customer. Nice.
So you all agree with this, given this conversation, uh, by the way, the data on the right, it's actually real data. It's not made up, right? So we, you know, our yes team seems to obsess over these numbers, but they ask these questions all the time.
And I'm pretty sure our customers are pretty pretty blunt. I don't think they lie. Um, we ask them specific questions of course, but, but it's actually quite amazing, right?
Once you have the right standardization in the company, the amount of time that you get back is actually, it's incredible, right? And then people realize why we were wasting a lot of time on a lot of stupid stuff. Mm-hmm.
So this list, of course, the deck is available to you, uh, you know, at your leisure piece, look at the, this, this is not intended to be, uh, an exhaustive list of features. When this slide was put together, a few of us sat around on a whiteboard and said, what are the most important things we can put in without the font getting too small? Mm-hmm.
And one more box would make the font too small. So we decided to stop here. Um, but it's a pretty, pretty exhaustive list.
Wait, so hold on. Back up just one second. So standardization has two meaning here, but one of the things that I've been noticing, including in, in a lot of the conversation that we're, we're, we're having with each other while you're presenting, is There's a, something going on.
And I'm not on it. Of course, The, it's to, to put it bluntly, there's always, there's kind of this tension between DIY and, and, and off the shelf. I, I don't sense that Rafa, Rafa is like completely off the shelf, nor should it be.
But, um, a whole lot of this stuff is already being done, including by folks in this room, uh, as well as by other vendors. Um, at least parts of it, you know, mean there's a Venn diagram and they don't all completely overlap. Um, so when you talk about standardization, I, I feel like what you're really saying is why invest in someone else providing you with this environment orchestration platform and managing it for you?
Why do that? Because standardization also means, you know, using standards like Kubernetes or, Or, and your standards Yeah. Or whatever.
Well, okay, yeah, that's right. So it has those, it has three meanings. I miss, I missed that one.
That's an interesting way to think about the problem. I mean, ultimately you're coming in and saying like, these are things you all want, want, wanted to do. Or you maybe in, in the process of doing.
I don't think, uh, there are many organizations that have already done a lot of this, whether it's using vCloud Director or vRealize or any of these other tools that have been around for a while, um, or using a competitor or doing it all, almost all homegrown, like you said. That's my experience as well, by the way, is that most shops size have fully built up some kind of, uh, um, a hydra, uh, you know, money headed monster of mm-hmm. Infrastructure or orchestration with a whole lot of people who all know exactly how to do those little things that they know how to do.
And you're, you're potentially disrupting that. So, so yes. So what do we do about it?
Right? So day one, we can't come in and say everything we've done is wrong because obviously it's not wrong. It's all great, right?
You've done great work. It just, it's become big, right? So where do we insert ourselves?
Where does standardization start? There's a very simple step we may take with a customer who has a bunch of Terraform, right? So we have this concept called environment templates in our platform.
Um, I'll just pick a simplistic one. Yeah, why not? So, okay, so you have some code, uh, resources.
Each resources are essentially a Terraform blo, uh, and, uh, you know what? Let's do something very simple, simple. Let's add a ServiceNow hook.
But before you launch your Terraform code, we're gonna ask a boss, should you do this? Number one use case. Makes sense.
Simple thing, right? Number one use case. Um, or, um, where's the, uh, schedule?
Okay, alright. You have a bunch of developers wasting a bunch of resources. What if every time they bring it up, it's gonna go away in seven days?
Number two, use case. We can do many, many things. I'm just speaking second of the fun use cases, right?
The point is, we have to deliver value immediately. I can't sit here and say, for one year we're gonna take your code and translate it to our world. Nobody's gonna buy a product now we're gonna have it right now.
AI use cases are different, right? And we talk about that soon. AI is very simple because it's a new, it's all green field.
So AI is very, very easy to sell to because they have nothing. It's all new. And I'll talk about that for sure.
But this part is really important because, uh, and the reason why I'm spending time on this is because, um, without these foundational technologies that we have built in this company, we could never have solved the AI problem anyway. 'cause without this, how can you run a cloud? Because I'm telling my customers, you can run a cloud, you're gonna be a CSP on RAs back, which means you gotta have all these things in place.
Now, think about a cloud, right? So, so in a large enterprise, if I have these policies where I can do the lateral escalation protection that we discussed earlier, right? We can do, uh, upgrade of a thousand clusters in a shot or set policies to tie into your internal controls like ServiceNow, JIRA, whatever, or set policies to remove environments or move environments or, or, or, or, and we looked at GitHubs before, right?
So we can do staggered updates of things, et cetera. So all of these things are not possible on top of what you've built. And the next step we take is we say, Hey, so how do you build your clusters again?
And then we start taking that back because the next time you have an upgrade coming up, we say, Hey, if you're gonna move from here to here anywhere, because most people, when they upgrade in, in the industry, by the way, they don't actually do inplace upgrades. They do blue green deployments, they put brand new clusters and then move everything over. You're gonna do that anyway.
Why don't you, why, why don't we make the, the green cluster be ra it's free. And we take the Kubernetes and then we say, oh, by the way, KVM, we have a pretty cool KVM solution free. We only focus on the governance layer.
'cause to me, everything else is free. Should be free. I really believe this.
I think Kubernetes should be free. Nobody should charge for it should pay me for the, for the policies and no, because I gotta make money, right? But Kubernetes should be free.
Nobody should make money on it, including, uh, red Hat to the prior point. Very expensive. But these tools now come together for me to deliver an actual cloud-like experience.
And if it's okay with you guys, I'd, I'd love to sort of show you a simplicity demo of what people can do with our, on the AI side of the house, given the tools that we have. But before I do that, I just wanna make sure, is there any in uh, area here that you'd like me to dig into? Think platform engineering.
Think big picture. Platform engineering problems.