FinOps and Machine Learning – JT Giri, nOps
NOps CEO JT Giri explains how financial operations (FinOps) is being automatically implemented using machine learning algorithms to rein in cloud computing costs that have spun out of control.
Transcript
This is Textron TV. Hey guys. Thanks for the throw.
We're here with JT Geary. Here's the CEO for nops. And we're talking about the cloud and how to automate the management of cloud infrastructure at a time.
When costs are much more on people's minds these days than average JT. Welcome the show. Thanks, Michael.
Thanks for having me. I might argue that we've allowed developers to kind of spend like drunken Sailors because no one ever really thought twice about looking hard at what they were provisioning and using in the cloud because it was all in the name of agility. But now that the economic headwinds are blown a little harder do you think folks are looking more at what the costs of cloud resources are and how developers are using that stuff?
Absolutely. So, you know just a little bit background on myself. You know, I I've been migrating companies to AWS ever since ec2 was in beta.
When I was working in a data center, you know, it used to take like months before we get a server. And we had to go through an approval process. I mean no one wants those days, right?
We we don't want to wait. several months to get resources but you know, the flip side of it is provisioning resources is very easy, whenever people provision resources, maybe they over provision and with the mindset of you know, I'll go back and right-size them, but they but you know, no one like has time to go back and you know pick the right resources, so we are in this sort of You know weird state where provisioning resources is a lot easy and very very easy to provision new resources. But optimizing those resources is not that easy.
So this is where automation comes in play. We have to provide a similar experience to developers. So it's as effortless to optimize your resources as it is to provision news resources.
So this is one of the areas we focus on we we focus more on automation. So automatically just like the experience developers are used to provisioning resources. We leveraged automation to you know, optimize the resources to optimize the cost and you're right Michael like right now everyone is trying to optimize, you know their clouds and in the last two years maybe that stuff, you know didn't matter that much but now it does so, you know, we're like seeing a lot of interest from companies.
To optimize their classmen. It doesn't seem like the cloud providers go out of their way to really make that easy. But do you think people are shopping clouds now because they have multiple clouds and there are they looking at the cost as they shop clouds for workloads.
Oh what I see based on our customer base, like normally people pick a cloud and then you know, they they leverage that primarily, you know, that's that's what I see. It's incredibly difficult to, you know, just run workloads across multiple clouds. So people definitely leverage multi-cloud, but I I noticed they will like primarily run most their like, you know workloads in one cloud or the other but sort of the mindset of optimizing wherever they're running their cloud is, you know, a lot more appealing to everyone right now and everyone's thinking of ways where you know, they can figure out a way to optimize their spend and that's why a lot of people are adopting like this finops sort of framework and this to us is like one of the best disciplines I've seen in the market.
Which drives sort of accountability theater Point earlier, you know, like people were provisioning resources, but no one took the ownership. So I'm seeing a lot more people adopting pin Ops and automation. So everyone takes kind of responsibility of their cloudsman.
What is different about phenobs than you know, we've had capacity management issues since the Mainframe way back when so what's ultimately different about the way we're approaching pin Ops today versus what we might have done in a premise environment to control costs. Yeah, so Cloud environment inherently is very complex. So, you know you like we mentioned earlier right?
You can provision resources within minutes and seconds. And sometimes you provision for testing for R&D sometimes production. Some of the costs is shared AWS has like, you know close to 200 Services.
These are built differently. And you know people who provision resources never actually see the bill and CFOs cannot make sense out of the actual bill because you know, it's it's to a second resolution and these are like millions and billions of you know the lines on your bill on your, you know building CSP. So it's just hard to Hard to make sense and and it does not require any approval process.
Right? Like anyone could just like provision resources very easily. So, you know, like one of the promises of penops is like starting with just building accountability like accounting Ever every single Cloud spend two different buckets to different teams, and then we're in Ops helps is like just more focus on automation like we we do automatic show backs and you know, we also automatically manage commitments.
We you know pause resources when they should not be running and and we're having a lot of success so to your point earlier Michael, I just feel like it's just inherently a lot more complex. They're a lot moving pieces pieces. There's multiple people provisioning resources and the billing is liking, you know, incredibly complicated.
How optimize can we get because cloud service providers will have everything from spot instances to reserve instances and there's a lot of different pricing schemes. So can I take advantage of that down to the minute or do I got to make some commitments but I'm still automating that and trying to optimize it I guess I want to know like how granular can we get? Yeah, so let's take the example of a reserve instances.
We strongly believe that Engineers should focus on Innovation rather than trying to understand either vs pricing plan. So one of the things nofs does is we look at your usage pattern and we automatically try to maximize your RI coverage your team doesn't have to do anything. So what we see is like companies have dedicated resources where they spend like, you know hours and hours trying to figure out you know, which RI instances they should buy and I honestly believe it's just hard to do it manually.
It's very complex, you know billing is always changing so I think by leveraging automation, you can actually maximize we have some customers where you know, we actually have a graph where you have on-demand resources and we reserve and the graph is like kind of matching and it's just hard to do that if you're if you're doing a manually And and workload change all the time Michael like you launch resources you shut down resources. It definitely is not a one-time activity. And this is where I feel like, you know, people leverage tools like like analogs to your point about spot instances again, you have to you know identify which resources Can run on spot and then you know, like there's some workloads maybe not a good fit for spot.
We also provide automation for that. So people could easily run any type of workload on on spot instances. So, you know, the the theme Here is like we want to provide all the Automation and we want to free up the engineers so they can focus on you know Innovation rather than trying to understand Ada vs pricing plan trying to you know, profile the application and trying to implement, you know, all this stuff manually.
We have seen this emergence of a cottage industry around fin Ops where there's all these different platforms. But do I really need a separate platform or should it just be all part of my automation experience in the first place because you know cost is just yet another factor in where I decided to put a workload, right? Yes, what we are seeing is that there's like a lot of on point solution in the market where you know, some of them focus on spot some of them maybe focus on managing your commitments.
There's some tools that do like showbacks and things like that and UPS to to my knowledge is the the first sort of complete Finance platform where we provide sort of all, you know, fin apps has three phases, you know, inform optimize and upgrade and we cover all these three areas. We provide show back functionality. Now, we are automatically manage commitments.
We automatically can pause resources when they're not being utilized. We have a really good scheduler. We leverage machine learning and and then we have a very nice integration with jira and other ticketing solution where customers could just like, you know track all the fin Ops related initiatives and the stakeholders could you know kind of track the progress and and we do see for a lot of the customers very appealing where you have like a complete automated Finance platform rather than looking for a bunch of OnPoint Solutions.
Mmm to your point there's a natural tendency to over provision infrastructure. Can I use your platform to kind of create a set of parameters that I don't want somebody to exceed and then what an alert come to me that says, you know, we're close to exceeding this so I can actually do something about it before something bad really happens. Yeah, we actually go one step further like I feel like there are a lot of tools out there where where they show like here are the recommendations but again Michael and we started the conversation provisioning resources, very very easy.
We have to provide similar experience to developers. So optimizing is as easy. So we we go one step further.
We focus more on automation. We have a very like nice integration with interviews event bridge for example and leveraging event Bridge customers could create policies and then based on that. We are actually, you know, taking actions and customers account.
We believe if you just show recommendations developers and Engineers consider that as work. We have to make it as effortless as possible. We also Look at the utilization of for example Auto scaling groups based on you know our ml we're able to detect like what should be the right configuration for this autoscaling group.
And we we tweak those auto-scaling groups. And in many cases were able to optimize customer spend by 40 to 50 percent because workloads are always changing right and Michael, you know, like workloads change. So Engineers, you know need to go back and maybe adopt the work the instances and resources based on that but they never have time right?
This is where our ml comes in handy where we're constantly profiling and based on the utilization. We're able to like, you know, make the changes and customers environment and kind of say optimize their spend on auto pilot. Driving this conversation and how actively involved is the finance team or are they just coming to the IT people and saying, you know cut budget by 20% and you go figure it out or is the finance person starting to use these tools to monitor the situation directly?
Yes, that's a great question. I just feel like in last two years Finance people didn't have much saying and things have changed right? So now Finance.
People does have a lot of saying that they're they're heavily motivated to you know, reduce and optimize the cloud spend. So I'm seeing a lot more conversations originating from like CFOs and VP of Finance type of folks. and I'm actually seeing a lot more aggressiveness like like hey, let's slash our bill by 50% you know type of initiatives and and then yeah, then like it's like VP of engineering who's actually like kind of leading these initiatives and normally they have a choice right like they can either Spend engineering cycle trying to figure out you know, how does AWS pricing Works, which should be what should be Auto scaling configuration.
What are these resources that we can run on spot? Or they can leverage platform like an app. So one of the things we do with anops is the platform is completely free.
We only charge customers if we we take a percentage of the savings essentially. So 100% focus on figuring out ways. How can we automatically optimize customer spend and if they optimize spend we make money so and we're like, you know, that's our number one focus on how do we make this stuff easy for for our customers?
So if you get a percentage of the savings you're kind of taking on some of the risks. So does that mean that if something goes wrong you'll fix that too for free or how does that work? Yeah, especially around managing commitments Reserve instances.
If we end up, you know buying more than more than what customer needed. We do refund our ml was pretty good. We're we get better at it every day, but sometimes workloads change right people or maybe doing some experimentation and we ended up, you know, our REI and ml platform maybe level provision more than it should have but yeah in those cases this completely, you know, we refund that on our scheduler.
Are like our recommendations are really good and if a customer does not agree. With the recommendation they can just go in and start those resources back up and we don't charge anything. So we give that full control.
And we do the same thing with kubernetes, you know, if customers accept our pull requests we charge percentage of that. So it's completely risk free. Is it getting easier to do this in the age of kubernetes?
Because one of its attributes is it's supposed to be able to let you scale up and down more automatically and that's one of the benefits of being Cloud native so is does this get easier or harder as we start to move towards this new era of microservice is based applications. It's hard kubernetes is incredibly hard and Also depends on the size of the company, right? If you're a big company you have resources kubernetes adds this layer of abstraction.
And so in order for us to you know, right size or autoscale kubernetes. We actually have to deploy an agent to get that understanding. So it is incredibly difficult and that ecosystem.
I mean every all of our customers are like adopting kubernetes every day. And and it yeah definitely does not get easier. So we you know, a lot of people leverage an Ops to to make sure they're been packing the right way.
Make sure that utilization is right and you know recommendations are getting better. And we also have an integration with GitHub where we send a pull request on our home chart and we make updates based on based on how we think you should you know configure these pods and then it's really up to the customer. They want to accept our pull requests or not.
But yeah trying to do this at yourself is not an easy task. Yeah, so what's that one thing you see customers is still doing that makes you shake your head and go. I can't believe that people are still doing this this way and and have we got stuck in a particular rut that you can look at and say folks if we just did this one thing differently the world would be a better place.
Yeah, I think you know I come from my engineering background early in my career. I always I was like I can do this myself, you know type of headed to you. I still see a lot of that where people are like, oh I'll just do this, you know, I'll spend like a whole week or month understanding.
It'll be as pricing plan and based on that how you know try to reserve instances or let me just write bunch of scripts and we'll stop these resources on the weekend and things like that. I just think all that time should could be better spent on innovation. In helping your customers and really leverage platform like nons because this is all we do every day, right?
We wake up every day trying to figure out how can we optimize customers spend? And and we focus a lot more on automation. So I think this is just this is something I see is changing.
A little bit now because there's this sense of urgency because the what's happening in the economy. But yeah last couple years like I just I saw a lot of that, you know people just trying to do this themselves. If you're like a Netflix, you know, you can have a team who's gonna focus on this but you know for a lot of other companies, I think you know, they should leverage my platform like an Ops.
All right. Well given the cost of Engineers do it yourself is always going to be an expensive option. Hey JT.
Thanks for being on the show. Thanks, man. Appreciate it.
All right back to you guys in the studio.