K8s Performance Announcement – Patrick Bergstrom, StormForge
Achieving the proper performance/cost optimization is a challenge faced by SREs running Kubernetes.CTO Patrick Bergstrom shares StormForge’s announcement about their new release of StormForge Optimize Live. The new features include bi-dimensional pod autoscaling, combining the benefits of horizontal and vertical autoscaling, to reduce container infrastructure costs without sacrificing reliability and performance.
Transcript
This is Textron TV. The great pleasure being joined by Patrick Bergstrom Patrick at CTO with stormforge had good to be talking with you today. I met great talking with you today you bet well for folks that may not know you tell us a little bit about yourself and also about stormforge.
Yeah absolutely happy to so, my name is Patrick Bergstrom. io. Fun fact, this is actually my first startup which is a little unusual.
I think for most people prior to this. com as well. So when I had my first couple conversations with storm Forge when I was getting ready to start looking for my next thing they told me what they did and I just kind of stopped and took a breath and said, okay I needed you for the last DEC.
Of my life. Where have you been? So needless to say it's a perfect fit with stormforge.
So with our product what we do is continuous optimization and automated fashion for kubernetes environments. So this is something that I've seen firsthand for years now, but whether it's with kubernetes or previously when we were doing on-prem or ec2 instances, there's a need to essentially make sure that you're scaling your application correctly in a way that's fluid in a way that meets demand spikes so that you can be there when your customers need you to be but then more importantly make sure that you're right sizing when they don't need you to be there. So one of the challenges that we have when running a large enterprise-sized kubernetes cluster is that by default we typically use like a cookie cutter recommendation for CPU and memory settings and then we deploy that across five six seven thousand main spaces and inevitably you end up being drastically over provisioned costing millions of dollars extra that you don't necessarily need to be spending.
And so with kubernetes in particular in the complexity, there it is possible to solve it by throwing human beings at the problem. But when you're operating at that scale like it the humans don't scale, right? So with stormforge what we do is we leverage machine learning to essentially right size each one of those individual workloads so that you're not wasting money and you're not sacrificing performance for the sake of savings instead you're getting both cost savings and increased capability when it comes to the performance of your application and then by extension reliability as well, so it's obviously a perfect for me and I love the organization and I love the mission what we're doing.
Well, it's perfect timing. Yeah. I swear this is totally true.
I just got off a panel with five sres talking about the horrific challenge of Performance Tuning kubernetes. So it's like right after you're out you should absolutely yeah good. Well, so exciting news tell folks about the the product and then what the EXC News is about yeah, absolutely.
So the biggest thing that we do and kind of our tagline is as we speaking of sres, we enable you to take that observability data that you're collecting and turn it into actionability. So there's a lot of capabilities out there. I think that will highlight where you're spending a lot in kubernetes or maybe where most of your compute resources are tied up, but they don't necessarily give you the specific actions to take they just say go going back to the human thing.
They say hey have some human beings look at this. Right? And so what we do is we actually collect the observability data straight from your monitoring tools and we turn that into actionability by taking action on your behalf, or if you are at a big sensitive Enterprise organization, we will we will show you the recommendations and you can manually approve those or tie them into your ci/cd pipeline.
There's a million different options and originally what our product did was you could think of it as as a really really smart vpa, right? So we essentially replaced the vertical plot Auto scalar within kubernetes and we directly made changes to the recommendations for how you deploy your pods into your environment. And as anybody who's really good with kubernetes, especially at the Enterprise scale knows there's basically two autoscalers with kubernetes, right?
You have your vertical pod autoscaler and your horizontal pod on a scalar and never the two should meet especially if you're if you're out of scaling based on default metrics like CPU and you're trying to scale both vertically and horizontally on those CPU metrics it will cause just a world of hurt and so the exciting news is that with our new products. We are actually now Rec now able to recommend by dimensional Auto scaling for your kubernetes environment. So what that means is that actually allows us to reduce the verticality of your pod down even further and enable the HPA and then we can make smart recommendations and adjustments to the configuration of the HPA itself, which allows us to actually make sure that you're you're still meeting your needs for meantime to scale when it comes to bursting out horizontally, but at the same time you're doing it on a reduced compute memory workload, which allows you to save money allows you to save carbon credits and Still guaranteeing the the excellent performance and reliability of your application without having to worry about no longer having all the extra CPU and memory waste that most organizations are wasting today.
So if I'm understanding correctly know that this is also with the products called optimized live, right? Yeah, absolutely and if I'm understand the bi-directional part of it is you kind of squash down the CPU and memory and all the resource requirements to what you really need and then you can scale it where you need to, you know, Auto scale it to absolutely over sizes appropriate for given situation in Miami track there. You're absolutely on track and that's that's the wonderful thing about our machine learning too is that you were not just allowing you to use the HPA while using our system as well.
We're actually making recommendations in an automated fashion for for the verticality and the horizontal pot out of scale itself. So we're actually making smart recommendations on what that scale Point needs to be if you're using the HPA and we take that into consideration when we're making our recommendations for the the vertical size of your pot as well No machine learning usually if it's working if it's unsupervised machine learning. It's doing it on data, right that's Collective.
Running kubernetes and the environment Etc. How much do you have to run kubernetes with stormforge to start to get those intelligent recommendations really great question and that's that's one of the things that leads into like the better your algorithm. Right?
And the more data that you have the more you can do with it the better the recommendation that you can make especially when you're talking about true machine learning that's leveraging custom algorithms like ours is versus you know, everyone else that tried to sell me machine learning when I was at uhg and it turned out to just be I like to make the joke, but it turned out to be like three linear regressions and a trench coat right like dressed up like it was machine case statement. Yeah. Exactly exactly.
Why are there just if then I'll statements in here? That's right. But anyway, no so with our system.
Like I said, we're using observability data, one of the things that you might expect is that we install our system and then we just have to sit and watch it so that we can collect data before we make that first recommendation. However, one of the things that we can do is because we partner with a lot of the monitoring capabilities out there. A we integrate directly with Prometheus.
We've actually just announced a partnership with datadog as well. I think I was last week we can also get access to historical data. So when you install our tool we can actually go back and look at the last two weeks worth of data.
And so then within a matter of hours not days and weeks we can have that first recommendation for you on what a right-sized pod looks like for your application. That's the same day. And really excited about it.
I'm really curious how much variability do you see on the scaling side with the bi-directional? Is that something that you tune, you know, like on a weekly basis or you are you tuning things more frequently than that? Yeah, a lot of that depends on the customer and what they're actually use case is we can do any for the bi-dimensional we can actually do everything from down to recommendation every single hour if you want us to all the way up to like usually the max we recommend is a recommendation once a week and then we can do anywhere in between as well.
Really. It depends on the seasonality of your workloads. Like if it's changing day after day week after week or if it's you know stable for a month and you don't really anticipate changes you can dial that back to where the recommendations are happening once a week, but ultimately depends on the customer use case.
When was cool. It's not a point in time recommendation because you make that change now, you're analyzing based on exactly figuration that set up and you know, if that's performing the way you want to or you need to Adjusted some more exactly it might personal recommendation is I usually look in I'll ask customers like how frequently are you releasing code for instance? Because as an SRE, one of the things that I always said was as soon as you make a change to an environment, whether that's your your increasing traffic, you're releasing new code, you've done a change to how you control your infrastructure something the previous analysis that you've done whether that's a performance test or monitoring or what have you.
It's no longer valid, right? You have to start over and watch it again. That's why we always say, you know do performance testing every single week or in the case with something like stormforge you'd be monitoring for those recommendations as frequently as you can because you never know how often your environment is changing to where you might need to make a tweet to get really get the best performance out of your application.
Well talk about that a little bit. I'm curious. What is an answer.
He's a life look like when they're using storage storm Forge as opposed to when they're not I'm assuming it doesn't solve every performance, you know, testing problem in the world, obviously. Yeah. It sounds like a big chunk of it is is gonna be happening on Basis, yeah, I think it is and one of the things that I've seen in my past in particular with sres if you give them call it unfeterated access to configure a pod an application or application in kubernetes.
The first thing I would always do is say okay. How high does the CPU dial go? Right because the more CPU I add the more memory I add that's more reliability for me.
And I think the thing that a lot of people don't understand especially from the operation side of the house when you don't control that p&l and you don't care what you're spending on it, right? You're gonna crank it as high as you can go but there is a point at which you you reach those diminishing returns beyond that, right? And so what stormforge does is we find that exact point for you where you get the most bang for your buck and then whether you're using our our optimized live platform and the observability side of the house or whether you're doing the experimentation side of the house with optimize Pro, which is typically the the spot where we see the sres use it the most because it does supplement like your performance testing schedule.
For instance. We actually can help you understand what That level of diminishing return looks like because if you don't necessarily care about how much you're spending you care about getting the fastest possible 95th centile response. I'm out of your application.
We'll show you that configuration in nine times out of 10. We can still save you 15 20% off of your compute, but if latency isn't really important to you and you're okay with the you know, 500 millisecond 95th percentile we can oftentimes drop your savings down to 60 70 80% below what they were before and so we found that this is really powerful not just for the sres because even if you keep costs the same we can help you make your application faster that we've also found that the CFO is love it too because we can help make the application faster and save money on your your Cloud compute Bill. And so we run into that situation where you no longer have the cfo's banging on the door of the operations house saying hey, like why are you spending so much on cloud?
This is ridiculous. And it's one of those those really neat tool sets where we just about make everybody happy that Into contact with it and you don't see that very often which is nice and how many times do the technical folks go to the CFO and say I saved you money, right? It's usually you need to save us money.
Yeah, when usually the advice I usually give them is don't tell them what you saved right like spending on something else. You'll be fine. But now you save some of that budget.
Well, I think there's a lot of legacy to throw as much CPU in memory out as possible. That's on that stands. You're right after absolutely a server, right?
Yeah. Exactly. Very good anything else to highlight in the release.
I mean, obviously that's a that's a big big announcement and super exciting. Yeah. It's it's a huge announcement and we're absolutely thrilled for it because it is one of those things where like I said before this it just this doesn't exist anymore.
Right? io and grab a free trial and try it out for Yourself depending on when this launches it will have been out already. I believe I don't know when we're publishing this on recording this right before you're launching so perfect.
Yeah. Yeah, absolutely. So definitely check out the free trial and yeah, it's we're like I said, we're just we're over the moon with it.
We're really excited. Well, I have to say it. Yeah, I know it's not just a bunch of really smart AI folks right?
You have a lot of kubernetes experts. Yeah that's on your staff to be able to do this. I mean absolutely and that's one of the things that we're really proud of as well is just the even if you just look at the raw numbers from the the sheer number of certified kubernetes up experts that we have in the organization.
It's it's very exciting and a couple of them or even part of the core contribution group to kubernetes, which is really neat. io and chat with one of them as well. A lot of them are in our our technical organization.
And as you can imagine they just love talking kubernetes all day every day. Yeah, probably can't get them to stop but that's good. That's what they do.
io and and use that account and go play with it. Go use it see what you can do. I mean to kind of savage you can get sounds pretty very exciting.
All right. Thanks Patrick. We'll see you again soon.
Yeah. Thanks for having me next. You bet.