Kubernetes Deployment Validation in Two Minutes
Kubernetes deployment validation can eat up hours of an SRE’s week. Alan Shimel welcomes Sai Joshitha Kathari, Senior Site Reliability Engineer at Visa. Furthermore, she recently wrote about cutting validation time from 45 minutes to two minutes for Cloud Native Now.
A green deployment is not a healthy application
Sai has spent more than six years as an SRE working on cloud infrastructure and distributed systems. In addition, she focuses on making deployments safer and removing repetitive manual work. A deployment can report success while pods still crash or fail to come up. Consequently, engineers often log into every cluster and check every dashboard by hand.
Building Kubernetes deployment validation into CI/CD
Sai moved those manual checks directly into the Jenkins pipeline. Meanwhile, the pipeline captures the state of each cluster before the rollout begins. After the deployment, it checks pod readiness, failure rates and workload stability. As a result, the pipeline flags crash loops and failed pods right in the Jenkins logs.
Teams running 20 or 30 clusters no longer have to inspect each one. Therefore, engineers can go straight to the cluster that failed with the logs already in hand. Furthermore, the time saved can go toward troubleshooting and new automation. Sai stresses that a deployment is not finished until validation also passes. As a result, every SRE team can benefit from automating the checks it already runs by hand.
Sharing knowledge both ways
Alan praises Sai for sharing her experience with the wider tech community. In addition, Sai hopes other SREs and DevOps teams will adopt the approach and suggest improvements. Consequently, she sees reliability work as a two way exchange of ideas. Alan calls it one of the easiest interviews he has ever done.
Explore more DevOps coverage and the latest Techstrong TV interviews on Kubernetes deployment validation and cloud native operations.
For more information please visit cloudnativenow.com
Transcript
Hey everyone. Welcome back here to Techstrong TV. I want to introduce you to Sai Joshitha Kothari, who's a senior site reliability engineer, and she's going to talk to us, and she wrote about this on Cloud Native, but she's going to talk to us about from 45 minutes to two minutes, what she actually validates before calling a Kubernetes deployment healthy.
Sai, welcome to Techstrong TV, and thanks for joining us. Thank you, Alan, for the introduction and my work. So first of all, I'm Joshitha.
I'm working as a senior site ability engineer, so I have over six-plus years of experience as an SRE. So my work mostly focus on cloud infrastructure, distributed systems, and platform reliability side. So a big part of my work has been around making deployment safer and reducing the repetitive operational work through automation.
So every engineer will understand when I say the repetitive operational work because our day-to-day life includes almost so many manual things. So we wanted to automate those things to reduce the manual work and the impact of the applications. So that's where I focused mostly on.
So one area I have focused on, Kubernetes technology. Everyone is, I think, working on this particular technology right now. So as an SRE, performing deployment is my primary thing or primary responsibility I can say.
So there's the reason my focus is on discussing about these Kubernetes deployment validation. So this work led to my article about reducing the validation time. So because after every deployment, what we do as an SRE is we go and check every each dashboard, or we go and log into each and every cluster.
What if there's a big organization, and what if they're maintaining almost 20 to 30 clusters? They can't go each and every cluster. It might take so much time.
So that's where my article talks about. So reducing the validation work from 45 minutes to a few minutes, that reduces the manual intervention. So today, I'm excited to share especially about this, where I learned from my experience, and more importantly, I'm very much excited how SREs and DevOps teams can apply the same approach and suggest more and more approaches from the things they have developed.
So, if I talk about my article, as I mentioned, it's a problem we often see with Kubernetes deployments. It always says a clean deployment, but that doesn't mean that application is completely deployed, and the application is completely healthy. So, I wrote about how I approach that gap by automating these particular validation checks every engineers were previously doing manually.
Instead of checking whether the deployment is completed or the processing checks, things like pod readiness, failure rates, and whether the workload remains stable or not. So these keywords that I'm using as every Kubernetes who are working, everyone will understand. So this is what I'm talking about because after every deployment, every engineer will go to that particular cluster and checks the port state and how this particular application is running after every deployment.
So instead of going to each and every cluster, each and every port, checking the logs, so here I have come up with the automation where this particular port process can be included during the deployment itself. For example, the Jenkins pipeline. That is one of the technology that everyone uses during the deployments like CI/CD.
So where instead of checking the validation manually, we can also include these particular manual checks during the CI/CD itself. Since the starting, when the deployment starts, it automatically, CI/CD checks go to the cluster and just checks what is the status of the cluster before the deployment. It gives the particular report.
This is the current state of these particular clusters. Then when deployment starts, it will automatically deploy that particular artifact, a particular version. So during the deployment, if it is deployment is successful, the back end, we never know what is happening.
But deployment is showing successful, and we will think the deployment is successful. After the deployment is successful, instead of going to the each cluster, the deployment pipeline itself will say, even though deployment is successful, these particular pods are going into this particular crash loop. But this particular pod went down.
This particular deployment did not came up. This list will be displayed in the CI/CD Jenkins logs itself instead of going further and logging into the cluster. And here the advantage is instead of going to each and every cluster, the deployment CI/CD will show which particular cluster is failing so that the engineer can go to that particular cluster instead of all other 20 clusters, which can save us most of our time during the validation because as an SRE, deployment itself is not the end thing.
Because even after deployment successful, when the validation's also successful, then only we say all the applications are successfully run. So deployment validation plays a very important role in our daily responsibilities. Every SRE engineers, I think, will do these validations after every deployment of their application.
So that's the idea came up. So instead of the manual checks, we can have all the things automated in the deployment itself instead of the manually logging into that. So that's my article tells about where we can reduce the manual work and where we can include these automations when we are doing deployment so that every engineer can save time and focus most of their time on some other troubleshooting things.
So my article tells it's easy way to log in. It's mostly it will say where it is failing instead of saying only deployment is successful. It also tells us where it is failing and where to go and where to can check, and it also gives the logs directly so that engineers no need to worry about to logging into the all other applications.
They can directly go to that particular application and troubleshoot that so that they can save most of their time, and they can think more efficiently how it can be reduced later in. So that, yeah. So I feel like the time we have saved can be contributed somewhere else, which is efficiency of that applications and some other automation things.
So that's why this article came up. I got to tell you, this is the easiest interview I've ever done. I didn't have to say much.
But Joshitha, thank you for doing this, first of all. Yeah. You know what?
I think part of what makes the tech community great are people like you sharing your knowledge, your expertise, your experience for other people and maybe let them build on that, right? Because, one, you pay it forward. One generation to the next.
And so this is a great thing that you've done here. We appreciate it. We appreciate the article.
Keep it up, and we'd love to hear more about what you'll be doing in the future. Oh, thank you so much, Alan. And thank you for this opportunity.
I really wanted to talk what my work is so that other engineers can incorporate it, and so that I can also get many suggestions from engineers so that I can incorporate from them. So this is, as an SRE, it's a two-way give and take idea so that we can implement in our systems, and they can implement in their work. So, this is the most motive of my contributions where I can contribute my work so that every engineer can also share their knowledge and share both ways.
So yeah, thank you again for this wonderful- Thank you ... opportunity. Yeah.
Thank you. Thank you. Sai Joshitha Kothari here on Techstrong TV.
We're going to take a break. We'll be back.