Jonah Jones – AWS Demo: GitOps Modern Operations for EKS
GitOps is a standardized workflow for how to deploy, configure, monitor, update and manage Kubernetes, its components, operations, and all the applications that run on it. GitOps lets teams continuously deploy new versions to hundreds of EKS environments simultaneously. With a clear deployment pipeline any failures are instrumented and changes can be reverted immediately.
In this session you will learn how to evolve DevOps to GitOps, and how to install Flux onto an Amazon EKS cluster to enable zero effort deployments in Kubernetes.
Transcript
Hi there and welcome to this demo session. I'm John Jones. I'm a container partner solutions architect at Amazon Web Services.
And I'm going to be showing you today in a few minutes how you can set up Amazon EKS getups with progressive delivery using aptness. So, yeah, let's jump into it. So today for our panel, we're going to be showing you how to set up get up.
So I've done a few things ahead of time just to kind of make sure that we can get everything in this demo and kind of save a little bit of time here. But normally, if you're setting up get getups, you need to make an X cluster. And so you can do that using settlor and the other tools.
And you would also need a git repository. And the first thing you do when you are making the repository for it with getups is you need to make an SS HQ here. And so when I make this, I'm using it with my email for GitHub.
And what I'm going to do is I'm going to take a look at the public key here. I know that was a private Q there I just showed you and we would take that or you could add it to our deployed keys and our GitHub repo we made here so you can substantiate in your repo and you would add one. So this would be a workshop.
The next thing we would be doing as we would need to be creating a namespace for flux. Flux is the getups engine created by We Work. It's an open source project and it utilizes a few tools to help handle getups and progressive delivery.
So we're going to create a namespace for blocks and we're going to add the flux helme repo here and then we would install flux. So here we are putting it in the flux namespace. We're making sure to call out our GitHub rebels.
So I just made this one right beforehand. It's called GetUps Workshop. That's the one where I put the assets in and we're going to get the pull interval really low.
So we're sending it to one minute just so that we can try to get through this workshop here. But normally you could set this five minutes or three minutes or whatever you felt was better for you. And so this is how I think we'll pull the GitHub repo to look for changes and declaratively apply them to your cluster.
Next, we're going to just take a look at the flux key that's created so the flux will create a secret, a secret under the hood. And then we need to take this key and add this to our GitHub repo as well. So this is going to be helping make sure that things can talk in addition to our HQ.
So here I would do the command to pull the key down. I'm not actually going to show you guys my key. Better luck next time.
But this would be the workshop secret here that we're applying. Awesome. Now we're going to apply the flux helm operating CRD.
And this is just helping make sure that everything holds well, this is part of the whole get off set up here. So we installed this also making sure we call out the SNH secret name. So this is the community secret that was just created.
And the home version we want to do is check and make sure our pods are up. Awesome. So we have the home operator and the operator.
Now we're going to make a name directory, first name spaces, and so we're going to be installing the Amazon cloud watch container insights, Prometheus influent exporter, so we can get some sort of visualization and performance metrics when we deploy this. And then we're going to deploy Ammash and then the sample app as well here. So here we're going to start with app.
So we're going to create the namespace here. So I am creating a file that is going to substantiate my namespace and then we're going to apply this just to save some time. If you're using normal good apps, you could just do a get and get push and the flux operator using its one minute pull interval would pull this apply automatically for you.
We're going to check make sure we have the atmosphere in space, we do. We're going to make a directory for our objects and we're going to pull down the Atlas controller that's made like eight of us. And we are going to apply this as well.
Awesome. So here using Ops, I created this Khubani manifest object and here is where I would add comment and push and using the magic and begin ops, we can wait a second and things under the hood are going to pull. So to pull the game, repo using those keys, we instantiate it and we should be able to see it in a second here.
So I'm going to create the injector object as well. So this is the next. So when you're namespace space has a label, same with Isdell, can automatically inject an invoice side car on applications that are adhering to that label.
Awesome, so we're going to check our next system and we can see that we had the controller and the engine running and we're also going to create the Prometheus object. You can probably see I created some of these ahead of time just to save time. But this is how you would go about this.
And so this is the object that's going to provide open telemetry data and permit this format to watch container insights. So we're going to use the same getups method of get and connect and push. And now things are sinking from our repo and to the communities cluster.
Sorry. Awesome, so we're checking everything is is good, and now since we added our cloud watch sorry, our Amazon adventure controller, we can actually go and take a look and see that we are getting some app measures created and so we can see them from the Scilly or from the Cube CTL get matched. So it's going to get mesh on the object.
Awesome. Now we're going to create a namespace for our cloud watch object. So these are the cloud watch at Sportage.
They're going to send that telemetry data back up to the cloud so we can get a nice visualization later. We're going to do the gig and commit question as well. We're going to check to make sure the Amazon cloud launch is there and you're all set, and we are going to pull from Amazon the cloud watch container insights and agent.
And I'm just saving a little time here. I'm doing a search and replace command with Kluster name with the name of our EKS Cluster and region name with us W two. And then we're also going to be pulling an agent that is going to be handling the Prometheus metrics here.
And so here, when you pull these objects, normally they would actually each of them has a name space on the top of it, and we have two files and both of them are going to try to create the same exact name space. So I'm just pointing out here that if you were following along with the demo, you would need to go into the actual file itself and you need to edit out this namespace like this. So I did a live Kearl command here just to show you how it would be done.
So we would go and we would pull out this entire space like this and say. Awesome. We would do our get add comment and push and the sense of get ups and so get UPS has a few tenants here while we have a little bit of time.
And so, like the first one is declarative configuration that is at the heart of an office. So these are the five tenants that we've has prescribed when it comes to jobs. So it's all resources are managed through that process that's expressed declaratively on the second one is version control, obviously, where you can get hub get lab works.
So any sort of some version of source control is your immutable storage for your back in configuration. And these declarative descriptions are what support, immutability, versioning and version history. Right.
And this goes a lot towards auditing as well. Automated delivery. So delivery of the declarative descriptions from the repository to your runtime environment is fully automated.
So that's all done through flux, right? We have software agents. So are we can we can Skylar's deploy and maintain resources described in our declarative figure configuration?
It's a mouthful, but that's basically breaking down that our agents are the things that are reconciling that one minute goal interval and redeploying and then the last ones that it has a closed loop. So its actions are performed on divergence between the version controlled declarative configuration and the actual state of the target system. Again, a mouthful, but that's so that if I go and I start manually editing these cute manifests and I go and I change the node porked or I change a cluster IP to a load balancer type getups is going to basically use its pull interval and it's going to change it back to whatever it is and get.
That's a source of truth here. So these are kind of the getups tennants or laws that we abide by. And so I figure I just go over these.
Why we have a little time here, get back to them. I consider getting something cool here. So we're going to check our pods from the Amazon cloud watch.
We can see that they're all created. So we have the cloud watch agent, the Prometheus agent in the cloud launcher. Now we're going to create our Apennine space.
And like we were talking about earlier, the injector label is on. So this is going to make sure that Sidecars is attached to the deployment's. Are going to.
Get and commit and push again, and then we're going to make it. And we are going to save some time. So we're going to be using a premade app called Pod Info here.
And don't mind the failed commands. Again, these are preinstalled ahead of time, so it's just it's already had an already has something there. Don't worry about that.
We are going to get and comment and push. So now these are pushed from our local up to get hub. And now we should be able to get pods and we should be able to see our pod service and namespace for the names of these apps and see that they're up and running right.
And real time isn't actually take about a minute, a minute and a half. Again, I'm trying to save time to make sure we can get to the cool stuff in this. So we have pod back in front end, etc..
We're going to check our aptness components so that she a virtual services and virtual nodes as well. So you can see we have the virtual node is going to correspond to the pod back and info and the virtual service is going to correspond to the service. We're going to check to make sure we have our virtual router up for aptness.
We can see that it's up and running, which is perfect, is created by the controller and we're going to check our roots here and just make sure our roots are also up. So those are those should be all the components that we need here. So we can see we have a root here and using one hundred percent of the weight to the pod to here.
Awesome. So we're going to export our front end name here like this, so we're just getting the name of our pod. And what we're actually going to do is we're going to pop into our our pod, our pod, and we're going to loop and hit the back and see what is actually happening behind the scenes.
And so what we would want to do is we would want progressive delivery here. So when getups is pulling a new deployment, we want it to use cannery deployment strategy and shift some of the traffic to the new deployment so we can test it out here. So what I'm going to do is I'm going to get the name of it here.
And I'm going to exact into the pod here. And from here, I'm going to curl to test to see that I can hit the back end and I'm getting a response back. So here we can actually see that it's using a random hash.
And so this one is using the V2 as well as the random patch. And so when we looked up here, we can see one hundred percent of our weight is going to be two apps. So what we're going to do is we're going to do a half second curl here.
That's going to loop and we're going to see that is continuously hitting the two. But what we'll do is we will update the. He started the mesh roots and we're going to change the loading to 50 percent, and what that will do is that will allow us to simulate what it would be like if we were testing a deployment.
So what I would do here is I'm going to cancel out of this demo. I'm actually going to run these commands here, what will happen often in a typical application deployment is when you have one of your service or two or the old one, you have one hundred percent of traffic on there. And so what we're going to do is we're going to be ending our virtual service here and we're going to send 50 percent of the traffic to our new version.
In this case, it's calling the new one he wants. It's not too confusing, but what hopefully we'll see is that our loop is going to start sending half the traffic, the new one. Or if I look down here, I can start to see that is now alternating between the two and V1.
Right. Again, I've done a few of these things ahead of time just to make sure. But what we'll do is we can push it together and push.
And it is now Guinn's. But in the case of a typical application deployment or all of times and people work at a software company, sometimes that first deployment doesn't actually work. So we're going to actually go off the script a little bit here.
And what we will do is we're going to go straight into this app's virtual service. And we're going to show you how you would reset this back and so what you could do is you could straight up change this so you could do the role forward strategy, which is a pretty common strategy, so you can straight up change it back to 100 percent and then you can do it and get that changes and get the push and get off to it. Also sync it.
Or you could go straight and do a get reset hard head like this, and so what I will do is what we're going to stand here and look at this for a few seconds and hopefully outflux will pick up our changes here and we'll start to see one hundred percent of this go back to the two so we can actually wait here a second. And while we're waiting here, I think it'd be a good time to go take a look at these metrics. So this whole time we installed the cloud watch Prometheus game and set in for container insights so we can actually get some telemetry data about our cluster and the services.
So if we go back to our firewall console here, we can take a look at our performance monitoring tab. So you get there by going to Container Insight's performance monitoring and changes is Ammash. And you can actually see the pods we have.
We're running the two back in pods in the one front and pod in your quest for second and a whole lot of other measures. So like inbound traffic, your average up time, you can get a graphical representation of how all your data is going, as well as a breakdown of the pods and the traffic by pod here. And so this is the nice advantage, it automatically is collecting that telemetry data for us.
And so if we use the Prometheus Damasak, we can get that pole and send and get some data here. And so now we go back to our deployment and we can see this colonel loop is still running here and we are now only getting the two because we had a bad deployment. We will forward, push all the data back to the two on the virtual service.
And now it's sending a the traffic backups so we can see our loop is running here, but only sending the be here. And you can tell it's running because I keep having to scroll down every time I go back up to this. So test.
So, yeah. So this was a demo about how we can use Get Off and AppNexus to do some progressive delivery and to our cluster, hopefully learned a thing or two. I will post the reffo and the top abstract and hopefully you guys can follow along and try it out yourselves.
And it's been a pleasure walking through how to use graphs with the maps and the progressive application delivery. And if you want to learn more, please check out the eight of us Booth, where we have we have members. They can talk to you in case studies and blog posts and also check out the repo in the abstract if you want to try this yourself.
Thank you so much for watching.