Accelerating DORA Maturity: How Observability Drives the DevOps Experience at SKILup Days 2024
Gain Practical Insights: Learn how to leverage observability to enhance key DORA metrics, leading to more efficient software delivery and improved operational performance.
Get a Clear Implementation Guide: Discover the Observability Maturity Model and how it provides a structured roadmap for adopting and scaling observability practices in your DevOps workflows.
Drive Continuous Improvement: Understand how real-time insights from observability can be used for continuous optimization, allowing you to refine processes and achieve better results over time.
Transcript
A, a recent report published by, uh, DevOps, uh, the, the Google's state of DevOps report for this year. Uh, so they found, I think they found that, uh, there's a 50% ROI increase for the organizations, uh, who are following mature dev practices and followed by some very good eye-opening numbers. Uh, these MA organizations or the light performance, uh, they are having around 1 27 next faster, uh, lead time to deployments.
And they are doing like a massive amount of fast, uh, deployments during a period of time. So what this all sources is speed is everything the business need is speed. And then technology has to provide or enable this speed.
And in order to do that, you have to align with some of these Dora best practices that can have a direct correlator to your revenue. So with that, uh, welcome to observability drives. Modern best dev was practices skill update.
My name is ind and I'm going to walk you through about how we can accelerate DevOps maturity by leveraging observability. So I will discuss about typical observability based practices and how we can leverage those best s to uplift or amplify DORA maturity. 0, which is a new concept and how that we can leverage as well.
So in high level, the agenda. So I will go through the DORA metrics, why it's matter. I think I already shown couple of, uh, figures, which are, uh, generally eye openers, which says lot.
And why do the DORA metrics or the Dora mat is very important. And we'll go a little deep into DORA metrics and then we'll discuss about Dora and the relationship with observability. 0, which is a new trend imaging, which is about shifting left because observability people associated more with production and what is happening, uh, at the deep end or age location.
But we also want to know, can we give this benefits to the developers who actually develop code so that they are better prepared in doing their work? And we'll also follow up looking at how observable generally improved ORA metrics. We look at the implementation roadmap and we'll finish up with some of the best practices, what I have learned from my experience.
So moving on a little bit of myself, I'm based out of Columbus, Sri Lanka. I'm a solution architect, is specialized in site reliability, engineering, DevOps, observability, AIOps, and I do some little work on generative AI as well. So I'm employed at Virtu.
I am kind of like involved in technical delivery and capability development at my organization. I'm also very passionate technical trainer. I do a lot of trainings, uh, related to SRE, DevOps, observability, AIOps.
And I continuously been publishing, uh, the, the technical blog post I'm writing. I can, you can find me at dev two. And I'm a very proud ambassador at DevOps Institute and also AWS community builder.
So with that, uh, let's dive into our topic today. So why Dora Metrics? So why it's important and why it has been for decade.
Uh, people are still getting into the this metrics. So why we track Dora metrics, because if you can see it provide a lot of, uh, data driven insight. It provide organization with data.
So we are, the organizations are making decisions based on data, uh, and it's also allowing us to benchmark ourself and set some goals like based on the, uh, the tech stack I'm coming in or the business domain, we'll be able to benchmark ourself and improve on that. And it's obviously the continuous improvement. When we have metrics, it's easy.
And then we can start coming, coming up with the short term, medium term and long term objectives and drive. And it's about prioritizing where the investments and other things has to go in. And we can look at the business ROI and it's pretty much, uh, need and year on year.
The organizations are adapting, uh, the DORA maturity and trying to get some of these benefits I have listed down in right hand side, it's about return on investment, like fast time to market. So this day and age, every business, when they have a concept, they want to get that concept out to the market as quickly as possible. That's very important.
And it's about enhancing customer experience. It's about providing quicker deliveries and keep on innovating. Innovation is key.
Innovation is now part of the blood of any organization is DNA. And this require us to do continuous, uh, changes into our environments, production environments. And it's about increasing efficiencies in improving organizational performance.
And also it's about reducing risk, right? So this day and age, a lot of organization takes some level of risk, embracing risk, but we also want to ensure the risk we are taking is controlled as well. And this all justify why Dora metrics are still super important and what sort of ROI it's generating.
And if I go a little deep, I'm pretty sure everyone is aware just to set the context clear. So when it comes to Dora, we have four metrics. One is about, uh, change lead time.
It's about when you get a requirement, how quickly you can code and how quickly you can test, how you can package it, and how quickly you can deploy it into production environment. It's about how frequently do deployments. So do you have that trigger, uh, to do, uh, frequent deployments?
And when you are increasing the deployment frequency, can you do that in right? First it's about, uh, reducing or minimizing change failure rate, and it's about the recovery in case there's a failed deployment, how quickly I can, uh, fix it and roll it forward, or how quickly I can roll it back so that I minimize the impact and once deployed, if there are issues in the production environment due to that, how quickly I can identify how quickly I can fix it. So we sometimes call it meantime to resolve as well, but overall it's about how quickly we can get this business concept into the, uh, production and how quickly how, how fast we can run successfully.
And in case of any failure, how quickly we can restore things. So Dora maturity and observability go hand in hand. So observability, the definition is can you understand internal system state by looking at the telemetry data, which is your logs, metrics and traces With Dora in mind, it's about can we leverage observability to enable teams to optimize performance and achieve high maturity of Dora metrics?
So how can we achieve, uh, the ability to, uh, do a, uh, minimize the lead time and ability to faster deployments, ability to manage or reduce deployment failures and ability to understand issues quickly, detect quickly and resolve things much faster. So those are some of the, uh, value added observable bringing in. So nowadays, uh, every development team, if you're a developer, you need to have visibility.
So you want to understand what's happening and you especially, you want to understand what's happening in your system and you have to know what are the bottlenecks, what are the issues in my pipeline and what can I do to understand, oh, identify some of these failures much faster. So those are kind of like a very important requirements for any development team. And when we build this observability into our systems that pretty much facilitate these requirements, uh, uh, of our development teams.
So now setting the context clear observability is key part of, uh, enabling DORA maturity. You had to observe what's happening and that will actually lead you to better outcomes and talking about it. So when we used to say observability, uh, we had to also understand how exactly it's going to impact some of these, uh, DORA metrics as well.
So if you go through the metrics, if you look at deployment frequency, so observability enable us to understand some of the bottlenecks in the deployment pipelines, understand why sometimes deployment pipelines are going in a little sluggish performance, optimize them, optimize workloads and ensure that we have better capturing management of these pipelines. So that can help us to improve some of this deployment frequency and then lead time for change. Uh, the observable can provide more visualization into our development cycles, understand the more the process of deployments and provide a lot of data which can allow us to understand, uh, how it is going.
It can be about the area of the change, it can be about area of testing we had to do. It can be about the, uh, what is required, uh, for particular code and how it is behaving. So if you have more data, understand what's happening when you deploy that code in your development environment, you are better prepared to un and fast track that deployment to, uh, production and also, uh, observable to help us to understand the failures, uh, especially being proactive when it comes to failures.
And then, uh, super accelerate some of these fixes, which is about time to restore. So we are able to use observability logs, metrics and traces to develop, uh, especially metric we use using metric. We can develop better alerts using, uh, logs.
We can look at it's kind of audit frame. We can exactly go through and understand what's happening in the system. And using traces you can find out what's your code doing, right?
So with three of these your logs, metric traces, you have the power of identifying things much faster and fixing things much faster because you can getting into that root cause much quicker. So overall, it's a combination of everything at observability providing which help you to improve your adora metrics such as deployment frequency, lead time for change, change, failure rate, and time to restock. And we'll also look at, uh, one of the newer concept or trend happening.
So when we see observability, we think observability is more like, uh, uh, at the production then. So it's about can I make my systems more observability observable, which is front end user facing. And we think observability is only related to the production systems.
And because of various reason, like, because when it comes to login, we kind of like, uh, sometimes login is expensive, so we just enable a certain level of logging and production. Sometimes, uh, the sample, the traces or sampling, we use sampling to control, uh, in production because observability is also costly as well. The metrics we try to kind of like look at things in a more simple manner.
So because of that, because observability is also literally a cost we have to incur, we generally look at enabling observability in production systems only. And that is a classic, uh, mistake because observability is required in our development environments as well. 0 is more of a shift left, uh, philosophy where we want to enable observability at our environments where our developers are working so that developers have more observability or able to observe what's happening, the things which are developing.
So it's about acknowledging that observability is evolving and that all has to be reflected that, uh, uh, some of these, uh, low environments where the areas where developers working so that they can troubleshoot. So recently I spoke to Thomas Johnson, who is the CTO and co-founder of multiplayer. 0 driven products to provide better capabilities to developers so that developers can also leverage this observability advantages bringing into the table.
So generally what's happening is observability is more for site reliability engineers or DevOps engineers, but we also want to get this, uh, those advantages, those capabilities to our developer friends as well. So some of the key benefits. 0 is able to provide real time context, risk rich insights and let developers understand this unknown unknowns in, uh, the codes they are reading.
And especially when this code is running not only in production environment or in their test environments as well, and we want to enable fast debugging. So as I said, the logs, metrics and traces are very important and it help us to understand and understand root causes. And sometimes even though you may go for more of, uh, flexible the log level or deep locks, uh, level, uh, but you might still not go with traces because sometimes traces can be a little expensive if you go with the observability tool or, but if you can open source it, it's much better, but it still, there can be some cost.
But if you can work around that and then enable full the capabilities of traces that can help developers do faster debugging because they now understand how they are code runs, like it's all about the code, how fast the code running, what are the bottlenecks, and then, uh, it's about understanding, uh, your systems better. So it's not only been about the wisdom of production, but it's about wisdom of your, uh, the test environments or developer environments as well. How clear how clarity you have related to these environments.
So this is about observable. 0 is all about enhancing developer experience and naturally taking this data and insights and able to troubleshoot. So this will all align with Dora maturity.
00, uh, uh, enhanced developer experience, that will naturally result in high DORA maturity as well. So with that, uh, I used to think when you are looking at observability, it's always better to have a plan, uh, with the plan, you know, where are you and you know the destinations you want to go. So observability, uh, it's around maybe around five level.
It's about keeping the lights on. It's about observing things. It's about enabling your, the tracers or real user monitoring service maps and enabling DOA metrics.
It's about measuring a customer experience, defining service level objectives and those things. And then after that, you go into more of a correlated level. You look at things in a holistic fashion, you look at metric anomalies, log anomalies, it's about using AI to understand, uh, reduce noise.
It's about, uh, correlating things much better. It's about, uh, rule-based issue resolution. It's about baseline, uh, the and correlating things so that uh, you are little high-end at what you do.
And when you go to predictive, it's about using AI to do self diagnostic and then self feeling and powered by gene AI as well. So in this maturity model, uh, generally what you do is you improve your observability maturity from left to right, and there are areas where you will focus. So we have to look at CI/CD observability and we had to look at can I improve run type called performance monitoring and enabling DORA matrix plugging to your observability setup as well.
So all of these things are very important. Once you do not only you get match at observability naturally, one of the ripple effect is your DORA metrics get to high end. So you naturally becomes a a, a light DORA practitioner, right?
So if you see the DORA maturity get increased while you go through this journey, so I'm requesting all of you, when you are starting your roadmap as observability roadmap, build something like this, try to understand where you are and based on your business requirement, based on your technical requirements and what is possible, develop a plan like this and, and also try to focus some of these Dora specific things as well. Naturally, re regardless you focus or not observability will provide you that benefit. But if you can also bringing this concept like CI/CD observability, the core performance run tide core performance, and this especially specifically enable this DORA metrics and sometimes some organization have this challenge, it's good, but how can I enable it, right?
How can I get IT systems to reflect it? Because you don't need manually calculating this metrics and do that manual work because it's a overhead, it'll not reflect the true picture and it's not automated and it has lot of drawbacks. What you want is you want Dora XB visualize and plugging in and you get this data available quickly.
So if you look at it, this is a sample dashboard of Dora visualization provided by Datadog. So you can connect your, uh, Datadog instance or observability tool with the deployment pipelines and other systems Git and your code reports and other system. With that, you will see this data in your fingertips.
So now with this, you can go and start, uh, looking at where the bottlenecks are happening. You can start drilling it down to different areas you can look at, look and set some goals, and then you can go through your roadmap. So this is a nice way for you to, uh, visualize things and always visualization is better.
And this allows you, uh, to centralize this Dora capturing. So this is one of the key thing, part of observability you can enable. And then the drill down is pretty straightforward.
If you see you are struggling with your lead time or change lead time, you can drill it down, figure it out. If you see meantime to restore struggling, you can go inside, understand the incidents, understand why it has taken time, understand how, uh, we could have done better, right? And probably it's about awareness.
It's about getting everyone awareness into how to use observability, how to use this data to better, uh, provide, uh, a faster resolution or improve overall system performance. So it's important that when you are building your observability framework, be mindful of this and also enable this kind of visualization as well. And finally, before I wrap it up, there are some best practices.
You have to be mindful. So it's about tracking the metrics. So you have to understand what are the metrics you need.
It's, it's all about, uh, developing metrics which are correlated with your end user experience so that you understand what's happening, you have a clear ROI and you understand what is the benefit your end users are getting. And then after that you have to start enabling some of these Dora metrics like deployment frequency, leave time for change, change failure rates and other things. So that way you can go deep into some of these metrics.
And also when you are doing that, you have to understand sometimes we go with sampling or, uh, the the ation. So that will not give you some of these data points or you can miss out some of these things. So have a balance of, uh, limiting or sampling.
So that's always important. 0, which is about improving developer experience, right? So shift left the observability so that these good things not only happen at your production, but this is happening in entire your, uh, ecosystem from your, uh, local develop environments to test environment, to E two environment, to use a accept an environments to, uh, staging or, uh, pre-live environments to production.
So that way that you have the full stack or the full environment visibility and developers has these capabilities to enhance and improve themselves. And while you do this, you have to look at what automated alerts you can bring in. So it's all about metrics.
And once you have metrics on top of that, you can build a comprehensive, uh, alert mechanism. So you have to get that advantage. And it's about bringing this end-to-end visibility.
Not only the production, as I said, generally people's always thing observability is with production. 0 enabling this, uh, developer experience, shifting it left into invisibility so that a developer knows when he, the time they start coding. And until that code runs in production, what's happening.
So that is, uh, very rich, uh, experience and very rich data as someone can see. So those data, the visibility we had to provide end to end to everyone in our project teams organization should have that view so that we are all in one place and we are all able to plan and improve together. It's about bringing Observ built into your CI/CD pipelines.
So your pipelines are more transparent. You see the speed of deployments and you see the bottlenecks and you see the failures and you can improve as well. Sometimes those are silent killers.
There are a lot of areas where you can improve and that can, uh, provide better ROI. And it's about data. So observability is all about bringing in rich data.
Either it's metrics, logs, traces, so that data is very important that will provide you a lot of insights why you, why you are struggling with lead time for change, why you are developers taking time. Is it about the coming up with design and coding or it's about troubleshooting issues or it's about trying to figure out the, how that code is plugging in or it's about that one issue took entire your time that last week, right? So it's all about looking at data and how this data can help you to do better coding, understand the issues, coming up with fixes and speeding up deployments and fast track things into production.
And once you have that data, the good thing is on top of that, you can build and bringing in ai. So you can use ai, uh, ops or, uh, artificial intelligence for IT operations. It's about running AI models on top of this data so you can plan it.
So you, you can use anomaly detection or you can use forecasting or you can use correlation or noise reduction. You can get, I mean when it comes to ai, almost all the benefits of AI running at other end, all the production is well known, but you can use it in some of your DORA metrics as well. You can use some things like forecast into critic, like based on the, uh, the scope of the coding you do and uh, what you think, like what is the realistic forecast, how quickly you can go and what are the bottlenecks and what are the anomalies, right?
And anomalies of forecasting apply into door metrics, lead time for change or the deployment frequencies and other metrics that will give you, uh, a sense of being on top of this game, right? And you can definitely use AI and especially gene AI as well to improve some of these areas. And it's about gen, it's all about creating new content.
You can use this and leverage to do powerful things. And then finally, not to forget about autonomous operations, use self-healing when you have issues. Let, let's get, uh, uh, systems to heal itself.
And that can help you in long run maturing your system, especially if you remember, or if you can go back to the time where, how much time you spend troubleshooting some of these regression issues and how much you, uh, time, uh, to spend to bring up the system because system just crash once you de had that last deployment. So if you, you want to build this self feeling and the remediation capabilities not only in production but in your low environment as well all together will give you more benefits and achieving and going your door, uh, journey. 0, you enhance the developer experience.
This will provide you, uh, uh, ability to do much, uh, faster, uh, requirements from your end users, develop, test and package and deploy faster into production. Ride first, reduce failures and in case if there's any failures, of course there can be you detect quickly, fix it quickly so we can run it much faster. So thank you very much for spending this time.
0, get observability in front of your developers so that you can be a, a Dora Mature organization. Thank you for taking this time to listen. If you have any question, you can put it to the chat.
Probably I'll try my best to get back to you. If there's anything you can probably try to reach out me as well. So there are a lot of, uh, good presenters presenting various topic related to this area.
So take time to listen and it was my pleasure presenting to you. Thank you.