Metric Mischief: An Amusing Adventure in the DevOps Universe | DevOps Experience 2023
It’s been more than a decade since the term DevOps has become popular. As of today, there is probably not a single IT organization who has not embraced DevOps in some form or fashion. One question is often asked by business and senior leaders: “What should we measure to know if we are doing well?” The world of DevOps metrics can be both perplexing and amusing. There are many metrics sets, measurement techniques and frameworks that have been documented over the years – DORA metrics, Flow metrics, SRE metrics, DevOps maturity models etc., to name a few. But large enterprises still struggle to establish a uniform set of metrics to measure their “DevOps Success” across all teams.
In this session, I will explore the confusing side of DevOps metrics, acknowledging the inherent short-comings and occasional misadventures that arise along the way. Through entertaining anecdotes and relatable stories, I will showcase the challenges faced by all when trying to make sense of elusive data points, bewildering performance indicators and unexpected correlations that can turn out to be erroneous.
I will also add a few additional measurements that can be used in conjunction with, or in place of, the existing ones. I will discuss some key metrics that matter in the DevOps context, and how they can be aligned with overall objectives to drive success. Moreover, I will emphasize the importance of contextualizing metrics within the broader organizational and cultural context, avoiding the pitfalls of vanity metrics and focusing on meaningful, actionable measurements.
Transcript
Hello, hello, hello. Uh, good afternoon, good evening, or good morning from where we are. So thank you for joining my live talk at d o e 2023.
com. Uh, thanks Strong, uh, for inviting me. Really feel honored to be here.
And my talk is going to be, uh, on, on, uh, seemingly funny enough article, uh, topic, which is, uh, about DevOps Metric. And the title of my talk is Metrics Mischief and Amusing Adventure in the DevOps Universe. Now, uh, I'll tell you that I did not choose the title myself.
I actually asked Chad g p t to generate a title for me. So that's where you are. Uh, so it is, uh, title generated by chat G p t.
Before I start, I want to put out some disclaimer or disclaimers. First of all, this stock has nothing to do with my current employer Fidelity Investments, and it is just a point of view based on my knowledge that I gathered from my industry friends. Uh, and many of you are already on this, uh, live talk today.
So thank you so much for my learning sessions. I do not know which metric work yet, but I do know that the, that there are certain things that do not work or should not work, or you should not be using, and I'm gonna talk about those. And I also know that something, uh, uh, that I have in my mind that may work for you.
So it's up to you whether you want to use that or not, and I'm gonna talk about that too. Now, my name is Topo Pal. My actual name is toa.
Uh, I go by Topo. Currently I'm a vice president in, uh, fidelity Investments Enterprise Architecture Group. However, as I said, this talk has nothing to do with my, uh, current work.
It's just my own point of view. Um, after I, uh, I did my PhD in semiconductor physics. I spent some time in researching in Nanomaterial area.
Uh, I was actually teaching, uh, physics in a, in a, uh, major university or college, if you will, in India. But then I switched to it, and, um, and if you ask me why, uh, I'll tell you that's a different story. We'll skip that for now.
Uh, I joined, uh, uh, my first, uh, IT company that I joined was, uh, called Circuit City, is no longer there anymore. Uh, I actually had, uh, all kinds of hats, uh, while being at Circuit City. I proudly own them.
Um, uh, and I was a developer. I was in operations support. I was an architect and many other things.
Uh, and then I worked at a, a large retail company for a while, and then joined almost, uh, a startupy healthcare software company. Um, and then I joined Capital One prior to joining ca joining Fidelity, I was there, uh, for, for about 10 years where I, uh, started my DevOps journey at Capital One. And, uh, uh, I did that until I left Capital One in, in 2021.
Uh, I was always interested in some sort of automations in the software delivery lifecycle. So, uh, I'm, I was very happy that, uh, I was given that role at Capital One and also at Fidelity Investments, uh, there are a few things that I'm still proud of. One is this, uh, HI DevOps dashboard, which actually measures, uh, uh, many things around your DevOps practices across your, your pipeline, or across your, your, uh, storyboard and your deployment, uh, environment.
Uh, so that was open sourced in 2025, uh, 2015. The next thing that I'm, uh, proud of is, uh, a paper that, um, I was a co-author of, uh, it's called dig Ops, automated Governance Reference Architecture. It was published by IT Revolution.
Uh, it came out of a group, uh, working session, uh, that, uh, it revolution hosts, uh, yearly, uh, uh, at, at Portland, and it's called DevOps Enterprise Summit Forum. So back in 2019, um, a group of us got together and started documenting our writing white paper or reference architecture around DevOps, automated governance that came out to be a book that was published last year around, uh, September. Actually, September 13th was the, was the right date.
It's a novel about DevOps, security audit compliance and, and, and how to thrive in the digital, uh, age. Uh, there are nine authors, and I was, uh, one of them. Now, having said that, let me kind of set the context of my talk, which is around DevOps measurements and metrics.
Now, DevOps is not a new world. Uh, it's, it's there for more than a decade. However, if even now today, if you ask anyone, what there was is you get, uh, an answer, which, um, kind of looks like this.
It's a big elephant. And depending upon who you ask, you get different answers. So the confusion around the DevOps definition is, is going to be there.
It has been there, uh, and it's not going to change. So instead of asking, what is Dev DevOps, I, I started asking myself and answering the question, of course, why is DevOps the answer that I came up with? Is DevOps is needed because we want to deliver high quality software faster.
Now, I think that is something missing here, um, in, in, in terms of, you know, delivering high quality software faster. There are certain cultural aspects to it that was missing from here. So my friend John Smart came up with this one better value, sooner, safer, happier, and I think this one covers everything.
This, that DevOps is for value, speed, safety, and above all people's happiness. Having said that, I'll tell you that every enterprise is different. There are different sectors like financial, retail, healthcare, technology, energy, defense, and so on.
Even within a sector, say financial, different enterprises have different structures or unit that are put together in different ways within any such unit. They develop and operate different systems based on various combinations of technology stacks. Every enterprise has unique set of cultures that set them apart from their counterpart competitors.
Uh, and these cultures are not, uh, around geographical locations. Uh, it's beyond that. It's about different work environments, different mindsets and different business goals and different, uh, technology stack, et cetera.
Many such enterprises were established more than 50 years ago, or 75 or even a hundred years ago. And on the other hand, we have companies that are born on the cloud. On one hand, we have mainframes and cobol, and on the other hand, we have serverless, cloud, and rust programming languages and all of that.
So most large traditional enterprises have everything in between these two extremes and including those two. So I started thinking about how to actually fit DevOps in into, uh, an enterprise. And the conclusion that I came up with is this not one size fits all.
And we all know that given the enormous, unbelievable diversity, there's no way that a single DevOps model or agile framework or a unified platform can satisfy every need. Not even close. DevOps in a box or agile on a page does not exist.
I mean, they exist, but I do think that they don't work. So how to actually, uh, get onto DevOps practices from a large enterprise perspective. I think the answer is this, build your own DevOps, uh, copying other transformation models into yours will not work as is.
There's nothing like our DevOps in the box. We have to build our own DevOps. I've seen in many instances, more than once, when a company a tried to copy company B'S models and failed, it goes something like this.
Uh, Spotify has a model working for them, for working for them. Let's use that. Or Google puts all of all of its source code in one single code repository, so we should do that too.
Or HC does 10 deploys per day, and we should do that too. Or Capital One being a bank is going all in cloud and should do too. You know, all these type of discussion and debates.
The biggest thing is that we all miss the context. Is your company trying to solve the same problem that Spotify wanted to solve? Can you do everything else that Google does besides putting all the code in a model depository?
Do you know that Capital One is just 25 years old and it does not have a lot of mainframe applications? And so I think the context really, really matters here. However, when you build your own DevOps, one set of questions are going to be asked on, and they're going to be common.
Like, how do you know we have done the right thing? How are you doing? Where are we in our journey?
Yes, we need measurements. We need to know how we are doing. We need to know how to improve.
And there comes this thing that measuring nothing is not good. But again, the question is what should we measure? Measure?
There are wide ranges of metrics and measurements. Uh, and I'm going to go through some of these and they all measure different aspects of DevOps and agile, agile practices. Dora accelerate metrics, and I'm gonna talk, uh, uh, in great detail on that one.
Space metrics flow framework bond down of velocity charts, et cetera. Developer productivity metrics. This is the latest one that is getting a lot of focus on Everybody's talking about developer productivity metrics, and I know that there are some, uh, threads going on, on, on, on, uh, Twitter or, uh, currently known x, uh, or, or even LinkedIn about some companies publishing ways of measuring developers productivity.
There are pros and cons in many aspects, but I go to my fundamental question about developer productivity. Do we have a list of the things that we want our development teams or deli developers to produce? If we don't, then how can we measure productivity?
But again, this thing is new and lot to be seen, um, as to what comes, comes up in the industry. Metrics is another one. And DevOps maturity models.
Yes, I know that many companies use them. Um, there are, uh, pros and cons of using maturity models, but if that works for you, so be it. However, in these all, in all these measurements, I always go back to my, my, uh, science basics and I do have fundamental questions.
Uh, these are the questions. What, why, who, when, how, when, and, and whom. I think every measurement has a time and precondition in that context.
There is no point measuring lead time, let's say, when there is no CI pipeline or source control. Uh, we need to be careful about what we are trying to measure when we are trying to measure why we are trying to measure who we're trying to measure. And, and more important, most importantly in my mind, how we are going to measure.
So let's go through these one by one. What's the context? Are we trying to measure technical ability or business value or somebody's skill?
Unless we know the context, we will be measuring something and we don't know what that measure means. Next is why do we want to measure? Because someone told us is that some vendor told us that you should, uh, use DOA metrics or develop a productivity metrics.
Or is it something that coming from within us, somebody within our, our, our enterprise, uh, or in a company, uh, uh, domain is asking that, Hey, I think we are lacking in something. Maybe we should stand up some measurements and, and, and try to check. Or do you want to actually improve something?
The measurement tends to be always to show off in my mind as opposed to improve something we want to look good as opposed to look bad and try to improve from there. The next thing is who are we measuring? Is it some individual?
Is it the team? Is it the product? Is it the business unit?
Is it the company? Because many of these measurements go loosely on exactly what you are trying, trying to measure or, or who you are trying to measure. If we do not know who to improve or what to improve, uh, we do not.
We will not know what actually to measure. Uh, I hope nobody's me measuring individual, uh, productivity or individual skill level or anything of that sort in current day. Uh, I do know that people are trying to measure teams or the teams that are, that are, uh, underneath the product umbrella or within the business unit.
Now, there's another thing that I want to, uh, discuss is, is when should we measure? Now if I draw parallel to, uh, human life, human life is born, then it grows up, then it match yours and starts working, and then the, you know, the, uh, we, we we retail and then someday we die. Uh, a typical application actually follows almost the same life cycle, right?
An application is born and then people start, uh, developing on that, and then, uh, it gets matured. At some point the application ages off, and then it needs to be repair and replaced with something else. Now, depending upon where you are in this lifecycle, I believe that your measurement should be, ought to be different.
Because if I drop arrow to human lifecycle, again, when I was born, the things that were measured on me, I don't know what they were, I don't remember them, are not the same things that I measure on myself every day. Um, and, and, and, and, you know, take any lifecycle, uh, stages, and you will see that there are different things to be measured at different stages of the lifecycle. So is it the right time to measure certain things?
Are there preconditions that, hey, if I have to measure, let's say deployment frequency, are there any preconditions to that measurement without those preconditions to be made? Or if it is not on the right time cycle or, or right, uh, life cycle, the measurement would be meaningless or actually be false positive in, in many times. So the last important thing that I call, I want to call out is how should we measure?
Are there standard measurement techniques for that measurement? And I'll actually hit up on this thing when I go through some of the donor metrics, uh, uh, metrics, uh, items. So the thing that I wanna call out here is that unless you know how to actually measure these things, your measurement may be a false positive.
Uh, there has to be an unit, there has to be a unit of, of counting some of these, uh, metric item. So wanted to call that out loud and clear. The next thing is, who is it for?
Are you measuring something for the managers, for the developers, or for the executives? And what are they going to do with that measurement? For example, let's say my CI pipeline bill or execution time.
Well, there's no point measuring that because developers already know what they are. And if you are saying that it needs to be measured for the managers, my question would be what are they going to do with that measurement? Or even executives?
What are they going to do by knowing that, uh, pipeline execution? I, uh, execution time is 10 minutes versus 15 minutes, are these actionable? And then the other thing that, uh, comes to my mind is, are you trying to measure something or are you trying to get something that is an indicator of something?
For example, if I go to a furniture store and if I want a dinner table, you know, there are two ways that I can express what I want. One is, hey, I want a dinner table, which is like, let's say six feet in diameter. If it is a round, round one, or I can say, I need a de table for eight people, right?
I don't know which one is easier for you, but for me, hey, I need a dinner table for eight people, is the right way to go? Uh, and then there's another thing you need to keep be careful about is, is the observer effect. Uh, observer effect is nothing but disturbing an observed system by the act of observation, meaning that when you say that I'm gonna measure something, are you actually disturbing the system so that your observation itself becomes wrong?
Uh, for example, let's say I use lines of codes as some kind of productivity metrics, then just by saying to the developers that I'm measuring that maybe I'm asking the developers to do work around or gain the system by putting in more lines than it is needed, thereby actually producing, uh, or score than it needs to be. So that observative effect is, is, is, uh, you know, a very much, uh, uh, a thing that we need to be careful about in our mind before we start measuring the other side of this is true. Also, just by saying that I'm observing something, you may bring out some good practices.
That is true too. Now, with that, I want to say that every measurement that we do needs to follow the way the verbiage or the pattern that we use for writing agile stories as, or something, I want to measure something. So that's something which basically means that we are actually saying here who the consumer of the measurement is or metric is.
And I want to very clearly state what I want to measure so that I have something in my mind behind that measurement. I want improve something, I want to change something, I want to reduce something I want to accelerate, uh, or, or, uh, make something more, more meaningful or more visible, whatever that is. Now with that in mind, I want to go through the Dora metrics, for example here and start on answering what, why, who, when, how.
All these questions that I just asked as fundamental questions before I go there, I'm wanna go, uh, just show you the data metrics. Uh, again, uh, for those who do not remember, I know everybody knows about them. Um, it's deployment frequency, lead time for changes, time to restore service and change failure date.
And there are three other or four other columns on the right side of it. Um, there is a benchmark for elite performer, high performer, medium performer, and low performer. Now, I'll come back to this, this, uh, benchmark for Elliot High and medium and low.
When I go through my, uh, uh, basic questions on what, why, uh, how and all those things, the fundamental questions, but remember that, uh, I wanted to call out these columns because I have something in my mind. I want to point once and, and kind of, uh, point out to you why you should be careful about using these metrics. Now, before I start, uh, uh, uh, the critical side of of, of these Dometic, I want to call out that accelerated DOA or DevOps, uh, research is, is this book, this is the book I'm showing here, right here.
You probably have seen this book, read this book many times. I actually keep it on side of my working days so that I can read it every now and then. I think it is one of the best books in DevOps period.
Why? Because it has 250 pages full of data science reference recipes and all that. And it covers all aspects of DevOps culture, leadership, technology management practices.
It has 24 capabilities listed with all good explanations that can help drive improvement across the enterprise in any, any, any, uh, size of enterprise. My, my, my problem is that we all ignore all these useful information in this whole book, 250 pages. Uh, and instead of, uh, focusing on these 24 capabilities, somehow we focus on only those four things and say that only those are important.
Rest for the book is not important. And the last thing I wanted to call out is I was the first customer of this research. So I still remember, uh, why I wanted this research, and I think it's, it's very useful even today and it'll be useful for tomorrow and, and days to come.
Now again, this is DODA metrics, those four keys, and I want to talk about those and, and balance that or balance that across, uh, my, my fundamental questions, deployment frequency. I think it's very much undefined. Uh, when I talk about deployment frequency, there is a notion that I need to count them.
Now, the things that I don't know is what is a deployment? Is it one application deployed across 500 servers? Is it a network change?
Is it a feature flag change? Is it all of those deployment of what, uh, is it a deployment of application, the whole stack or the part of the stack or just an infrastructure change? How do you count them?
If I deploy one application to 500 servers, is that 500 or is it one? And this is important because based on the counting, I could be either Elliot or medium or low or good performer. So unless you know how to count them, you know, I tend to gain the system and say that if one application is deployed to 500 servers, I deploy 500 times a day.
But is that right? I don't know the answer to that. Next thing is who are we measuring?
Is it the team or the teams that are actually behind those applications? Development of those, of those applications? When should you measure it?
And again, I talked about this development life cycle or application life cycle application is born, it gets matured, it gets used, and then at some point it gets retired. When you try measuring deployment frequency in the early stage, then probably you have a good number, like 500 or thousand. But when you actually try to return that application, I should not be deploying that after maybe some patches, maybe security vulnerabilities.
And that's it. That does not mean that the teams that are supporting all the, all the application itself is, is gone from a to low, low performer when you should not measure this. As I said, when you are trying to, uh, retire an application, then stop measuring it.
And again, if it can figure out the push answer to the first three questions, then probably it's a good thing for you to measure. The next thing is lead time. Commit to deploy.
I think it's very much incomplete because it only measures the pipeline execution time, which is commit to the release branch, to deployment to production. It can, it does not, uh, uh, incorporate the time span before it gets mar before code gets MAR merged the release branch. And we all know, all the developers know, we spend a lot of time in designing, actually making the code work, writing the unity cases and all that.
So if we completely disregard that, my lead time is nothing but my pipeline execution time, and if it is that, just say so that it is pipeline execution time. And then I would argue, does it matter if it is 15 minutes versus one hour? Maybe it does when there's a big incident, but overall it does not, uh, time to restore.
I think DevOps are not, this is a classic metric. DevOps are not. It, it has been used in used for decades in IT industry.
This is nothing new. Change failure rate. I think it's a good measure, but can be confusing because there are some fundamental question, if a blue deployment fails, is it a failure?
It can be hard to pinpoint in a dynamic and complex, uh, system when actually you are deploying multiple times and suddenly some bug pops out. It's very hard to pinpoint your particular release. Uh, if that happens, and this is the real life Dora metrics, and, and again, uh, uh, this is not high resolution.
I just took a screen cap from one of the internet sites. Uh, by looking at it, I actually do not find anything that is actionable. So I think this may work.
If we really want shortcut, as I say, if you don't want shortcut, please read this book, accelerate all the 24 capabilities and follow them word by word. If you want shortcut. And if you're thinking data metrics is the shortcut, I will say that there is only one measure that could accurately reflect organization's technical ability, and that is production access.
How many people have production access? Because if we say that ideally zero production access, that means you are insuring all DevOps, technical abilities, production change only for automated pipeline patching the automated pipeline infrastructure deployment via infrastructure pipeline. Everything is source code log streamed up.
All these are good DevOps practice, and you can ensure that by just targeting zero production access for anyone. In fact, I'll say that you should probably, uh, establish something called production access day. Just make up your own rule.
Let's say by default, no one has production access. That is zero debt. Zero debt, and that is ideal.
Each persistent production access costs, 25 points every break, glass to temporary read access costs, one point and so on and so forth. So you can actually establish this system for yourself and actually gamify the production access. And, and the idea is if you can ensure there is zero production access, you have insured, you have good, uh, DevOps practices, two balancing measures, just so that people don't sit idle, you also need to, um, uh, you know, measure the number of tickets that are remaining open or use four metrics for that matter, or, and along with that, you need to make sure that your, your, your actual, uh, customers are not impacted.
So measure your downtime before, beyond budget or use, you know, any of the S L I O metrics. I think there are other missed opportunity that it could, it could have number of hours spent in meetings, number of standards that are more than 15 minutes, how many teams must be involved in one feature release of on product, this kind of thing. But I think there's a new indicator and that is this.
How long did it take for you to fix lock for Shell in production? Because we have a research that says that enterprises is based lock for shell responses, have all these things. And these are all good DevOps spec metrics.
I don't want to go through all these, but I think you can take a look at that and, and see for yourself if, if this makes sense. Thank you so much. And I have about two, three minutes for any questions.





