MoM is Dead, Long Live Observability: A New Era of Digital Operation – SKILup Days 2024
This session delves into practical insights on how our team successfully drove Full-Stack Observability adoption in a dynamic region where legacy paradigms like MoM still hold sway. The UAE, a hub entrenched with every technology vendor claiming Full-Stack Observability (FSO) and AIOps, presented unique challenges. We specifically targeted the MoM concept, a longstanding paradigm in on-premises infrastructure management, to shift the mindset of technology leaders in large commercial and government enterprises towards Unified Observability.
Transcript
Welcome to our talk. Mom is Dead Long Live Observability. Uh, today I will be discussing with you a very interesting topic in controversy.
However, it's very familiar to most of us, uh, as we are embarking in digital transformation. Uh, many of us are encountering a shift between what used to work and what is working. And today we'll be tackling one of those topics, mom, manager of managers.
As we all know, mom has been around for almost 25 years, or as long as I've been in technology, and it was originally brought to connect various systems and integrate them into a cohesive way. However, as we moved away from the central computing into cloud, and then further now into digital and information we find ourselves running into issues with more. So today I'll be introducing you to observability.
Uh, it is what used to be monitoring. However, as we all know, monitoring used to suffer from siloed frameworks, from separated technologies, and more importantly, uh, a multitude of tools that made our ability to connect those contexts very difficult. Uh, observability was introduced for the past 15 years, but more seriously the past five years through many players and vendors in the market, and maybe the most famous and the most accomplished is Dynatrace.
It's been around for most 14 years in a row as the top of Gartner quadrants, uh, observability is a construct that allows you to connect what typically we have as monitor data along with the context of vertical and horizontal integration. And Dynatrace is one of those platforms that has proven to be unique in how it tackles observability across the enterprise. It provides end-to-end visibility with the ability to automate capabilities and to monitor systems continuously.
Let's talk a bit about Dynatrace, just to get you all familiar with it. Uh, Dynatrace is at the core of it is based on cloud hybrid technology that allows you to store a multitude of data across various systems. So as we collect data from various traces, metrics, topologies, network and problems, we combine that through various connectors when agent being the central collection that Dynatrace utilizes, and of course various APIs and various open telemetry based systems.
That data is processed through a powered and a massively app powered processing lake house that RA has mastered over the last, uh, five years. That data is further brought into analysis. Uh, when we talk about AI and ml, there is nothing better of an example than what Dynatrace does with AIOps.
So the data is brought in, correlations are established, and context has been fed further. At the end of observability, what you're looking for is an ability to monitor, is an ability to observe behavior of the system and to automate capabilities as you proceed. Uh, one of the key differentiators between monitoring and observability is the construct of vertical and horizontal stack, so what we call as full stack observability, where you are going on the horizontal level at a service level between applications and the service flow and vertically between the application and the underlying infrastructure.
That allows you to have a holistic view of what is going on with your system in relation to the user and in relation to the transactions, of course, since we're talking about Dynatrace specifically, but what I'm gonna describe applies to any mature observability platform, uh, at the heart of observability is this construct of being able to balance between causal analysis, which is related to what's actually happening, and predictive analysis, which is more of forecasting and more of plotting anomalies in the future. Uh, Dynatrace has mastered the technology by introducing a higher productivity rate through Davis copilot, which is at the heart of generative AI implementation of Dynatrace. Uh, this platform has allowed us, uh, in some of the use cases I've described to you today, to really master a very complex landscape of enterprise in an easy way to start, but allow it to evolve with the enterprise itself.
When we talk about observability and where mom has sometimes fallen short is, is what is important. And within observability, there are key four aspects that are important. Uh, I'll cover some of them quickly before I move into the details of those use cases.
Latency, uh, the energy of performance, especially when you're dealing with hybrid systems, is this idea that transactions will take a longer time as distance traveled and as complex systems are in the picture. Uh, additionally, as transactions are proceeding, the error rates and the ability for the transaction data to transmit through layers is another measure of metrics that observability focuses on greatly. Uh, at the end of the day, uh, when Dynatrace is introduced in an ecosystem, it is seeking automation opportunities, and that's where the AI becomes very useful in taking some of the mechanical or manual steps that we typically spend hours scripting and automating that.
Uh, when we think about observability, one of the questions I usually get is why, and the simple answer is, as you're building complex digital systems, you have to create a sense of immunity within that system where the system can actually respond to changes, can be protected from laws of violation in terms of security, and at the same time being able to adapt to changes in its topology and infrastructure. And that's what typically a digital immunity system is driven by. Uh, proactive monitoring is one core capability of mature digital systems where it allows you to forecast what would or could happen in the future with some eye into capacity and availability planning.
Uh, I wanted to share with you a few, uh, use cases. Uh, those were two customers that we've had the fortune to work with, uh, in the company I work for Your Compass and Dynatrace was a powerful platform that allowed those two organizations to do what they couldn't do prior. First one is a global media company.
I could not get, uh, uh, a disclosure to use their names, so I'll refer to them as a global media company. Uh, needless to say, from an infrastructure, they are spanned across four clouds and an on-premise data center. Uh, the problem they were tackling is they had great monitoring, however, it was siloed.
It was siloed across the environment that was siloed within the environment itself. Uh, when we introduced Dynatrace, which took us around six months, three months to implement it, and another three months to bring some of the processes online, we noticed an amazing, uh, reduction of the MDTR, which is the meantime to recovery by around 40%. 99.
This occurred not because we increased their cloud capacity or we automated some of their core processes, but simply by increasing the resiliency of the systems through knowing and observing its behavior and forecasting or predicting some of the outages prior to their occurrence. Uh, as a whole, this company has seen, uh, a multitude of return investment over the past one and a half year, and now they are actually in the process of acquisition. So this was the perfect opportunity to increase the resiliency before this scale.
Second case study is a UAE government entity, and we've spent the past two years in the UAE in Dubai lobby working with a multitude of agencies on implementing observability. And, uh, for me, somebody who was coming from North America, I was expecting maturity in the private sector to exceed government and to be truthful, I found the opposite. Uh, government entities in the UAE have really embarked on digital transformation in a serious fashion.
Uh, this, uh, government entity had around 700 services that span individual consumers as well as organizations and businesses. And again, the ability to see through this dynamic multi-cloud environment was limited. Uh, we implemented, uh, Dynatrace across their clouds, and we saw a huge reduction in downtime, which used to be one of the issues, uh, these services encounter, especially that they were not all centrally managed and owned and above and all response time to incidents.
So a faster incident response time that has increased by 25% in less than eight months. Uh, overall, this really helped their operational efficiency and really, uh, made a dent in their cost of operating these clouds and as a result allowed them to expand their services. Uh, when we talk about these examples of operational efficiency or cost efficiency, uh, at the end of the day, uh, the value behind digital, uh, transformation lies in being able to provide capabilities in a secure way.
And there is nothing more, uh, I would say risk inducing than a highly sophisticated, highly automated system suffering from security issues. Dynatrace especially implements security as part of what it does in observability. So at all layers of monitoring at all layers of detection, security is in the context, whether it's application security or it's integration between layers.
And that allows you to actually provide proactive threat mitigation that starts before vulnerabilities are detected on the outside and it continues throughout the lifecycle of the system. At the heart of the system is, uh, a famous or to become famous, uh, backend called grail. Grail is a re-envisioned, uh, highly scalable data lake that Dynatrace has reimagined from the old R-D-B-M-S to really an unstructured data store that allows you to manage both the transactional data as well as the analytical data in one place.
This really allowed you to start thinking and rethinking even some of what is possible using transactional data versus, uh, analytics. Some of the key lessons, uh, that I wanted to share with everybody, uh, with respect to if you're embarking with this journey of observability, is start small. Uh, platforms such as Dynatrace allow you to do that, what you are targeting, certain parts of your ecosystem, and gaining some understanding and familiarity with observability.
Once that is in place, then you can scale forward. Second part is complexity. So when you're talking about observing, it already exists in highly complex system.
There is no way beyond going through autodiscovery, you cannot manually map out these systems. So you have to look at a platform that actually does that, and Dynatrace really does that in a superior way, where after enabling agents on servers, you as we say, watch the magic happen, where they actually start discovering various parts of the system and start looking at the logical relationships between these systems horizontally as well as vertically. Uh, this is a change program.
So as you introduce observability in a traditional, uh, ITSM or IT organization expect challenges and be prepared to overcome those challenges. This is adoption, uh, challenge. This is new technology and old habits.
So you have to be willing to look at some of those challenges and approach them intentionally, approach them heads on you cannot avoid some of the, uh, process changes that must happen and above all skilling. Uh, one of the things I've seen, uh, as a big challenge in organizations is when you are faced with new technology and old skills and human nature does resist change in all forms. So anticipating that through upskilling is really one of the key, uh, strategies we found work within our journey in observability.
Uh, I wanted to also share with you some lessons learned. Uh, Dynatrace or any other platform for that matter, should not be looked at as one size fits all. It is not a silver bullet, but it is one part of the puzzle.
And if you choose the right observability platform, and in my case last two years, it's been Dynatrace, you realize that a platform that takes integration as a primary concern is really equipped for the enterprise. You're not gonna replace all the tools, nor should you, but you have to be able to start leveraging the data, coming from all the systems and be able to inject insights on all these systems. We've seen that with the security systems.
When we enable that and trace, we see that with automation. We've seen that with DevOps, with DevSecOps, so definitely that one of the lessons we've learned is seek integration from the beginning. This also improves your adoption and actually gives you a better return investment.
And thank you.