Revolutionizing Cybersecurity in Regulated Industries with Dynatrace’s Alois Reitbauer
Dynatrace is breaking ground in cybersecurity by integrating causal, predictive, and generative AI to deliver actionable insights into complex IT environments. Alois can discuss lessons from recent outages like CrowdStrike’s, where Dynatrace’s AI tools supported rapid recovery, as well as the challenge of balancing AI innovation with regulatory compliance. He’ll also cover how Dynatrace builds trusted, reliable GenAI environments for regulated industries, ensuring safe, accurate outputs.
Transcript
This is Textron tv. Hey everyone, welcome back here to techron tv. I'm God first timer on our show right now.
He's never been on before, so I'm really excited to Amont. It's AOS Wright Bauer. AOS is the Chief technology strategist at Dynatrace, LOS.
Welcome to Tech Trunk tv. Thank you for having me. It's our pleasure.
So, chief Technology strategist, that's not a title we hear every day. We hear chief technology officer, we hear Chief strategy Officer, but chief Technology strategist, tell us a little bit about that role and, and maybe a little bit about yourself, how you got to sit in this chair at Dynatrace. Uh, so yeah, I think the role is kind of related to what you were talking about.
Um, but that's how we, how we came, um, to introduce the title. It's somewhere like in between what you do and otherwise it's very much focused as the title says. So very often chief strategy of officer is very much focused like on the overall business and it's more of a business focus, um, uh, that the CTO like focused overall.
And what I do in my role is as the chief technology strategist is more or less work on new product innovations, which we want to integrate into the product. Also, from my background, uh, I come from a computer science background, but I have worked in pre-sales, in post-sales. I've done product management, uh, released several generations also of the Dynatrace product have actually been with the company for like roughly 17 years, which is like really very long, but it doesn't, doesn't feel that long simply because it changed pretty much every two years what I was working on.
But the, the trend was always kind of the same. There was some new technology coming along and we as a company have to figure out what are we going to do? Like what's our tech strategy there, like kind of implying the technology strategy piece there.
Uh, what's our offering going to look like in this direction? Do we even go there is now the right time is maybe we should wait a year or more. What are the problems that the people are actually seeing?
And the key for this role is really to kind of like be a jack of all traits. So you have to have the technology understanding to understand what needs to be done to build a product offering. You have this need to have, to have this product management mindset to like kind of put this like on a high max is and what needs to be done by when.
But you also need to think of topics like how can we actually sell this? How can we make customers adopt this technology? Also working closely with customers, are we really solving the problem they have right now?
Where are they moving? Or maybe the product is a bit too early in the market. So this is kind of the work that, um, Adam mostly focus on helping, introducing new, uh, capabilities to the market.
And just to share some examples, uh, cloud was one of them, like in the very beginning of cloud, we thought, okay, how does cloud change the monitoring market? Like what is changing? Then we came into, oh, okay, like this whole container thing, um, what is changing?
And if you remember back into container days, you had to take some decisions. We had obviously docker there, uh, very dominantly, but then we saw like more orchestration layers coming along and we didn't know which one should we really focus on? Where should we go first?
Uh, which one will be the one that customers are mostly going to adopt? And would say the same for ai, like everybody these days is doing ai, but what's your strategy? How you to create value, how do you differentiate yourself from your competitions?
So these are like usual questions, uh, that really come up in my role. And then work closely with like the core product teams, uh, with our CTO, but also with our, uh, solutions engineering team in the field and like even salespeople, uh, being in sales calls, talking to customers. So it's a very wide role on uh, on how to, uh, on how to interact with people.
Uh, very, I'm sorry. I'm sorry, go ahead. No, a very interesting base then, because you never know what's going to happen.
That's what I was gonna say. It sounds like a really interesting job though. 'cause you've got your fingers in a lot of pots there and you're, you know, you're taking input from a lot of places and trying to fashion that into, okay, how do we, how do we capitalize on this?
How do you productize, how do you, how do you, you know, move into this from, based upon the inputs we're getting, I'm not quite sure. How do you, how does one get prepared for this kind of role? Uh, you're gonna be a little bit of a technologist, a little bit of a futurist, a little bit, you know, a little of this and that.
How, tell a little bit your background there. Yeah, I think that the 17 years of Dynatrace, uh, definitely helped there, uh, because I, I started a company was like roughly 30 people. And you know, if, if, if it's a company of only 30 people in the beginning, you kind of do everything.
I think I ran my first Dynatrace P Cs two weeks into the job. Mm-Hmm. Working in pre-sales.
I was working with the product teams and new feature, uh, was obviously also helping customers adopt technology, position us in the market, which you would now refer to as EF Rail. So it's like having like a lot of chops along, um, that journey and learning kind of what's essential in each of those jobs. What you need, how different people in different roles, uh, think because you can build the greatest technology is you targeting a totally different audience than with your current product.
Your sales team will have a very hard time selling this product. Or if you see that, okay, this is, you have a great solution for a problem, but then you realize by engaging with customers that the customers are too early in their adoption, uh, challenge right now. Um, then it's also not the right time to roll it out because one key top thing in this role is it's also de-risking, uh, technology investments.
Because if you go all in on a technology, you have that whole organization, um, more or less following through and like putting a lot of bets stare. So you want to start small is usually from the idea or even the question on, uh, what can we do all the way then to gradually move into like getting their first 10 customers on board, building something repeatedly until you Right. Roll it out wider.
So one key aspect there is also de-risking it and also knowing when you might not want to fall in one direction, uh, but in, in another one. And like how much investment kind of makes sense. Just to give, to give you an idea, like a also I think also need a long breath because sometimes you have like this big vision.
Uh, I remember we did a, a video a long time ago, it's, uh, like seven years ago, uh, how the future of observability monitoring and automation could look like and how AI would prob would actually start to automatically detect, uh, problems, find solutions and deploy them and what this future could look like. So now you have this great video, you have this vision that you can interact on, but then you have to work in a product organization. How can you deliver, uh, actual product value every quarter, even knowing that the technology isn't there yet?
Like what's possible in the next year, what's possible the year after, like breaking it down into, into baby steps. So that's where obviously the product management, how I help as well combined with this customer consulting type of mindset. I think you have to be both, you have to love technology, but you also have to be realistic, uh, very realistic of how technology can be used and how much the world or your customers in this case, uh, are are ready to to, to adopt it.
Excellent. I always, I can't think of a better person to ask about what does Dynatrace do that someone who's been there for 17 years since there were only 30 people. And in your role briefly, I I look a lot of our audience has heard of Dynatrace and quite frankly Dynatrace has done a lot of acquisitions dive there.
They've changed over the years. How would you describe the Dynatrace product line? What does Dynatrace do for our audience?
So the space that we are in today is the observability space, which had many names over the years and, uh, what it was used to be called a PM. And we're gonna have a longer discussion where it's the same and how it's different. But basically what we help people to do is to run their applications reliably and effectively and help them to, to resolve problems.
What we added over time was obviously new areas that kind of fall into this realm. The first one was user experience. The very beginning when we talked about the old application performance management days, it was very much just about the backend systems, but not what was happening in the browser near mobile application.
So we are covering this all the way to the front end so that you understand which, which users have an issue with the replication, why do they have it? And going there also beyond like tracing metric and logging by even having capabilities. But uh, we refer to as visual replay.
We can actually see a user clicking through your website or your mobile app and understanding why things are not working. And what recently also came along actually a couple years back already, is also adding security in there because interestingly, we saw that our customers were using us for some security use cases even before we had an official security offering. Uh, but we started to build out one where we detect vulnerabilities.
I'm not able to combine us with observability data to not just telling you that you have a vulnerability, but how critical is it really? What are, uh, the potential attack vectors in your replication? How can you work with them?
And then the, the other, the core aspect there is, is also ai. Because we started to realize, uh, 10 years ago, uh, when we were teaching people on a day per day basis how to do, um, operationally manage their applications best, how to overcome issues, how to remediate them. We were asking ourselves the question, why do we still need to train people?
It is a very well-defined domain. It's well understood all the data is there. Can we build a system that is intelligent enough to understand the customer's environment and do that type of reasoning for them?
In the beginning we even didn't call this ai also, we lose used a lot of those principles, um, that, that we relate to ai. That was the next evolution. And then automation came in because once we had this, a lot of people started to ask us, uh, when we built this automatic root cause analysis.
So if you already know what the problem is, why don't you fix it or at least propose to fix. And this is when we moved into, when we introduced automation, basically allowing you to automatically, uh, react to problems by triggering the right remediation actions. And the the next level that we are moving into right now is even getting more into a preventive or predictive case where we see that application behavior is kind of not as, usually it goes in another direction and preventively take steps that things are not going wrong.
So to summarize all of this, we help as, um, also stated like in our, uh, mission and vision, we help customers to like, uh, help customers to run their applications, uh, perfectly. So the goal is to have applications that work perfectly. Uh, it, it's really, uh, the sweet category is worthy.
Like they're very long term vision from my point of like self-healing, self-managing and self securing applications and uh, self optimizing. So the idea is they figure out where they're broken at a given idea how to fix them. They find a security issue and they tell you how can remediate they find some efficiencies, whether it's cost storage, it's carbon footprint, whatever you pick and give you ideas on how to remediate them.
And that's also where I see the space moving. Having been there for 17 years, we really started out, okay, this is your data. And back then that was a hard challenge, but that went away.
Now we can tell you then the next step was, okay, we can make sense of this data for you because it's too much for you to manually analyze it. And now it's, this is the action we would be taken. And by the way, we get you 80% there, um, uh, ahead of time.
So everybody basically in the across industries who has to, uh, run very large software estates that their customers rely on, um, these are the companies out there who are helping to on the one and ensure that their replications are, uh, well performing, are secure, but also help those companies to do this in an efficient way. Excellent. LOS is a great segue into what I wanted to talk to you about today a little bit more.
And that is, look, everything today is ai, right? We're all talking about ai. If we, sometimes we talk more about the future with AI than we talk about the present with ai.
But nevertheless, AI in many, in its many shapes and forms is, is making a, an impact, uh, over a Dynatrace. You guys are working on combining the power of sort of three different AI frameworks, three different implementations of ai. One is causal, one is predictive, one is generative to give better observability, better insights into these complex systems that we're running.
Talk to us about that. How, first of all, we should define what do we mean by causal predictive and generative ai? And then how are you guys working with all of them to, to help your customers?
Yeah, so let's see. As you mentioned like different types of AI functionalities that we integrated with the product. Uh, actually we started 10 years ago, um, uh, with predictive and causal ai.
Um, so what was the basic idea? Being predictive ai as the name implies, it's predicting the future behavior of some component that you have in your system. This might be the number of users on your system, uh, response times and so forth.
So you want to understand what is the most likely behavior of my system in the future. Um, and what this also allows you to do is, which is very important for any operational task, is anomaly detection. Anomaly detection basically means you're not meeting that prediction.
You were predicting based on past data how your system is supposed to behave and you have an anomaly, meaning the system behaving differently. Obviously there might be a lot of reasons why your system is behaving differently. One might be that like really something is wrong might be that then you new deployment, the application change.
But this is usually what you want to know when you are responsible for an application, what changed, what is different now, uh, from from on the way it was before. And this is really where predictive is helping you finding more or less those spots inside, uh, your, your application, your IT landscape that suddenly expose a different behavior. This is very different from SLE management.
Uh, just to give a very concrete example, your SLA might be your application has to respond, uh, within 200 milliseconds. And currently the app, the application, the application used to respond at like 15 milliseconds. That's now it's up to 100.
So you would not alert on that SLA because you're still well within this SLA boundaries. Um, however, from an anoma EU would still know I want to know about this. I mean it's almost doubled.
This is an anomaly in my system, although it doesn't necessarily have impact on the overall world application and we still stay in the box that we have with SLAs 'cause that that's what you sometimes see people confusing. Like one with the other. Not every um, anomaly necessarily means, uh, that there is an SLA violation or critical alert in your system, but you still want to know what changed in your system.
It would feel kind of weird that a monitoring system that wouldn't change tell you about those changes. So what we have by now, we know our system is supposed to behave and which components are not behaving as they want. Okay, that's helpful to some extent, but um, whether we alert on it or not shouldn't be that relevant, but just assume that you alert, uh, the biggest problems with alerts, you have a lot of them.
And the biggest challenge with alerting is one of, depending on the tools you use, deduplicating them, but putting them into an order that makes sense. And the number one thing you want to understand is what's the cause and what's the effect. A very simple example, you have three services, A, B, and C that the usual technology, creativity, technologist, creativity of naming services by the alphabet.
A is calling, B is calling C, C is having an issue. And obviously B and A are most likely impacted as well. If you would just alert on the individual service, it would get al alerts.
We systems have an issue. So people would start to fix all of those three services, although it's very opposite, only one of them is broken and two are just impacted but are not the root cause. And this is really where causal AI now comes in.
What causal AI does, it really looks at this fault domain and starts to isolate what is the root cause and what is the impact by following this tweet exactly what we would do as a human. Okay, if these two services are calling the other one and this is having a problem, yes, sure that then the two are the impacted one. And what is the root cause like for those three services?
It sounds very simple, but if we think about modern, uh, container environments, it's a whole different source. When you're talking about, uh, millions of individual services, this is a whole new story, uh, because then suddenly as a human, you uh, start to get too slow actually figuring out which service is failing simply because you need to look at the data. And that's where causal AI is really helping because you get an answer in just a couple of milliseconds versus this taking minutes, hours or in some key cases, even days depending on how complex, uh, the environment is and how the dynamic the environment is also dynamically the environment is, is also changing.
So that's the key you get of causal AI understanding how different events in the system are related to each other. And you also get an explanation why the ai, so we talked about the explainable AI here, why the system believes that this is what actually went wrong. What we started to introduce, uh, when we, uh, started to work on causal AI is what we call back then the visual resolution path.
It's more or less you can replay how the the behavior of the system changed and you could follow this causal change. It's almost like watching a video of the system fail and recover to like really get back this understanding why things are the way they are. So what causal now, uh, allows us to do is really knowing what actually went wrong in the system.
So now we know for is a component failing in misbehaving or not. Now we know what do we really have to fix? And that's coming back to that the example, the only way actually to know, to automate the fixing is you have a causal AI in place.
Because if you would only react on the alert and then would trigger some remediation workflow, a static one, you would trigger it for all of those three services. But that would be wrong because only one of them is actually failing. So you would usually try to over remediate on what's happening in that system.
We only need to fix C at the rest goes back to normal. There's no need to fix A and B. So very good.
Now we know what the problem is and now we want to find a solution. And I think especially with the rise of clouds, uh, also Kubernetes, but basically infrastructure as a service, uh, we now have this great opportunity to change the configuration of a system with markup, whatever it is, the list of terraform. Let's take some other type of scripts, use Kubernetes manifest, and guess what?
This is text. So we have text basically with some semantics behind it and we want to change it to do something differently. And here that suddenly generative AI comes in perfectly because what generative AI cannot do, we would not know what needs to be fixed.
And again, we can do this if an error occurs or jumping back to the predictive capabilities. If we already know that we are expecting, uh, twice the load we running right now, we could actively change infrastructure, risk code and start to automatically deploy these changes. So you, we see that actually predictive and causal of feeding generative fi.
And this is a thing that you see with generative AI overall. Uh, generative is only as good obviously as the model that you're using. Uh, plus obviously the fine tuning but also to the input, whether it's like rag where you add documents or prompts.
Very simple example from the real world. If you're asking for a nice restaurant recommendation in your hometown, you might usually give it a restaurant that you like most. Uh, but then I might come up, Hey, I'm, I'm actually vegan and you might come up with a different restaurant, but yeah, and my hotel is there and I'd like it to be close to my hotel because I want to walk.
So the more information I provide, the better the answer gets. And it's not that you would not know that, give me the perfect answer, but the more information you have, the better the answer gets. And that's why we like really saw the need to combine, um, these three forms of AI together into what we call hyper modal ai, where we even enriching this information.
Not saying this is the number 15. So no, this is actually a response time of a service. It used to be 10.
And by the way, when the response time is like this, that service also has like a replica count or a number of instances like this. And by the way, there also was a deployment. So we are giving it more meaning.
And by enriching the meaning, like changing more or less the modality from a pure me to something like a response time, giving it semantic and also reaching it how it like, uh, interacts with the overall system, we can much better feed, uh, generate the FI for remediation. That's one area of of the work we we do on the gen EI side. The the other area is that we allow people to, uh, easier interact with systems.
And one of my favorite examples is building a dashboard. Nobody loves to build the dashboard. I yet have to make the person who gets up excitedly in the morning saying, well today have to build two or three dashboards for production.
And so usually we build them once and they don't change them. And you have to put a lot of thinking in there. And an example that we're working with is we say, okay, uh, build me a dashboard for Black Friday, which is coming up soon, right?
So this is nothing you can template because it's very much depends on your application, how it's behaving, what kind of what it's functionality, it's exposing to the end users. You would have to put a lot of work into this. That's usually when we built these artifacts, they have to stay and survive for a very long time because there was just so much work going into it.
What we can now do though is we take this question, build me something for Blake Friday, go back to like a JI model that, what does this mean? It means basically show me all services that are relevant for my business and might react, uh, in a weird way or an unexpected way once load gets too high. Okay?
But that's basically everything chain AI can figure out. Now it has to go back to causal ai, understanding causal dependencies and what we collect the Smartscape model as a semantic model of environment and asking, okay, what are all conversion points in my application? What are all the steps people need to take along that conversion path?
Now that we understand the steps, we go one level deeper and understand all the services that are required now that we need to have all those services, we go to predictive VI and figure out how do they behave under very high load or how will this behavior most likely change? And then we get a list of services, we get a list of key metrics we need to watch. We also need, we'll understand which of those were issues, uh, what were the root cause of similar issues in the past.
Now we package all of this information up in a way back to like our gen AI components, say, build me a dashboard that exactly contains this, this, this, this, this, this and this. Um, and all of that, like heavy lifting is now done by, uh, AI hierarchy, like hyper AI behind the scenes. I think what this example also clearly shows while gen AI is instrumental being the user interface of the productivity tool, a lot of that heavy lifting is actually coming from, uh, the generative part as I, sorry, from the, from the causal part and from the predictive part.
And that's where, where we see like all of this playing together. And last but not least, there's another very important component to this. When we talk about ai, how the AI is actually able to, uh, work with the data it needs to like feed all of this information.
And, and for this we have built a dedicated, uh, data store underneath, which works without the schema and without an index and you can query it in real time. The grail, that's what the data store is called. It basically doesn't have an index.
You can send any data read that you want and can query it in real time or have varied large amounts of data and you don't have a schema either, because as you can see, the queries change depending on the use case, depending on the um, situation the system currently is in. And if you would go for a traditional approach where we would've to re-index the data, the answer would've come in way too late. And the AI by itself would not have a lot of value because if you realize, oh, that one field is key, but that end don't have an index on it.
Um, and it would have to wait that it take forever or the answers take forever. And also if I realized, okay, I need a different type of data structure, data format supported, and then the database would not support it, you couldn't use the data to like provide these answers as well. So that's like how the whole thing plays together at how we like really envision that the technical underpinnings for, for the future of those systems, again now looping back because this really enables us then to automatically go for self feeling for our self-protecting like we detect the vulnerability, uh, making some changes to the system or proposing them.
And uh, this the same obviously for the optimization part. Alois, I didn't wanna stop you because you were rolling and you were doing a great job. This thing we, we went way over time, but I think it was well worth it for our audience.
Just real quickly 'cause we've gotta run, what website can they go get more information here on this? com, but anything within the main website that we talk about this? com, pick whatever it's like most important to your business.
Again, we see people starting at different points depending on the journey. More backend, more user experience, more the automation side. Great.
Thank you so much for coming on here and joining us at textron tv. This has been great. com, check it out.
A thank you. We'll be in touch. For now, this is Alan Shimel.
Bye-Bye.