Bringing the Mainframe to the Whole Enterprise with Broadcom WatchTower
In this presentation at Tech Field Day Extra at SHARE Cleveland 2025, Angelika Heinrich, Product Manager for Broadcom’s WatchTower real-time streaming capability, discusses the growing imperative for organizations to improve customer satisfaction and operational efficiency through enhanced observability and service reliability. As digital interactions increasingly define customer experiences, enterprises must quickly detect and respond to issues that impact application performance—especially ones involving critical back-end systems like the mainframe. Heinrich emphasizes that WatchTower was developed to align IT visibility with these business needs, enabling technical teams to better understand and respond to end user experiences in real time.
Through the session, Heinrich explains how the traditional siloing of mainframe systems poses challenges in unified observability. While modern DevOps and SRE teams rely on platforms like New Relic, Datadog, and Grafana for monitoring, these tools often lack native support for mainframe environments. WatchTower bridges this gap by streaming mainframe telemetry—encompassing traces, metrics, and logs—in the open telemetry format, a widely adopted standard across observability platforms. This allows SREs, without deep mainframe expertise, to access actionable data insights and correlate performance metrics across distributed and mainframe systems. By doing so, teams can analyze trace-level data, understand latency issues in applications such as KIX transactions, and relate these to environmental conditions or infrastructure constraints.
Heinrich further demonstrates how this unified observability framework empowers operational roles to detect and respond to service degradation more effectively using service level objectives (SLOs). Rather than relying on static thresholds, SLOs allow enterprises to evaluate “good” versus “bad” events through dynamic measurement of user experiences. With WatchTower, contextualized insights—including detailed trace information, correlated logs, and relevant error documentation—all become readily accessible within mainstream observability tools. This not only facilitates faster root cause analysis by SREs but also enables mainframe SMEs to prioritize their efforts on platform innovation rather than acting solely as data interpreters. Ultimately, Broadcom’s WatchTower creates a modern, integrated approach to mainframe observability, democratizing access to critical insights across the entire enterprise.
Recorded live on August 19, 2025 at the SHARE conference in Cleveland, Ohio as part of Tech Field Day Extra. Watch the entire presentation at https://techfieldday.com/appearance/broadcom-presents-at-tech-field-day-extra-at-share-cleveland-2025/ or visit https://www.broadcom.com/solutions/mainframe/observability or https://TechFieldDay.com/sharecle2025/ for more information.
Transcript
Hi guys. My name is Angelika Heinrich. Uh, please call me Angie.
That's what I go by. I'm the product manager for the Watchtower realtime streaming capability. And for those of you watching us on the livestream, if you happen to miss some of the previous talks, uh, Michael Kehl covered Watchtower and all of its available capabilities.
So please check out that, uh, recording later if you can. When we talk about the Watchtower realtime streaming capability, I like to always start with the why. What is the need that customers are trying to address when they install and activate, um, Watchtower Tower?
And really it all comes down to business goals, right? Our customers business organizations are looking for ways to grow, uh, increase profits. And maybe one of the business goals that they have as part of the strategy is how can we increase customer satisfaction?
Because if we have a happy, happy customer base, we can increase their customer base and we can increase our profits over time. From a business perspective, we are more involved, reliant on software and applications because it's a digital world. So the performance of applications and software are more and more critical to our business organization.
So that means the CTOs at our customer sites are saying, how can I, from a technology organization understand how, what the customer experience is like when they're engaging with our digital services? So let's think about this as an example. If you're making a payment, maybe you're trying to send your friend some, some cash from the dinner last night, or you've ordered something online, we have multiple options available to us, right?
We have maybe a mobile banking app, maybe we have an online banking portal. Maybe we have multiple bank accounts. So there's different ways we could do that.
We have, uh, peer to peer payment services available to us. How do we choose which of those options we're gonna use? Now, maybe convenience is a factor, right?
What's the right use case? But maybe we're also thinking about the experience we've had with a specific payment services. Like for example, when I press pay, am I expecting pay done?
Am I expecting pay wait and wait? Yeah. Right?
And then maybe it works. And what's even worse is you press pay and you wait, and then you get an error. So as a, as a user myself, maybe I try again.
'cause I think maybe I didn't click the button, right? Like, let me, let me press extra hard this time. Maybe it'll go through and then the experience is the same.
It fails. I'm probably gonna switch over to a different payment service, right? I'm not gonna keep trying.
I'm not gonna call and raise a, a ticket with the bank and be like, Hey, this didn't work because why I have, I have a couple other options already there, and the next time I need to make the same kind of payment, I'm probably not gonna go back to that service because I'm gonna remember that experience. So how can we make sure that we're able to understand the experience our customers are having, our end users to our business services and react. If I'm losing one customer or thousands, hundreds of thousands of customers within seconds, how can we reduce that impact?
Is site reliability engineering a term you're hearing from customers, from, from organizations you work with? Yes. I've, I'll just jump in real quick.
Please go ahead. Um, platform engineering platform, engineering's outcome focus is enabling, uh, the develop The developer. Mm-hmm.
Um, my sense of SRE when that came in, it was, it was, I mean, each seems to be a glorified version of the previous Yeah. Title. Um, and SRE was a glorified version of the, um, you know, uh, infrastructure operations team, right?
Whose focus was as the term implied, um, reliability and performance of the infrastructure. So you have a different design point, so to speak, for each function. Yeah.
And, um, I, you know, mainframe in general, in my sense has been very application oriented from the beginning. So, so, you know, I think it even has, has, has a really good point. Mm-hmm.
That's the trend where it's going. But if you're serving an SRE function, you know, maybe you are serving the developer within this particular corner of, of, of tech. I don't know.
That's a good point. That's a good point. From the SRE side at Broadcom, when we talk about the SRE, we're talking about the role of understanding what the end user experiences they're working with tools, different kinds of tools to monitor that experience, to triage.
When they see anomalies with the experiences that we should be having, they wanna be able to alert stakeholders, they wanna be able to analyze and ask questions on the data live right? As issues are occurring, how can I more quickly figure out what the root cause is? And when we've resolved that problem, is there a way we could automate, uh, that resolution next time?
So this is a little bit to what Steve was saying, is there is a little bit of platform engineering that is happening in the background as well. Now, all of these aspects within the site, reliability engineering pillar requires data. You can't do any of this if you don't get data from the different platforms and applications in your environment.
And that's where the watchtower realtime, uh, streaming capability comes in. And we're gonna talk a little bit more about that. So I said this SRE is using tools.
Can anyone think of different tools that SREs may be using at some of your organizations or your customer sites? Do we know what they would use? Maybe something like Datadog or Dynatrace is a tool that they're using to understand how applications are performing, what the end user experience, new Relic, NewRelic, new Relic, Grafana, multiple different kinds of tools are available.
And what's driving the customers to us is they say, we love New Relic, or we love Datadog, but we don't have support in that product for mainframe. So I can map out my entire business service until it calls the mainframe, and then it's an IP call and maybe see the latency, right? How long it took to get a response from that mainframe service.
But I'm not seeing what the mainframe service is actually doing. I don't understand what applications are being triggered. If something fails, where does it fail?
I rely very heavily then on my SMEs that, uh, some of my colleagues talked about earlier to gimme that information. And we know that mainframe SMEs, our CIS pros, our platform experts, they have a lot to do as it is. And having to always be this mediator to give information about how the mainframe is performing is maybe not the time well spent, right?
That skill should be working on maintaining, upgrading, and modernizing our platform. So how can we improve the life of an SME on mainframe, but also make it easier for someone on the SRE side to understand how the mainframe plays a role in an end user experience? So with Watchtower, we are streaming live mainframe, uh, telemetry, telemetry being traces, metrics and logs in an open telemetry format, which is an observability framework and standard, and is supported by the various tools that SRE employ today.
This is a great way to use a standard like open telemetry to educate an SRE on mainframe. They don't have to be a mainframe expert to understand what they're looking at when we're using a standard like open telemetry. What we're talking about here is, is, um, exposing in Datadog, Grafana, whatever, through open telemetry a, uh, z based, a mainframe based application as if it were just in Just any One of their containers somewhere.
Exactly. Or pods. Yeah, Exactly.
I didn't wanna use the term, but it is kind of democratizing the mainframe information in a way that is easy for a non mainframe expert to understand. Now, when we talk about open telemetry traces, metrics, and logs, I'll just go back to the screen. A trace on its own is gonna tell me how a request froze from an end user application through the system to a backend mainframe kicks region database access, all of that.
Now, a trace on its own gives me this request level insight, but what happens if I start to see the kicks request latency or the, the response times increasing? But when I'm looking in the trace, I'm not seeing metadata that's saying it's something specific to this kicks transaction. For example, how do I understand if there's an environmental impact, maybe the resources kicks is using, is impacting the way this kicks transaction is performing.
We would use things like metrics maybe to understand the environment, but we don't wanna just have a metric here and a trace here. We wanna be able to go from a trace of a specific instance of a kicks, uh, transaction and immediately within context have metrics around that trace or that specific, uh, uh, instance of the transaction. We may also want to have logs, right?
Maybe there is a specific event happening as well with some descriptive information of an event that's happening on the mainframe that could be impacting this kicks transaction. And I wanna talk a bit about that because each of those signals I've just described are powerful on their own, but they're more powerful together. And if you use open telemetry the right way, you can easily correlate those, that telemetry.
So you're giving someone kind of, uh, a three dimensional view of how that kicks transaction, for example, is performing. What's around that kicks transaction that could be impacting my, uh, my end user experience today. And that's something that we focused a lot on as we were developing, uh, the watchtower realtime streaming capability.
And, um, I'm going to just take you through a very short, uh, play through of what a, you know, an SRE may be doing to understand what is happening on the mainframe. Does anyone know what I'm looking at here as an SRE? Has anyone, does anyone see any indicator of what this may be?
I'm thinking the big bad red, the big, I'm thinking there's something bad. I don't know why my eye's drawn to the red, this big red thing. Like with the key that says bad events, maybe Something bad is happening.
That is so true. Oh my gosh. And I love that they just used really simplified language, good and bad.
I mean, let's not get technical about It. Well, I'm seeing an infrastructure, not a budget doesn't look good. Yeah.
I'm seeing a capacity and infrastructure stress because we got a whole, we have more good events mm-hmm. At that moment than anywhere else. Um, but there's clearly a capacity constraint.
There is some kind of capacity constraint. I wonder what it is. I wonder what it could be.
Has anyone heard of the term SLO service level objectives? Mm-hmm. Yeah.
Yeah. Do we know why we would use something like service level objectives, maybe in place of a static threshold alert? Does anyone know why we would do that?
A good question. So an SLO, um, SLA is the usual term, but it's, it's morphed, Right? So we, we might agree on an SLO, right?
We have an agreement, this is how services should perform. And the SLOs actually maybe the, you know, what do we want the experience to be for whatever it is, an end-to-end service, a specific KS region? Yeah.
I feel like SLOs acknowledge trade-offs mm-hmm. When failure is allowed. Yes.
Uh, because of the assessed cost of, of meeting, you know, something stricter. Yes, that is exactly it. We have, you see there on the, on the, on the side there, we have this error budget.
What is our trade off between, uh, reliability and availability at this point in time? A static threshold is great when we know what we want to measure. There is something specific that could go wrong and we know what it is.
But when we are talking about complex cross-platform services, we don't always know what might go wrong when it does go wrong, right? That means with a service level objective like this, we, we want to use two things. We wanna be able to measure good events.
We know what a good event is because we have a service level objective. We know what good performance or good experience should be. And we also know the total number of events or transactions, right?
So we can measure good transactions and total transactions. But the bad is we don't know, but we, we know it's not good. So this is where the service level objective is quite flexible.
The downside could be the first time you start to use service level objectives, it can be more complicated. So there is a bit of a lead time to understand how to properly employ this, but all of the data that we're using is coming from traces and metrics and logs, traces. There's ways to set an error to say a transaction has failed, it has not performed well.
Metrics also give us the ability to understand trends that are, uh, good or bad. And so here's just a snapshot of how an SRE could use data coming from, uh, the Broadcom, uh, watchtower solution that we, we commonly refer to as the iris. These are metrics that we are producing out of a Kix region, for example.
And here we can set what a good event is. We can set the total number of events that I wanna measure. I wanna have my five nine availability.
And when we have a breach, I wanna be able to notify someone about what that breach is. SLI, So are the service level indicators. So an indicator would actually be these metrics here.
Yeah, I know and No, but I is used for indicator and I never remember that it is. Okay. So now we're, we're feeding data from mainframe to a tool.
And SRE knows they know how to use these SLIs to mes measure SLOs according to SLAs. I love saying that. So again, the SRE doesn't need to be a mainframe expert to be able to use this data in a way that's gonna be useful to them.
They understand this platform well enough and they understand open telemetry, so they understand how to, um, set this up. But I was talking about correlation. So an SRE wants to go from that big red bar that we saw.
How can we understand more about what's actually causing this big red bar? Depending on what tool you have, you might have a really fancy little link that'll tell you, oh, let me see the related traces. And here we can see just like a snapshot of some traces coming in from this kicks region, um, that we are sending from Watchtower.
This is the kicks region name. It's a task owning region. This resource is actually the transaction, the SSPN.
You can see we have the duration of each execution of that, uh, transaction. And we can see that time is increasing, right? The, the amount of time it takes this transaction to complete seems to be increasing drastically.
At some points. You have four minutes. Can you imagine waiting four minutes for a payment?
So we're seeing, you know, just from this, uh, summarized view of these trace, this trace data, we're seeing some bad behavior. But again, we wanna give context, we wanna give some correlation. We wanna help find the root cause, at least an area that an SME on the mainframe should be focusing on to resolve the situation.
So this is just a snapshot of what a trace could look like. Again, a trace is distributed, it's not just mainframe photos includes distributed services as well. So we can see there is this queue a message probably went on that queue, it triggered that kicks transaction.
And we can see some MQ actions that, that, uh, kicks transaction is performing. Now in the trace data, if you've ever seen an open telemetry trace, there are multiple ways to describe the environment, the transaction, um, the, the system itself as well. But there's also ways to add very detailed information.
And so what we did with, uh, the realtime streaming capabilities, we said we can't just say there's an error because that might be helpful. But that means someone's gonna have to go maybe into sys log or into the kicks log to see what the error is. So let's extract that information.
Let's give them their event code and then let's also give them full documentation of what that means because maybe the SRE is gonna raise a ticket. Maybe the first person who looks at this is not familiar with what a SP seven is, instead of having to go Google it, it's right here in line with the end to end user experience view. And we can see more information about what exactly may be failing and impacting that kick transaction.
So what I want to say is it's not enough to only talk about tracing or we're creating traces. Now. We need to add blogs and we need to add metrics, and we need to do this in a contextualized way to help SREs better understand how the mainframe is performing, but also help the SRE helping in a mainframe.
SME. All this information here is useful to a mainframe SME to solve the problem. The region name, the time, the dates, the error codes, all of that information is useful.
So they will go into watchtower and know exactly where to start.