Expanding Value with Broadcom WatchTower’s Metric Analysis
At Tech Field Day Extra at SHARE Cleveland 2025, Broadcom presented their advancements in mainframe observability through the WatchTower platform, an evolution designed to improve how performance and optimization data is collected, correlated, and acted upon. The core problem addressed is that traditional domain-specific monitoring tools, like SysView for CICS or NetMaster for networks, operate in functional silos, leading to inefficiencies when diagnosing cross-domain performance issues. WatchTower eliminates this obstacle by aggregating and contextualizing alert data across different mainframe subsystems. As a result, users gain broader visibility through Alert and Contextual Insights—a holistic view that improves incident resolution speed and accuracy while making it easier to distinguish between isolated and system-wide failures.
One of the hallmark features of WatchTower is its intelligent dashboards, which provide both real-time and historical views across previously siloed data domains. These dashboards tap directly into established data collectors from key Broadcom tools such as Vantage, SysView, and NetMaster, ensuring that users can visualize telemetry data from CICS, DB2, MQ, IMS, and network components in one cohesive interface. The dashboards emphasize usability, allowing users to create custom views without writing complex queries. This flexibility accommodates both infrastructure-centric roles, like a CICS administrator, and higher-level application-focused personas. Complementing this visual capability is embedded machine learning (ML), which not only detects anomalies but tracks gradual performance degradations—informing users of potential issues before service level agreements (SLAs) are breached.
Broadcom’s ML Insights extends the usefulness of WatchTower by delivering proactive anomaly detection built on each organization’s unique dataset, rather than relying on static or generic thresholds. This approach tailors models dynamically based on ongoing operations, learning from historical patterns and adjusting its behavior to recognize events like seasonal peaks or daily load cycles. While immediate automation is not triggered by the platform directly, Broadcom allows integration with OpsMVS for customers who wish to design their own event-based playbooks or remediation rules. Although historical data cannot be retroactively loaded into the system at this time, WatchTower begins modeling from initial deployment onward, allowing it to curate dynamic baselines over time. These ML models, however, currently lack a formal feedback loop, though this capability is under consideration for future updates to further enhance adaptive learning.
Presented by Tom Quinn, AIOps Product Manager, and Machhindra Nale, AIOps Architect, Broadcom. Recorded live on August 19, 2025 at the SHARE conference in Cleveland, Ohio as part of Tech Field Day Extra. Watch the entire presentation at https://techfieldday.com/appearance/broadcom-presents-at-tech-field-day-extra-at-share-cleveland-2025/ or visit https://www.broadcom.com/solutions/mainframe/observability or https://TechFieldDay.com/sharecle2025/ for more information.
Transcript
Hello everybody. I'm Tom Quinn with Broadcom Mainframe Software. So I'm the product management lead over our automation, performance, and optimization products.
Uh, so that's a lot of the core products that Mike had talked about in a another session. So that's our SIS views, our net master, our ops, MVS mainframe application tuner mix. So these are all the core products that our, our customers are using every day.
Joined by my colleague today. Hello everyone. Uh, this is hin.
You can call me Mac. I'm the solution architect working on the Washout platform. So Mac and I are gonna talk to you today about how we are expanding the value of our core products, right?
So what the situations that we have today is we have those core products are used by our domain experts. So an example of sis view, for example, it will use or sis view monitors your CICS, your IMS, your mq. You might have multiple people that use sis view, use that specific piece that they're interested in.
They know you've got your network analysts, they're using Net Master and so on. So when a problem does occur and they get that problem inside of their product or through an alerting system, they're looking at their domain tool and they're only seeing their domain. And so we recognize this as, as kind of a problem, right?
To some degree. So what you end up with is that war room where everybody's looking at their own tool, everybody's pointing at each other. And one of the things I like to point out, uh, I've got a long history of performance management myself.
And one of the things that I thought was most important of not only finding where the problem is, it's finding where the problem's not right. So I like to be able to say, not my fault. That's just me.
So with, with Watchtower, and so with the observability platform that we've developed, we've integrated all of our products into the Watchtower capabilities that we offer. One of those capabilities, we'll start with here is what we call alert insights. So when you see a problem, say in CICS or you see that in mq, you can generate an alert, but that alert is not just in that product anymore.
You've now escalated that up into alert Insights, which by itself may not sound all that interesting, right? It's like, okay, great, an alerting capability. Where it becomes interesting is Alert Insights will now look back to all those, those core domain products and pull information from the time of that alert and correlate that information together.
So now you're not just seeing that CICS alert that maybe is interesting by itself, but what was happening in MQ at that time? What was happening in DB two? What was happening on the network?
Like, is this a bigger problem or is it just this particular piece that I have? So we call that contextual insight. So to the alert, that of itself, you're seeing the context of that alert.
You're bringing in all the information from all these different domain products. Now, once you got to that point, you see the alert, you see what's happening, you see the context of what's happening. We can take that then to a next step, pulling in data from additional resources to get to dashboards.
All right, so let me start with some challenges. What mainframe operation team today faces? And then we go about what we provide with the wash tower dashboards.
So the mainframe team operation teams today faces the challenge with respect to accessing, analyzing and visualizing the data across the silo domains. So we have tools already, which actually does this monitoring for DB two tools for CI ICS monitoring, network monitoring, storage monitoring. But all these deep tools actually provides one domain view.
But imagine picture this, that if you are facing a challenge with the CI ICS, there may be a problem with the network. There may be a problem with the DB two. So we would like to actually bring all this data, siloed data together in one central place, right?
So, uh, we want to actually see this data in near real time. And, uh, often the mainframe operation team want to see the historical data as well for trained analysis. That's where watch tower dashboard comes into picture.
It taps on the existing Broadcoms monitor, such as SIS view, Nate Master Vantage, uh, sis View four DB two. These monitors for last decade actually fine tune the collectors. So we collect a lot of KPIs around CICS, tb, two, mq, IMS, uh, network, T-C-P-I-P stack, uh, volume groups.
All these collected telemetry data is being streamed to the watch tower platform. And that's where the dashboard comes into picture to visualize. Now, we still catalog this particular, uh, telemetry data into our efficient data store and make it available to variety of personas, which includes, um, the system operators as well as the subject matter experts who are trying to solve an incident on the mainframe With the watched our dashboard.
You're not starting fresh. Uh, we have baked in some templates, which you get, you started immediately. And that's where our, uh, experts actually prepare those templates so that you get started, right?
Uh, when you install the product, uh, the dashboard, uh, correlates all the data from across the sis, uh, CICS DB to MQ network. You can imagine that you can build one dashboard which can cut across multiple domains and look from the application perspective, uh, the dashboard has, um, a easy intu way of building the custom, uh, dashboard if you want to. There is no query layer or nothing that you'll be writing complex queries such as prompt ql, but it is easy to pick the metric, pick the resource, and build the dashboard, share with your team, and so that at the time of incident, all of you will be looking at this particular dashboard collaboratively and pinpointing where the problem could be.
How do your customers typically set their BA dashboards up? Is it around a workload? Is it around, you know, do they have like several dashboards that are for different workloads that they need to keep their eyes on because the, their, um, SLAs are so high for those?
Or like how do they normally do it? Or is it maybe set according to like the user and their priorities? Absolutely.
So the question was about it is by persona, or is it by the application perspective? Both ways. So if I'm A-C-I-C-S administrator, I would be responsible for upkeeping the CICS system.
In that case, my view would be CICS dashboard only. Whereas if I'm looking from the application perspective, I'm not looking for the infrastructure, perhaps I'm looking for the golden KPIs from the application perspective. In that case, whatever makes that application, which includes, uh, subsystems, I will put them together.
So both ways you can create the custom dashboards. How do you see your customers doing it? Uh, currently the customers are focused more on the infrastructure side.
So if you look at the screen, uh, the CPU BZ percentage, this particular chart, what we call it, is a ML highway chart. There's a machine learning highway chart. This exactly does what you're talking about, that instead of me watching this particular metrics, the machine is watching this metrics and finding out if there are any anomalies in there.
And that's why we call these are intelligent dashboard, because we embed the machine learning in the visualization, we bring the alert exceptions within the visualization. So you are looking at alerts, you are looking across the domain data, you're looking for embedded machine learning within the dashboard. Uh, and so again, this is looking at the, the data that we're, we're seeing, the dashboards that we're building out.
This is extending the value of those core products that we have. So again, when you're looking at your SIS views, your net masters vantage, you are looking at your domain specific data. We're utilizing that same data for multiple purposes.
So it's not that you're installing, say, another product called Watchtower or Watchtower dashboards. You are installing this platform that's utilizing all the core data collection that's already occurring, which kind of segues into our machine learning piece. So with SIS view, with net master vantage sis view for DB two, we're collecting that that data, that's what our domain experts are using.
We're also making it available external to those particular products into what we call our ML insights. So our watchtower ML insights is our machine learning component utilizing that core data collection. It's the same data.
We're not installing another data collector to do all of this, right? We're utilizing that data that you're already getting, you're using on a, on a regular basis. And this is, uh, with my performance background, I, I will go into, I can go into de details with, you know, static thresholds and dynamic thresholds, right?
This is what we would call that really that dynamic thresholding what is happening as it relates to, you know, time. Like where is a Monday morning is different than a Wednesday afternoon, and this, this is where this anomaly detection comes into play. And we will look at your historical data and then the real time data is coming in, it's scored against that to determine is there something different that's happening that's, you know, not normally happening right now, uh, as opposed to a static threshold.
Is the number's five? Is it five or not? Is it above?
Is it below? So the machine learning piece is, is really taking that, all that knowledge and, and apply and Mac you wanna talk a little bit more about the machine learning? Yeah.
Where do the, where do the rules come from? Are these supplied by you or by the, the, the end user, Right? So I'll answer that.
Um, so you can imagine that the collectors that we have in the, um, currently with the CS view net master vantage, all monitors collect huge amount of metric data and then watching them and then putting the thresholds. It's hard job, isn't it? Your workload changes, your system changes, and hence both the Monday mornings are different than the Friday afternoons.
And so, so that means you'll be keep actually playing around with those kind of threshold rules. And that may be, it's not an efficient solution. So perhaps we need to shift from reactive to proactive approach.
And that's what this machine learning, uh, does for us. So instead of you watching hundreds and thousands of like metrics across subsystems, the system will watch them and then, then we employ the machine learning algorithm such as kernel density estimator, which is the customized algorithm to keep looking at the metric history, then give the probability or the predictions how this metric should be behave in the future, and then overlay with the real magic on the top and see if there is an abnormal behavior. So I think to answer Steve's question, you install this, yes, it learns for a few weeks, right?
And then it sets some thresholds. Is is that what you said? Or, um, is there something Well, it Sounds like what he said was that these like more dynamic than traditional thresholds, which is what I love about it.
So it'll learn based, yeah. Yes. So it'll learn based on your system.
Yes. Right? It learns, it learns from your system.
So as it, it, the more data it has, the more intelligent it gets, right? So as it experiences that first, you know, in the United States we have the, the Black Friday sales, you know, resources probably spike that day, right? So the next year it's gonna see that, it's gonna recognize that, hey, we had a spike on this particular day, right?
Or even on the, the Monday mornings when everybody's logging in from a long weekend into their, their brokerage accounts to see what, what they're gonna do this week, right? It, it's, that's maybe there's a, a peak there, but on a Wednesday afternoon nobody does that. So it's low.
So you're building out these bands of normal activity and as a, a new or realtime data comes in, it's looked at to that band of normal activity, right? And then it will see where does it land? Is it within that that band or is it way outside?
And the farther outside of that band of normal is the more severe the alert. But what's nice about this too is you also start to see slow creeps, right? So we'll hide those from you a little bit at first, but then once that slow creep continues, we'll let you know about it.
And so this is where there's that good tie between your static thresholds and your dynamic thresholds. So a lot of people think of static thresholds as being like, these are my SLAs, right? These are my agreed upon performance pro profiles for this kicks transaction, for example.
But I could be run, let's say it's five, I'm just gonna make up numbers here. It could be five. We're always running it.
Two, we go up to three, I still don't know about it. Machine learning will know about it. It'll start to see that slow creep that goes from two to three to four.
It's gonna tell you about it because you're outside of normal, your normal's way below your SLA, it's not a problem. But if you could bought yourself two hours on the clock to see that slow creep before it gets to be a problem, that's great. Right?
And so that's one of the other side advantages of machine learning is you're buying time on the clock, you're getting to know about a problem, a potential problem long before it ever is. So it only does reactive, it doesn't do, it doesn't make any decisions. All it does is inform you and say, Hey, there's a slow creep going on.
So you could build in automation. So there is opportunity to build in automation when it, when it detects an anomaly far enough that it would want to create an alert. That alert can have automation tied to it.
So if you did want to, like if it's a common problem, um, can't think of it off the top of my head, you know, add storage or something, or cancel a transaction. Like you could add automation to those tasks or to those alerts if you wanted to. Uh, I feel like, I feel like I at least am swimming in the deep end of the pool here because, so I, you're describing a reinforcement loop of that, right?
But you're not explaining what the, I mean, is there, is there a human, um, are you learning from human interaction? So, so I assume this model comes in with a baseline. Yes.
Right? So that's the first thing. Where does the baseline come from?
You have an opportunity like most vendors do to consolidate typical behaviors and have a human say, just use typical behaviors for this industry or for that industry or for this application, or that that can establish a baseline. Mm-hmm. But then you need a reinforcement loop in order to determine for the particular application scenario what is normal and what isn't.
And I love that it's learning and I love that it's using ML to do so you haven't described a human intervention there. So in other words, like Steven, what we're imagining is the operator's sitting there and gets an alert or Jeffrey and says, he says, oh, this is outta band. The alert says this is outta band, and the operator says, no, it's not.
This is normal. And after three months of that, or well, well almost immediately they may not get that alert anymore. That's the behavior just you're describing.
Is that, is that right? Yeah. I feel like I'm filling in blanks here.
I'd want you to fill in. Sure. So I'll, I'll help with that.
So the two questions that you asked about, where's the baseline coming from and is there any feedback loop involved in the machine learning? So the baseline, when we ship the product to the customer, the baseline gets automatically built for that customer. The models will be built for that customer.
There is no model that we ship to, uh, customers. Uh, as the data comes in, we look at the history and the model gets built for that particular metric and resource on the fly. And that's the model that will keep changing as the metric behavior changes.
Oh yeah. Okay. Yeah, of course.
Alright. So that's, that's the baseline. That's number one, number one.
And number two is the feedback that we don't have yet. It's on the possible roadmap that we are looking for the feedback from the customer. So today, when the ML alert comes in, it may be possible that the, the operator look at it and says, maybe it is not an issue for that today we ask operator to go and change the monitoring rule, which involves the machine learning and adjust that, which we could do automatically.
But that is human. I I would argue that is human feedback. Yes.
I, I described a a method of human feedback. But you're saying like the alert comes in and then the operator's action is actually changing a rule that changes the, the alert, the learning process, right? That's right.
So that's all right. Good. Can I ask a slightly more prosaic question?
You talked about the, the, the learning exercise that the tool will go through to give you the necessary historical data then, you know, to, to provide the heat map of like, well this is what normal looks like. Right? But you've also mentioned like there are, there are gonna be particular moments in the day or the week or even the year.
It is your suggestion or is it, is your sort of operation to go back at least a year to provide that? Because, you know, I can think of black Friday, Christmas, end of year reporting, all of that kind of fun stuff creates anomalous performance metrics that sure. You need to learn from, right?
Is that, is that basically how you're doing it going that far back? Yeah, so currently what we do is that we learn the cyclic behavior. That's an important part.
So the seasonality cycling behavior is already taken care by the algorithm. So with the aggregated data, you should be able to correlate from the past event to the current event That's possible. Yeah.
Okay. So the customers like they, I mean they, they have that data anyway, right? So it's no big deal just to go back however far back in time.
But it's That's right. You're going back quite a long way though to create that model. Right?
Excellent questions. Thank you. So This is like a diagnostic machine learning application, Diagnostic machine learning.
Can you expand a little bit on what you mean by that? So it just take basically everything that you've been saying, it can go and it can look at all of the different sources, bring those all together Sure. And then say, oh, here's where the problem was diagnosed, where the problem came.
Yeah. Yes. Right?
Yes. Or is it just alerting you to the problem? A little bit of both.
Okay. Right. So the, the machine learning component is going to alert you to a potential problem, right?
So there is a change in normal, right? So that, that's your alerting piece of it. Something is different than normal from a diagnostics perspective.
With building out the dashboards, if you put that what we call the ML highway, so the machine learning, the historical piece of that, that view of the data that can help with the diagnostics. So the diagnostics still has to come from the human, right? So it's have multiple sources.
So it's basically just giving you the, the here's here's what we see. He is reporting on it. Yeah.
Right. To visualize the issue. Yes.
Yep. Alright, great questions. Thank you.
That tooling for their workflows to see the data from multiple sources and find that root cause of the problem, uh, more quickly do some collaboration and get to the root of the problem. So Quick question, do how well do you play with all the tools? Because a lot, a lot of mainframe sites might not have the comp complete suite.
I might have a bit of a mixed bag from some of the other vendors. So right. Do, can we ingest data from other vendors?
You have an API where I could push stuff to you. And So long term, through the usage of open telemetry, that may become more possible, right? So if you can get the data out in an open telemetry format, your different signals, that may be something that we would pull in.
Can you comment a little bit on the, um, the data integration here from disparate sources? 'cause that's, uh, um, the, the, the 50% of machine learning that seems to be talked about 5% of the time is data quality and curation. So how are you, how are you accomplishing that?
Since, since the more and the more disparate sources you have, the better the uh, the better the, the alerting, Right? So the way that currently we work is that this collectors on the host on the mainframe, they are now collecting the data at the tune up, say one minute interval, and then you can change that interval if you will. And our collectors now talking in the same vocabulary so that the semantic convention of the resources remains the same.
And hence for the machine learning does not have to think through that. These are desperate sources, but it is one, uh, source like four db two and network. They are going to talk in the same semantic convention.
So that means if you are, uh, collecting data for Aze system, and another collector doesn't say that it is not aze system, but it is a host, both are saying it is aze system, so we should be able to correlate this in the machinery. It's a generative, uh, model so that it's taking this stream and predicting what's coming. Uh, no.
So we, we took care of the data collection and speaking in the same language. It's not a generative model. Yeah, yeah, I got That.
And, and once we collect this data, then we, we know that this is how, this is the plex, this is the zero system. This is what the network stack for the same plex, for the same zero system. This is the CI CS region.
So you see that there's a common vocabulary between these two that we are leveraging between the machine learning, but there is no LLM or model which, which brings the semantic solution on the top. Is there any plans for automation recommendations for next best action? I've done this thing nine times on the 10th time, automate it.
Can I build that level of automation based off an off the alerts? So this ML alerts, which actually has got integration with the Broadcoms apps M vs. Which is our event automation platform.
So that's that solution. Or with the integrations from the ML insights, our integration from si, Cisco, and other monitor products, you can actually write some sophisticated event rules to do some automation deduplication. Uh, yeah.
So you have to do that manually for sure. Um, when we talk with the customers, the reaction was that I don't want your observative platform to take the action to cancel my important CIS transaction. Let me human be in the loop, let me write the event and so that I know what I'm trying to do in order to take that action.
Typically they have a service management and an incident and all these policies in place in order to get something done. So we, we make sure that the contextual information is provided, event automation hook is provided, and then it is up to the customer to go write that rule for them. Okay?
So it could come back with reactions, do this, you could do this, you could do this, and then have somebody make a decision. Right? Okay.
Um, my next, my, I do have one more question and that is, uh, somebody that actually, let's say they're implementing this into the system. Uh, does it backlog, does it monitor or plan and monitor? So let's say that somebody's been running their, their mainframe for years now, and then all of a sudden we, they install this, this, uh, this watchtower software.
Will it be able to go back and look at the data and say, Hey, you know, on every year, on September 11th, this happens, um, every year, you know, these are the dips and these are the, uh, these are the, the peaks of things is, is, is does it collect, does it build up so it can create a model going forward? So currently we don't have the ability to go look at the historical data that you may be having in a different format and then preload in the model. But rather at this moment when you deploy the application and it starts stemming from the, and net master and other products, that's where it starts building model then onwards.
Okay? So it's a real time. And when you get installed, that's where it starts.