Greater Insights into Application Performance Management with Broadcom WatchTower
Broadcom’s presentation, delivered by Petr Klomfar at Tech Field Day Extra at SHARE Cleveland 2025, explored the emerging necessity for sophisticated Application Performance Management (APM) solutions to address modern IT challenges. Titled “Greater Insights into Application Performance Management with Broadcom WatchTower,” the session demonstrated how reduced visibility into mainframe applications, increasing resource costs, and complex modernization efforts are creating new demands for observability, optimization, and operational efficiency. Klomfar introduced Broadcom WatchTower and complementary tools like Application Performance for Z (AP4Z) and Mainframe Application Tuner (MAT), emphasizing their value in providing continuous, scalable, and low-overhead performance data to drive smarter and faster decision-making.
Klomfar began with a historical overview of computing trends, highlighting the shift from early, resource-constrained environments that necessitated deep optimization to an era where increased computing power led to more complacent development practices. However, with today’s economic pressures and pricing models like IBM’s tailor-fit pricing, there’s a renewed emphasis on efficiency. Mainframes now require high levels of visibility and analytical depth to optimize and modernize legacy environments. Klomfar identified challenges such as the “black box” nature of mainframe applications, difficulty in identifying ROI for performance improvements, and the complexity of root cause analysis. To address these, Broadcom’s AP4Z continuously monitors system performance with negligible (often 0.1%) overhead, supporting optimization, modernization planning, and troubleshooting with real-time, context-rich data.
A major focus of the presentation was the synergy between WatchTower and performance tools like AP4Z and MAT. AP4Z provides application profiling that can identify hot spots using Pareto principles, support modernization pathfinding by highlighting application components, and flag anomalies for root cause investigations. Meanwhile, MAT performs deeper, targeted analysis with slightly higher overhead (typically around 3–4%) and feeds into operational workflows via WatchTower alerts. This integrated architecture empowers organizations to decrease mean time to resolution (MTTR), improve SLAs, and make measurable reductions in costs through smarter resource usage. Klomfar concluded by championing the benefits of clear and accessible performance data for diverse roles, from system engineers to developers, affirming Broadcom’s commitment to continual improvement and client collaboration.
Recorded live on August 19, 2025 at the SHARE conference in Cleveland, Ohio as part of Tech Field Day Extra. Watch the entire presentation at https://techfieldday.com/appearance/broadcom-presents-at-tech-field-day-extra-at-share-cleveland-2025/ or visit https://www.broadcom.com/solutions/mainframe/observability or https://TechFieldDay.com/sharecle2025/ for more information.
Transcript
It's a pleasure to be here. My name is, uh, pet Kfar, uh, r and d engineering manager at Broadcom. And I have been around the application performance tooling, uh, for more than a decade now.
Today I'm gonna be talking about greater insights into the application performance management, and maybe allow me for starts have a short, hopefully short, um, philosophical window. So I think we will all agree that, uh, application performance management plays a substantial role these days with any healthy and competitive organization. Why do I say healthy and competitive?
Well, think of it from the perspective of the historical, um, it technology advancement right? Back in the days, um, where hardware was basically fresh in its NetApp piece and it was still evolving. Um, we had lack of resources where respectively the resources were precious.
Uh, you were forced by nature to actually have all your application programs to be, uh, you know, optimize down to the level of the instruction level because there was not instruction to spare. At that time. We did really good job, but then we hit a milestone where there was a light speed development in the technology, which basically made the computing power available to the masses and much more accessible at that point.
I dare to say that we have let our guard down because, uh, many people have chosen to beat the problems, simply beat the well with the amount of a cash and, uh, basically more power, uh, more CPU, more memory. You just paid it and everything was okay. Uh, it's not too far in a distance pass that, uh, we have been forced back into the trenches.
If you think about it from the perspective of how market has crystallized, how the, uh, competition, uh, needs to, you know, grasp for the market share, um, it reintroduces the hunger for optimization, and especially if you take a look on the mainframe in regards. So, for example, a uh, interaction of tailored fit pricing. This is, this is where, uh, every second and every application, um, you know, uh, that is running makes sense and counts, let's say.
So, um, every improvement and every innovation, uh, first we need to ask ourselves the basic premise question, right? Uh, we need to set up a baseline and understand the customer pain points and challenges. Uh, what we have been hearing from our clients in this regards is that mainframe is a black box when it comes to end-to-end visibility of, uh, the applications, mainframe applications, and specifically when you talk about the individual performance data.
Um, so, um, mainframe applications can be really high, complex individual pieces of work, and any attempt to modernize usually poses a high risk. And, um, it is typically a slow approach and costly, uh, uh, costly thing to do. Um, next one I have is the optimization costing the ROI.
So it is diff typically difficult to, uh, optimize the performance and resource utilization, especially hard to figure out what is your scope, especially hard to figure out, uh, what is your entry point, what are the applications mostly worth to optimize in general. Um, we also are talking here about the return of investment. I'm coming from the optimization space and typically the return of investment for all the applications that I worked with is pretty fast, but you have to be smart and reasonable with it.
So let me have a mini ezzo here. So, um, no application optimization measurement or optimization in general comes for free. It's always an act of balance.
I'm having in my mind, always the par rule, uh, the 80 20 division where you are basically saying that the 20% of your applications are typically responsible for 80% of your performance issues and 20% of your tuning efforts can result in 80% of the performance gains. Um, it's quite important for the later of my presentation as well, uh, when we are talking about the last, uh, challenge that we are seeing here, it's the, you know, the root cause analysis of poorly understood applications can be very difficult, and the subsequent issues are definitely impacting your resiliency and the cost of your operations. Now, um, how to get out of it.
What is the, what is the, you know, way how to address these challenges? So before I answer the question, what to do about it, I think it's worth actually taking a step back and looking at, uh, how we do things today. So optimization is a cultural multi-skill environment and where every organization is striving to be excellent in this space, you need to take, take a step back and look at your tech stack holistically, right?
So we typically talk here about the top down approach from this perspective. Um, in each of these stages you have a, uh, different roles, d disparate responsibilities, different different applications, uh, that on its own brings a sets of its own challenges. But let's take a look on the stack, right?
We starting at the top from the satellite point of view. So overseeing the machine utilization and capacity constraints is a must place to start, and it's a, a great way to actually get yourself going, right? But then you are coming to a pilot's view and the control tower.
So the pilot's view in, in its own is, uh, from perspective of looking systemwide your applications, understanding what is running and the control, um, tower point of view, this is where I see a basically a gap and it allows, um, it bridges the gap with the overlap in between of the ground level and the in-depth analysis versus the, uh, application optimization at scale. And this is actually where the application profiling and one of the products that I'm gonna be talking about shortly plays a significant role and possesses a unique place on the market. Uh, so application profile, uh, profiling, uh, and AP four Z application profile for Z in this instance is, um, a ap, uh, is a product which is built from ground up to be very lightweight, and it is, um, build up to monitor the application at scale 24 7 continuously and with very little impact to customer or the user resources.
When I say that the overhead typical of the application profiling in the sense is, uh, within about less than 1% of the workload as a overhead, but the measurements with the clients actually showed us it's, uh, around the ballpark of one 10th of a percent, which is a fantastic result for a product like this that is doing things at scale and, uh, uh, generally in the near real time. So Where are those numbers coming from? We are, we are coming, uh, we are having a reports from our clients in terms of how much MIPS they are using and versus what is being attributed to the P four Z.
1%. Thank you. Um, so I will still stay with the AP four Z for a second.
So AP four Z in general address three main use cases or three key use cases here. Um, they're related to the challenges that I mentioned in the be beginning, which is the, uh, optimization, troubleshooting and modernization. So from the optimization perspective, and I remember when we talked about the par rule, it helps you to, uh, really identify the scope of your operations, identify which applications are your hotspots, and, uh, to which you really, you know, extend your strength.
It allows you to, um, monitor the, uh, before and after optimization, uh, behavior of those applications. And, uh, it basically provides it in very easy, understandable way. Maldi, uh, deed, let's say the modernization, one of the challenges with the modernization was the complexity and the risks related to, you know, not understanding your application.
Again, the information that application profile can provide you is to understand your application more and have a better plan of attack for a modernization effort, essentially, um, um, boosting up the return of investment and lowering the risks. And, uh, one last thing that I will mention here is the troubleshooting for, so from the troubleshooting perspective, you know, application profiler works at the level of the modules and it can really, uh, assess the CPU usage and other metrics at the module level. And across the scale, it can, uh, really help you to detect potential security integrity issues or exposures because, um, it can see unexpected runs or, um, changes in your, uh, versions of the products in general.
And, uh, it also can provide you a good focal point for your, you know, application deployment for your QA people who, who will basically have a chance to see that this portion of the application is really the one that has been touched by the introduced change. I will still mention one thing that I really laugh about AP four Z and that's the, we call it first data, uh, first failure data point capture. And, um, as in contrast to what you have with the, uh, taking a dump, for example, when the transaction fails or events, a P four Z can actually record what LEDs to that point of view, and it can play it back.
So you have a trace from how the application or the transaction was behaving before it failed, which is actually a really good information now, uh, coming to the ground level point of view. So from this perspective, this is where the in-depth analysis are, uh, living. Uh, they are typically very powerful tools here.
Uh, I'm typically talking about mainframe application tuner. That's the one that my colleagues has mentioned in the previous segments as well. And the mainframe application tuner is a, uh, workload analysis tools, which is a load density and it is non-intrusive.
It doesn't alter the code in any way. Uh, it can monitor pretty much anything cable workload on a mainframe. So all of the IBM subsystems for glabas natural, IDMS, uh, unique services, distributed portions in terms of the Java applications, measurement and so on.
And the beauty here is it stay, it can, um, analyze the issues up to the statement level of the code. So it can really attribute the delays, uh, activities loop, uh, loopholes, bottlenecks, inefficiencies, hotspots up to the statement individual instructions. So that is it.
Yeah, this is, this is where the synergy of these products is really coming forth because where AP four Z as an application profile can be very proactive solution. The math, uh, is like more reactive and ad hoc, you know, dispatch me here and give me the information and I'm slowly but surely leading. What does it mean with the observability platform?
So still bear with me. So you say, okay, that's great, Peter, because you have lots of data for that are otherwise hard to get from a different sources, and it is, uh, you know, somehow connected together. But is this really the way that is basically being the wow factor?
Um, this is where Watch Tower, the incremental observability, uh, from Broadcom, the watch tower platform, uh, bears its fruits because it is a fantastic platform for, um, optimization products like these to integrate with. And finding the silver lining of the contextual information, uh, in comparison with also other watch tower players. You know, my colleagues has been representing different products here.
So, uh, talking about CCU or ops MVS, all of those products within the watch tower can actually utilize these solutions and say, Hey, I have a dynamic threshold break broken down or just breached. And it'll say, dispatch me the mainframe application tuner to get me the granular level of details. And one step back that I will say with AP four Z have said that the old pad for the at scale monitoring is less than 1% more like point over 1% with the granularity of information for, uh, the in-depth analysis, you would naturally expect that this is higher, right?
So, um, applica, uh, mainframe application to overhead while monitoring any of those workloads is typically somewhere close, around three, maybe 4% of the application run. And this is why you need to have it focused, and this is why the integrations of these products is really much helping the case. Um, so now coming to the, to the big wow, right?
So again, just to reiterate, typically the, um, mainframe teams are often lacking the expertise and time to optimize and modernize those applications. They don't necessarily have to link an idea where to start, and most of their code is older than a decade. Uh, and I'm being conservative here, right?
So, uh, again, watch Tower in the sense the products integrating with IT are allowing, uh, more closer collaboration in between of the products, um, cutting down the MTTR with any potential issues and resolving the critical problems, streamlining the processes. There's the big, uh, big thing as well. And, uh, allowing the individual, um, players within the observability to communicate better.
Simply when you are passing the information from stage to stage from, um, um, operator to subject matter expert optimization, expert performance engineer, and so on, you allow the next person to hit the ground running rather than, rather than starting from a scratch. This is, this is in my way, at least the multiplayer that has been mentioned in the previous segments here as well. So, uh, how these products like, uh, like AP four Z and Matt integrate with the watch tower at least these days.
So it's a continuous process, but with AP four Z, we can really, uh, call out a call graph or call three, um, of, uh, the individual module inter relations. And this is, this is similar to what, uh, the user has been experiencing. For example, with topology and other player from the watch star platform, it can visualize the trends in the CPU usage and elapse times as well as it has a very nice, uh, um, distinction of which languages your, uh, operation at scale is using on, uh, the mainframe.
And you can very much drill down even for a COBOL in the sense and see which of your COBOL applications are compiled with with version of a cobol. 5 10 to 20% CPU reduction usage on just a compilation, but I'm leaving out the, uh, thing that you can utilize the arch parameters and optimization parameters, and then in that case can even take you to the 75% savings on the, uh, COBO usage. So again, uh, a lot of those informations, uh, you can have a, in a modern way, easy way to consume, where typically the in-depth analysis from the user perspective requires a seasoned SSME to actually digest them because of the complexity.
The AP four Z in a sense is giving you the way in the relatively, relatively easy way, which enables not only the seasoned SMEs, but also architects, uh, developers and even new hires in the company to digest through the information. So I'm loving it. Uh, and um, I will take one more step to Matt before I just, you know, wrap up from my at least, uh, perspective.
So Matt integrates with, uh, Watchtower in a sense, uh, specifically through the, um, inter, um, alert insights, right? So Met can also, not only the other players within the portfolio, uh, which can invoke met and say, Hey, this is something is burning and dispatch me the application performance in depth analysis. But Matt, Matt can also watch for the, uh, CPU time elapse time XCCP or IO operations or service units.
And based on these thresholds which are dynamically refreshed and learned from the SMF records, it can, uh, intake, uh, it can issue an alert and provide a summary of that immediate measurement through the alert. So the silver lining is still being cascaded and passed from roll to roll and is beautiful. And I really like how these products are actually mutually benefiting from being part of the V tower ecosystem because thinking of it, uh, vow also allow us to, uh, do more on point and faster development of our applications because it allows us to collect more valid and quick feedback and validation with the client portfolio at the end.
So I'm personally looking forward what the optimization products can do and how great insights they can provide within the Watch Star platform going forth. But just a quick statement, Peter, you're absolutely right about the COBOL side. I think Tom Ross is here this week and he's talking about exactly that, you know, what version you should be using to compile and what performance me, Me absolutely expect to See.
So Yeah, there are Shows Only and only Cobo, right? Yeah. But the os vs co, you know, still a lot of it about, so, you know, it's an important point Absolutely.
Is reliability, uh, uh, uh, a big, what, what's, what's the, what's the performance in your, from your perspective? Um, performance can mean a lot of things, a whole ton of things. You're talking a lot about efficient resource utilization that leads to a, a typical meaning of performance just around sort of roughly sort of say speeds and feeds, but there's really a lot more to it, I think is, you know, especially with mission critical applications.
Um, so are there particular things that you would call out? I always associate reliability with these types of applications, reliability, availability, uptime, um, failovers, that sort of thing, but I'm, I can't tell if that's what you're talking about here as the pain point that leads people to say, Hey, we, we, we probably do need to, um, modernize what, you know, what would you say they are? Uh, my point of view on this one, that these tools are very much suitable, that, well, the quickest return of investment on these tools is typically in production, right?
But they can very much live in the test dev containers as well. So it depends whether you are approaching it from the AIOps perspective or operations or whether you are, uh, living in a DevOps space. And, um, I always, you know, in my mind, slip back to one of our, uh, VPs who had a article about the value of a second.
So respectively, let me step back. The performance in my case is typically boiling down to, you know, meeting the SLAs and getting the biggest bang for your box, because again, the competition out there is ferocious and you really need to stay on top if you want to win in the market and have the client user bases you need. And, uh, one of our VPs actually was saying, uh, about the value of the seconds, and I will try to be quick here, right?
So, so the whole article was about, you know, take a value of a second. For example, in sports, when you have a F1 car racing or someone competing for a gold medal, uh, even, you know, a split of a second difference, 10th of a second, 10th of a second can mean that you are not on the podium. And, and it's, it is a, it's a bad thing for you because you are training for it and everything like that.
But when you take it from the perspective of, of the, uh, sports, you typically have another season, another event to grasp for the victory. When it comes to business, it's much more worse because if you, uh, uh, for example, uh, take one of the big credit card companies, which is having a 10 trillion volume of was going annually for them, if you just do the math, it's like 300, a $300, three, sorry, $300,000 per second. Um, for any of the severity one reaction time before the people will cascade the information and react.
So the MTTR is really critical here, and there are much more other things related to that. It's a monetary impact. Obviously, millions of dollars are going out the door if you are not addressing your issues quickly.
And that's where watch towers observability being proactive can help rather than reactive. And uh, also you have a reputation hit and loss of a trust and so on and so on. So this is, this is really what performance boil downs, at least for me internally.
Performance is performance. Performance is clearly Broadcom. Um, uh, having customers come to them and say, we did the, we did the math and, and you know, we're, we're losing an average of, you know, a million dollars an hour or whatever it is, uh, because of this particular application that we are that's in production, we're concerned about, about updating and upgrading, come, come help us.
Um, and the terms of that have to do with, you said MTTR have to do with, you know, speed, which is the, which is the typical baseline meaning of the word performance mm-hmm. Really speed more than, you know, any other, any other elements of it. I think it's in, in a, in a mainframe context.
Uh, when we talk about, we also talk about, because if we, if we are in inefficient with the code, like using an old compiler that drives CPU usage, which drives software cost license, yeah. Now, so we, when we talk about performance, we're talking about the end user's experience, but also how efficiently we using the resources. And that all comes under that umbrella of performance.
And that, and that's where it can get a little blurred at times. 'cause if I'm the systems programmer, I wanna make sure I'm keeping to, you know, whether I've got tailor fit pricing or rolling four hour average. I wanna keep the utilization and machine down at a level that meets my budget aspirations, where the end user, the business wants subsecond response time and sometimes you can't have both.
So SLA is delaying costly hardware upgrades. Those are definitely things. And one, you are reactive to a problem.
One, you are ting the problem and the third time you are trying to utilize your box three x maximum potential. That's what, uh, optimization is. And it's a complex subject.
I know.