Minimizing Application Downtime with HPE OpsRamp
HPE’s presentation at Cloud Field Day 23, delivered by Juden Supapo, focused on minimizing application downtime using HPE OpsRamp. The demonstration began by displaying a dashboard that monitored a mission-critical ERP application. The speaker highlighted key performance indicators and the overall health of the application. He then presented a service map, a visual representation of the infrastructure supporting the application, including database nodes, servers, and network devices. The service map enables administrators to quickly identify issues by visualizing the relationships between various infrastructure components.
The presentation then illustrated how OpsRamp handles infrastructure issues. By simulating a database outage, the speaker demonstrated how the dashboard and service map responded in real-time. Alerts were triggered, and the service map indicated the location of the problem. He emphasized the alert correlation feature, which uses machine learning to group related alerts, identify probable root causes, and streamline troubleshooting. This grouping allows administrators to address the primary issue instead of dealing with numerous cascading alerts, thus saving time and improving operational efficiency.
Finally, the presentation concluded by showcasing automation and governance within OpsRamp. When the simulated outage occurred, the system automatically generated a task, including an email notification. Through a low-code, no-code process automation workflow, the speaker demonstrated how the system could trigger a script to attempt to restart the affected service. This showcased the platform’s capability to combine automated remediation with governance through approval processes. This combination of features minimizes downtime.
Presented by Juden Supapo, Solutions Architect, HPE. Recorded live in Millbrae, California, on June 5, 2025, as part of Cloud Field Day 23. Watch the entire presentation at https://techfieldday.com/appearance/hpe-opsramp-presents-at-cloud-field-day-23/ or https://techfieldday.com/event/cfd23/ for more information.
Transcript
So the first part of our demo, um, hello everyone. My name is Juden Sapo, and we're going to demonstrate minimizing application, do downtime using the HPE ops RAMP software, assuming you have the ops ramp platform fully instrumented, we've onboarded the devices that we want to monitor, right? We're pulling in those time series metrics, we're pulling in logs, um, and, um, right?
And so we we're gonna have a dashboard where you could look at, you know, the health and performance of your application. So I have a, I have this dashboard. It's looking at this mission critical ERP application.
Okay? Now, again, um, depending on, you know, what your persona is, if you're a network admin or a server admin, you're gonna have a dashboard that's curated for what you wanna see. Okay?
So I am, um, an admin for the ERP platform here. So I have this dashboard, it's telling me, you know, my, my ERP mission critical application is up and running. I can see some key performance indicators, uh, active users, current users, things like that.
Okay? And so now this is a mission critical application. So, um, on the RA platform, we've built what's called a service map.
And a service map is a kind of a, a logical way to, uh, visually display all the infrastructure, which is supporting this mission critical ERP application. You can see we have a database, a node here, we have ERP servers, web servers, we have network devices, right? Because ops ramp can natively monitor all this infrastructure.
Question uhhuh. Yeah. Uh, so you, you need to define this map by yourself, or you have discovery tool that is building this map For you.
Yeah, great question. Okay? So when, during the onboarding process, you're gonna go in and discover your infrastructure, right?
Your servers, your compute, your network, and your storage. Um, depending on the resource type, that discovery is usually automatic, right? Like we, we, we integrate with like the vCenter server, for example.
So we can discover all your ESX hosts, all its VMs and data stores. So infrastructure is automatically discovered, and if there is a topology, like network devices has a topology, we can visualize and draw that topology. That's all automatic.
And I can demonstrate that a little bit later when we talk about the networking. But this here, this, this is drawing an architecture for an application, but it, it does require manual setup. I do need to put the nodes here.
So for, because this is user defined. Mm-hmm. Okay?
So I'm gonna put a node here for networks. I'm gonna put a node here for database. Now, the way I allocate no, um, resources to these nodes is I specify a query.
So I can say, for example, um, if I wanna allocate routers here, I would say if the device type equals router and the IP address is this or whatever, we have discover, we discover attributes of a resource. So you can use some of those native attributes to do the automatic assignments here. Also, um, if there isn't a native attribute that we can use, you can use a custom tag.
You can say like, um, you know, this database is dedicated for supporting ERP. Okay? Now, we, we, we use these queries because that way it's dynamic.
If I onboard new infrastructure, like a new database, and I put the appropriate tagging to associate it with ERP, when that new database is discovered and monitored, it automatically gets associated to these nodes, right? That can kind of minimize the care and feeding that you have to do when you, you know, when your infrastructure changes, right? Infrastructure's going up, it's going down all the time.
So we wanna be able to, to kind of, um, automatically allocate these resources to the appropriate nodes dynamically. Or again, you can do it manually, but that's, that's, um, not, uh, uh, for best practices is not the best way to do it. Does that make sense?
So I, I have to do, I have to put the framework in there and then put queries so that I can appropriately associate the database in this database node, the web servers here. Now, I only wanna indicate infrastructure, which is supporting this ERP, okay? Because as, 'cause now as we're monitoring this, this entire infrastructure, if, and if, if these alerts start coming in, you're gonna see that this service map is gonna light up and this dashboard's gonna light up as well.
So I'm showing you right now the happy state, right? Everything is green. You know, my, my, my, uh, CTO, my VP of infrastructure, he is happy 'cause this application is healthy.
It's up and running. So what I'm gonna do here is I'm gonna introduce an infrastructure issue, and then we're gonna see how it lights up this dashboard lights up the service map, and we can very quickly pinpoint where the issue is. Okay?
So I have this little, uh, UI here where I can actually simulate taking down, or that it's actually simulating, it's actually, uh, it's a little script that goes in and takes down the server. I can take down the database and I can take down the network. So I'm gonna do this, I'm gonna take down the database, okay?
Now, there is a little bit of, of lag time because, okay, I, I ran, I, I ran this, um, a script is running now, it went in and, and shut down the database, but we're monitoring infrastructure at the most granular, we can do one minute monitoring, okay? So every minute we're monitoring the infrastructure. So I have to wait at least a minute, maybe a minute and a half before we start.
Oh, we already see some updates here. You see that the database looks like the database node went down. Um, so I'm just gonna wait a little bit more to, to fully let all the, the alerts come in.
And you see this, this little app that was running here, it's now gone because the database is gone. Okay? So now look, status is down, database is down.
And then if I drill back down into that service map, I can see now very clearly where the issue is, right? It's a database issue. So let me drill into that.
And notice there is one alert here. I'm gonna go ahead and click that alert. And notice, this is where we're in now, our alert browser, we're drilled into this one alert.
And if I, if I hover over this, there's some little icons here. This alert is what's called an inference alert. And what an inference is, let me, let me bring up another slide real quick just to show you what's going on.
So if I stepped in front of the screen, will the camera still pick me up? No. Okay, then that, then I won't do that.
Okay, so this is what's going on. I, I triggered an incident. It was a database, but in this example here, um, let's say the alert came from a network device right here, right?
And so this alert at the bottom, maybe this, um, port on this switch, number one went down, because that port is down, it's causing what's called cascading impact. It's causing this alert on switch number two. And then there's a ESX server here that's connected.
Using that port, it's gonna trigger an alert. And the VMs now lost connectivity. So our machine learning can ingest all of these alerts, and it can, it can, um, using the machine learning, it can group these together because it's causing cascading impact.
And we'll, we'll talk into a little bit of detail on that, but this is the kind of scenario we can, we can discover. So when we group it together, we create what's called an inference alert. And an inference alert is simply an alert data type, which, which links and groups, these five alerts together.
Now, that's what we're doing for alert correlation. Now, when I send out an alert notification or automatically create an incident, I'm not gonna create five separate alerts. My, my machine learning has already deemed, look, these are all five of these are related because it's causing cascading impact or common impact.
So we create one inference, and in that inference is a group of alerts, which we've correlated together. Does that, does that make sense, guys? So now we don't have to go in and individually triage these five alerts.
We give you probable root cause and it's usually the first alert in the grouping. So if you just address this alert, it should in theory, fix all these other ones that it caused. Does that make sense guys?
And in large infrastructures, you know, one root cause issue, like a switch port going down, that could trigger 5, 10, 15, 20 other alerts. Are you gonna go and triage all 20 of those alerts? No.
Correlate it together. Find a root cause, fix that root cause issue, and that should, in theory, fix all the other alerts that it caused. Okay?
That buys you a lot of operational efficiencies, because now what happens if you don't have a a, a mechanism that does this, your network team is gonna see these two alerts, and they may not even know that these two alerts are related, right? If it's a large network, they won't know that that's related. And then your virtualization server team, it's gonna see this alert.
They're not, and they, they, they, they're not seeing these alerts, right? 'cause they're only monitoring servers, right? And maybe the database team, they own these VMs, they're gonna see these alerts.
So you can see that's where you have to get into a war room and you have to have these three guys get together and find out, do that manual correlation in that war room, right? Does that make sense guys? Okay.
So that's what we're doing here. We made, this was the inference. And if I drill into that inference, you're gonna see there was actually six alerts that were triggered when I took that database down.
So here, there was the actual database going down. Let me make this a little wider here. This is the database going down there.
The, the, the platform node went down, the database node went down, um, looks like this was here, unable to connect to device 'cause the database is down. And there was even a synthetic monitor that failed. If I drilled into the service map, do a quick refresh, the synthetics are a little bit lagging.
So now they finally came in and look at that. Even the synthetics are affected, but not everything in the synthetics, the admin consult is still up, but the, that the dashboard is down for synthetics. Okay?
So now we have this, we have this inference correlated, these alerts. Now what did we do? We have automation running.
And the automation sent me an email. I'm logged into my email right now. This is the only time I went outside of the ops ramp platform.
All the screens that I showed you all inside the ops ramp platform. So when that alert triggered, it sent the team, which I'm included, it sent me an email because a task was created. So let me drill into that.
And this is what that task is saying. S-R-E-R-P, demo service SQL Server stopped on this host. And it was a task because this is saying I could potentially try to restart that service, but I have to give permission for the automation.
So what I'm demonstrating here is automation with governance. Because typically when you're run, you're working with production infrastructure, you don't necessarily want to run automation automatically, right? There's usually some governance, uh, involved.
And this, uh, proposition of, uh, automation, it's, um, generated by a system or you, you need to pre-configure it. Well, well, you, you have to set up, you have to set up what's called the process automation to do this. Okay?
So It's not proposing from, let's say it's own analysis. So like it's analyzing the infrastructure all the time and proposing some stuff. Oh, um, at this time we're not doing that.
Where I just get an alert and it says, based on this alert, maybe do this, this, this, uh, we don't do that right now. Um, on the optional platform, what we're analyzing is that alert correlation. So we're doing a, a, you know, correlation at the alert level.
Um, our, some of our other platforms like, um, HPE, Aruba Central, that will actually do some intelligence for you as these alerts on the network device comes in, they'll do that analysis. So they're a little bit further ahead and as, because they're only dealing with network devices, right? The ops ramp here, we're monitoring everything, compute, network storage, virtualization, containerization, databases, all of that.
But that would be, um, some, maybe some, um, enhancements that we can do. But right now, the way we're handling it is we have what's called a process automation. And I'm gonna show you that in a second.
But what I wanna do first is, um, so I showed you the happy state, and now this is the sad state, right? Mm-hmm. We have all these alerts, and then how am I gonna fix this?
So I jumped into my email, the automation created this email for me, and I said, look, click here to approve it and we'll fix your problems. Let me click there. So that email, when I clicked that email, there was a URL in that email.
This could have taken me to like, um, ServiceNow or to your ITSM. And I could have made the update there, or I could, this has taken me back into the, to the ops ramp. ITSM.
It doesn't matter 'cause the, the integration is bi-directional, okay? So it doesn't matter where you make the change. So if you look here under the conversation, um, this is what happened.
The SQL server stopped, please approve to try to restart this service. So I'm gonna, I, all I have to do is simply go here, click approve, and then hit save. Okay?
I gave the process automation permission to go ahead and run the script and try to restart it. But remember, I'm poll the, the service is probably back up and running now, but in order for the updates to change, I have to wait again, that polling interval, right? One minute, uh, like maybe a minute and a half.
But while we wait for that, the update here, let's drill into that process automation. So I'm gonna go under automation, under process automation. This is the process automation that sent the email, created the task.
Let me drill into this. Okay, so here is that process automation. Now what is, what is a process automation?
So the process automation is a low code, no code way. It's like a workflow engine, which, which I can use because what is the end-to-end workflow that you want to run to restart a window service? 'cause that's what happened.
The database was running as a window service and that database went down. And so we have a starting point where we look for the alert. You notice here for this alert, we have what's called a filter criteria.
So we're looking for an alert where a Windows service has gone down and it contains the subject contains of ERP. Okay? But it's specifically just for this ERP application.
When that alert came in, then the workflow says, okay, create a task. This was the task that was created. This was what was emailed to me, right?
That task was created. And then that task has a, has a switch here. If I approve it, which I did, then it's gonna go ahead and run that script, a script to try to restart it.
If I had not approved it, it would've taken this code path and just done nothing. Okay? So, so this process automation is a layer that we put above the scripts.
We're not calling the script directly. We, we, we wanted to put a approval process in here. That's why we have this.
If I wanted the automation to run completely without the approval, I could have just gotten rid of this. And once that alert is detected, it would've gone here and ran that script. And the script that we're running, this is the script that we're running.
Let me, let me bring it up. It's under scripts. So that whole workflow mm-hmm.
Has to be pre-configured for different alerts that might come up. Yes. You would have a, so for typically what we do is we identify with the customer.
Are there any issues that you guys are seeing repeatedly come up? And is there a way to automate it? Right?
We, we wanna tackle the low hanging fruit. Maybe there's one typical issue that's always happening. Um, you know, automate those.
'cause you already know what to look for and you can just, um, you know, if you have a power shell, shell or, um, a Python script to run, you can just call that. So this is the script that, that process automation is calling. It's called restart service.
And if I edit this, it's a simple, oops, not this, uh, edit here. The script we're calling is a simple PowerShell script and we're just getting what was the service that went down? Remember the alert that triggered, we got the host name from the alert and we got the service, the SQL service went down.
So we grabbed that and we're just simply trying to do a restart of that service. Okay? So this process automation is what provides intelligence to your automation, right?
We don't have to like, manually run these scripts. Kind of weird to manually run automation. If it, if it, if it hits this filter criteria, it will run this process automation.
And I could have made this. So here, after I run that restart service, what else does it do? It posts a comment, it'll put whether it successfully restarted the service or not.
It'll do a wait here and then it will go in and close the alert. Uh, but I could have had this, this process be data driven. I could have called script one, and then depending on the output of script one, I could have ran another script, right?
So you could automate complex workflows using this process automation. Okay, so now it's been way more than one minute. Lemme just take a quick refresh over here.
Notice now everything is magically back up and running green status is up, database is back up. If I drill back down into the service map, that's green as well. I'm gonna pause there.
Questions or comments about anything that we see here? Do you wanna, do you wanna drill down on any aspect of what I showed? Yeah.
Is email the only way to get, uh, responses out of the system? Nope. We can send, so we have the con, we send a notification, right?
If you wanna be notified about a alert, then you as the user can configure how you wanna receive that notification. For example, if you're a network guy, if the alert is on a network device and it's critical, I can have it set to send me notification that notification via email, text, or VO and or voice. If it's critical.
If it's not critical, then only send me email. What about, uh, webhooks into Slack teams? Yep, We have an integration for that.
Okay. Uh, oh. Going back to our, uh, um, documentation site here.
Um, hit that and search, search for, and I'll put this in the, in the link here, search for integrations and you'll see, uh, a full list of all our integrations that we have available on the platform. So on the left side here, you can see all the different compute that we can integrate with all the different network vendors that we support. So the ops rep platform was developed, uh, back in 2014 by a managed service provider for managed service provider.
So that's why from our very beginning we have to have, um, multi-tenancy support and role-based access support as well. Mm-hmm. How much of this, if any, is, uh, exposed via API?
Like what, what could we do with this if we didn't want to purely consume your interface? Great question. So let me put in the chat real quick.
Um, that link to this integrations page here. That way you can kind of peruse all our integrations. com in the, the home you see here, there's a link for API click on there and this documents our APIs.
So there's an API to talk to the ops ramp platform in the cloud, and there's APIs to talk to our gateway. So you can use these APIs and a lot of these APIs are, you know, use, you know, to get metrics to do monitoring. Um, whatever you can do manually, you can do through an API also, right?
Some people wanna script it rather than do things manually. The question here is, do you have any, uh, templates of, uh, these automations based on your many years of experience? You know, if you see that it's S six, maybe it's proposing you to use that kind of automation.
Yeah. So there, there is some risk when, um, you know, you, you, you know, get, grab a script from somebody, right? Uh, Like a template.
I'm not saying that you just click and use it, it just take it from, We we do, we do have a library of, um, some of these scripts that we use. Um, but you know, we, we caution our, our, our customers like, you know, here's some scripts that we have available to, to restart a window service or to do a, a disc cleanup on a system. But we also have templates for those process automations as well.
But, um, we don't, we don't, um, give them out of the box. Um, typically when you're doing implementation of the op option platform, you're gonna be working with our professional services team. They will have a library available for the customer.
So it's really, um, on a, on a, on on demand base. So It's on, on the, uh, integration moment. Like professional services do it for you.
You cannot take by yourself from this catalog and internally change them and do, uh, and prepare them for your infrastructure. Um, you always need to take a partner who have access to it and they sell you this as a service. Um, no, it, well, if, if you, if you, if you know how to, um, configure the integrations, I mean, you can go to setup here under account and, uh, to install integrations.
We have a integration, um, library here. So here under integrations, I can click there. And so here are kind of the ca the category of, of integration.
So for integrate with automation, compute network, uh, network security, we have integrations with like single sign-on. Uh, right? So yeah, you can, you can configure these integrations yourself.
Um, but some of these integrations are a lot more, um, um, technical and time consuming, like setting up that ITSM integration, right? You kind of have to know how to integrate Necessary. I'm talking about integration more about those workflows, uh, that, you know, that I think it's that, uh, actually automate the process to recover from something, from something like restart my service, right?
I I believe that it's very generic Yeah. Template for like 99% of customers who have Windows will use it, right? Yeah.
Also, because you mentioned at the, at the beginning the AI interaction. So I mean, this kind of, um, of automation would be even simpler than, for example, Nagios interaction. So, uh, I think that the question he did, Yeah, so, so, but the, the, so yeah, to, to answer the question really is, um, we give the customers the ability to leverage any existing scripts that they may already be using, right?
I mean, here, uh, under automation, we have scripts. So here they can, um, upload any PowerShell shell scripts or, you know, Python scripts that they have. We don't have like, you know, uh, an exhaustive library like Ansible does, right?
Ansible has a huge list of, uh, automation to automate, uh, network devices, but we give the customer the infrastructure to put their scripts here, and then now they can, uh, design process automation to leverage those scripts. But our process automation allows us to kind of like run one script and depending on the output of that script, call another script. So we can really kind of, um, create these customized flow, um, based on your scripts, right?
It's, it's a more, um, structured way to kind of execute your scripts, if that makes sense. Is it, is it typical for a SaaS application to require professional services to install? Um, sometimes, yeah, because some of these integrations are, are pretty, um, uh, it requires some subject matter expert with the device that you're integrating with, right?
There's a, because when we integrate, let, let's say if you're ServiceNow, ITSM, right? We gotta, you know, set up the, the API, the URLs and make sure the permissions are, are set, uh, appropriately. And that professional services is part of the Yeah.
Part of the implementation process when we, when we go through the implementation process. Yeah. Yeah.
But I, who's paying for it? I guess, where is it paid for? The, The, the customer typically will, will pay for that separately.
Yep. Separately.