Observe – Analyze – Act. An introduction to HPE OpsRamp
HPE’s presentation at Cloud Field Day 23 introduced OpsRamp, a SaaS platform designed to address the challenges of modern IT environments. OpsRamp provides a unified approach to managing diverse infrastructure by focusing on observing, analyzing, and acting on collected data. This involves ingesting data from various sources, such as applications, servers, and cloud environments, into a central tool, enabling users to access all their data in one place.
The platform’s key features include robust analytics and automation capabilities. OpsRamp utilizes machine learning to assist users in analyzing data, identifying issues, and automating corrective actions. This automation streamlines issue resolution, potentially reducing resolution times significantly through integrations with over 3,000 systems. Furthermore, OpsRamp offers both agent-based and agentless monitoring options, providing flexibility depending on the type of resources being monitored.
OpsRamp differentiates itself by offering full-stack monitoring and an AI-powered analytics engine that can integrate with existing monitoring tools to correlate alerts across disparate tools. It provides both broad monitoring capabilities and integration of existing tools. The platform’s licensing model is subscription-based, determined by the number of monitored resources and the volume of metrics collected, with data retention policies tailored to different data types.
Presented by Cato Grace, Technical Marketing, HPE. Recorded live in Millbrae, California, on June 5, 2025, as part of Cloud Field Day 23. Watch the entire presentation at https://techfieldday.com/appearance/hpe-opsramp-presents-at-cloud-field-day-23/ or https://techfieldday.com/event/cfd23/ for more information.
Transcript
Thank you everybody. Uh, today, my name's Cato Grace, and I'm gonna be talking about Ops Ramp. I'm just gonna give you a really quick, uh, brief overview of what Ops Ramp is, and then we're gonna dive into a couple of detailed live demos.
So give you guys a, a real nice feel for the product, but just wanted to make sure that everybody has an idea of what Ops RAMP is. Um, before we get into that, so everybody's aware of, you know, all of the challenges that modern IT faces, you know, multiple environments in multiple different places, tons of data, too much data to really deal with, uh, you know, on-prem, off-prem, all kinds of different things that are coming at you. And Ops RAMP really aims to help you deal with all of those.
And we kind of see that as, as a a three-pronged, uh, strategy. So the one thing is, is getting all of your data into one tool. So all of the, the logs and traces and information that comes from, you know, applications, servers, network devices, uh, cloud environments, data from everywhere flowing into that one single tool, um, gives you that, uh, a way of having access to all of that data in one place and then allowing you to actually do something with it.
So, analyze it and tools to help you analyze it so that, you know, uh, the specific user operator isn't having going to have to figure out everything for themselves. They're able to get help with that, you know, machine learning, all the standard buzzword stuff, you know, machine learning, uh, ai, but really helping the user to figure out what's going on, get a better idea, a handle on things so that they're better able to, uh, work at getting things back to the way that they want them to be. And on top of that, one of the things that that Ops Ramp does that's kind of special is giving you an ability to actually automate that corrective behavior.
So when you, you detect that something is going wrong, you get all access to all of that data. You fi help correlate and figure out what that specific thing that is going wrong, and then giving you an ability to automate the repair, the correction, the, the fix of whatever that issue is. Um, so how that presents itself is, you know, you go from that typical, how long it takes to resolve an issue of, you know, hours to potentially days, taking that down to as little as an hour, just in terms of the resolution standpoint.
And then if you layer that automation on top of that, you can potentially lower that even more and get that, you know, get things back up and running even faster. But one of the, part of the way that we do that is that we have, you know, a huge array of integrations with Ops Ramp. So beyond, uh, all the capabilities that we're gonna talk about, we have integration with a whole bunch.
I mean, almost anything that's out there, 3000 plus integrations. Um, and then really quick, just how we do this or like, what the underlying architecture of it looks like is we have both an ability to install agents on devices, and obviously agents is gonna be something that you'd install, you know, on your, you know, windows machine, Linux machine, something like that, that would support installation of an agent. And then beyond that, we also have the ability for, to have a gateway that acts as a collection point for SNMP, uh, potentially APIs, things like that, that it can act as a, uh, a centralized point within an environment and then sending that out to the ops ramp cloud where, you know, the additional analysis and, uh, correlation and all of those other things are done.
And then ops Ramp also on the other end is able to interact with, you know, ITSM solutions with cloud providers, uh, orchestration engines beyond just what Ops RAMP can do to help, again, with getting things back to the way that you want them. And we have, uh, ops ramp environments, uh, you know, we call them pods, but, uh, obviously all over the world, anywhere where you would need them. Um, and really that gives you kind of a, a nice, yeah, sure.
Go ahead. SaaS solution. Yes.
Okay. Yeah, great question. SA Yes, yes.
SaaS solution Agentless, or you need to, you need to have agent, so You don't have to have the agents. Uh, it depends on what you, what you wanna monitor. You know, if you're monitoring, uh, a Windows machine, a Linux machine, an agent is going to be beneficial.
Yeah. Uh, but if you just wanted to monitor it with, you know, SNMP or something like that, you could do that as well. So, So it's monitoring and, uh, collection of logs.
Yeah. Monitoring collection of logs With a C cm, durable, sorry. Uh, like a CM sol, uh, solution, something like this plus analyzes on top of that.
Yeah. Uh, well, Well, um, so ops wrap is I guys, uh, Genea, Papo also with ops wrap, H-E-E-H-P-E. So, um, ops Ramp has the capability to ingest logs, ingest traces, and do analysis of that.
Um, a lot of seen platforms also have that capability, but when we're parsing and analyzing these logs and traces, we're doing it more from a perspective of, uh, performance, um, and infrastructure and things like that. Right. But, um, you know, we do have the option to ingest these logs and traces via open telemetry.
Mm-hmm. And, uh, yeah, for, so, and for the monitoring capabilities, again, we do agentless monitoring using that ops ramp gateway as Cato mentioned. So again, we can do agentless monitoring of Windows or Linux, uh, for Windows, we're gonna go, the gateway's gonna go in via WMI or the Linux platform.
We're gonna go in via SSH. Now, as kiddo mentioned, we have the option to download and install an RA agent on your Windows or Linux machines. That's an option, right?
We, we understand a lot of customers are agent adverse because they have tons of different software, and the server ends up running like 5, 10, 15 different agents, right? And it bogs down the, the server. But, um, our agent is very small, very compact, very lightweight, is actually written in the go language.
So it could run on a very bare minimum of, uh, memory and CPU. And our agent, actually, we don't install a lot of code to do a lot of the things that we need, like the monitoring and to do automation. We're leveraging what's called remote script execution.
So we're just leveraging, you know, what's built in, like we can, if our agent, you can automate patch management. If you wanna do patch management at the operating system level for Windows or Linux, we just call those remote scripts to do the, the, the scanning, the patching and remediation, right? Sure.
Mm-hmm. Okay. So, um, I understand what it is.
Mm-hmm. Kind of, um, so there is a lot of, maybe not similar, but from the same, uh, family of software, uh, your, your competitor solutions. What is your value proposition?
Yep. Great question. And this is, I'm gonna revert you back to the, the wheel diagram here.
Okay? So, um, the first aspect of that platform is that unified observability, it's essentially monitoring, right? You guys all are aware of different monitoring platforms, uh, SolarWinds, science Logic, um, um, logic Monitor, but, but, but monitoring, um, pulling in those metrics, pulling in those events, pulling in those logs and traces, right?
Some platforms are better at monitoring network, for example, some platforms are better at monitoring storage and some in compute. What Ops Shop is trying to do is we wanna, we wanna unify all of that monitoring full stack, right? From compute network storage, virtualization, containerization for infrastructure on-prem in the public cloud.
We'll support monitoring AWS Azure, GCP. We even support Alibaba Cloud and Oracle Cloud, right? So we're equally strong at monitoring and observability for on-prem infrastructure and cloud infrastructure, right?
Some platforms are better at, you know, one or the other. And, but we don't, I mean, that's just the first aspect of it is that observability and monitoring, we can also go in our AI powered analytics. What we're doing there is because we can monitor full stack, we can, you know, um, generate these events and alerts that SRA is generating.
But Ops Ramp can also integrate with existing monitoring tools like your SolarWinds or Science Logic. Because if you don't want to, you know, replace your monitoring tool right now, maybe you still have some years on the contract and say, look, maybe I'll keep it around for another year. SRA can still coexist with that platform using a web hook.
We will ingest the alerts that those third party monitoring tools are generating. Uh, or even like an A PM, if you're running an application performance monitoring tool, like, um, Dynatrace, um, new Relic, right? Ops Ramp can integrate with those platforms, uh, by again, using a web hook to ingest the alert.
So in that way, ops ramp can act as a monitor of monitors, right? And then we can do, um, once we ingest all of those alerts, now we can apply our machine learning to do alert correlation, right? Because the problem, uh, that a lot of our customers and potential customers are having, because they have siloed tools, they have a separate, they have a tool just to monitor network.
They have a separate tool to monitor storage, another tool to monitor, compute, another tool for database, another one for cloud, because they have those disparate tools. Those tools are only looking at their alerts. Typically, when you have a root cause issue, like let's say a network switch port goes down, that one issue on the network is gonna cause multiple issues up and down the stack, right?
And if you have disparate tools that only look at a portion of these alerts, it's almost impossible to, for you to do a correlation. What ends up happening is people get on war rooms, right? You have the network guy, the server guy, the compute guy, everyone's there, and they're all looking at their separate tools and they're trying to correlate the issue.
They're doing correlation manually. Very expensive, very complicated, very time consuming where it can do that correlation in our platform. Again, because we're seeing all the alerts, because we're natively monitoring or we're bringing it in from their existing monitoring tools.
Do you see the security team, uh, as a consumer of the tool or as a receiver of what comes out of this? Uh, both really, right? Yeah.
Um, I, ideally ops RAMP wants to, to ingest data from everywhere and, but if you want to, um, ingest, uh, or take some of the data that Srap is generating, we have APIs where we can push our metrics, our alerts into their platform, right? But typically, ops RAMP is the platform that's gonna monitor your full stack, compute, network storage, virtualization, anything. So that, you know, you know, we can do that or correlation and do automation.
Yes. What's the process like to create that event correlation? You mentioned like the machine learning can do it, like I, I'm sure that everybody would love for it to be automated, but also there needs to be validation that yeah, these correlations are correct.
Yep. And how do we figure that out? Great Questions.
Um, when I go through my actual demo, we'll, we'll, we'll, uh, deep dive into that. Great. Alright.
Does that answer the questions? Sure. Can we go ahead and as we run through the demo, I'm gonna try to, uh, demonstrate each of these and then we can deep dive on how we do monitoring, uh, how we ingest the logs, um, and we can deep dive into that, um, machine learning that alert correlation, and deep dive into our automation as well.
So there are two, two distinct products, ops ramps separate from op ramp mod automation. Is that how I understand It? No, it's one, it's one platform.
It's just, um, ops ramp platform. And so it does all of that in that platform. So you don't have to purchase a new automation portion or license for it or something like that.
It's all part of the system. Well, So we are, you remember I mentioned agents, we have agents that you can install on your servers, right? If you have our agents installed on the server, then we can run automation natively because we're talking to our agent, right?
But on, if you wanted to run automation, let's say on a network device or on a storage device, we don't have an agent for those devices. So if you wanted to run like automation, let's say on a network device, typically we would integrate with something like Ansible, right? Take, uh, take advantage of Ansible and all their playbooks that they have.
Um, now, ops rep recently acquired a company called Morpheus Data. I dunno if you guys are familiar with Morpheus data. Okay.
So HH oh, I'm sorry, not LP has been acquired. H-P-E-H-P-E acquired Morpheus data about four months, four or five months ago. So Morpheus data is about automation.
It's like automation on steroids. It's basically does automated orchestration and automation. So, um, um, in the very, very near future, you're gonna see that, we'll, we'll, we're gonna have, um, an integration ops shop will have a Morpheus integration so that now we can leverage, um, Morpheus automation.
And again, Morpheus can act on, uh, compute network or storage. Okay. Can I, can I have a question before we will go to the demos?
Yeah. So, uh, the analytics and prediction, is there any sort of a access or connection between this tool and let's say your supports teams? So like, you know, HP has a, like, you know, better insight into their clients or customers structure.
Are you talking about like that? What is that thing? Um, info on that?
Um, Yeah, I mean, like, you know, let's say like, you know, I'm a customer. I'm running HP servers and like, you know, like some of my, this drives are dying. Like, you know, like, you know, is there any sort of like, you know, proactiveness that your team will pick it up automatically, you know, in terms of the support Or Oh, yeah.
Oh yeah. Okay. So, so, um, when you deploy the app platform, we do have integration.
So that, uh, if you wanted to, for example, automatically create an incident when a particular alert, uh, triggers, we can do that, but we integrate with whatever ITSM platform the customer's using, let's say they're using ServiceNow. Mm-hmm. We can buy, you know, uh, configure a bi-directional ITSM integration with ServiceNow, but it doesn't right now, does not link directly with HPE support.
Right? Okay. Not, not at this time, but that, And HP support team doesn't have the access into it.
So like, you know, when you're troubleshooting something, Um, they potentially can, I mean, but you know, the, if they, they give access to the support team, right? Because when, when they, um, purchase the platform, they own the platform, and so they own, you know, the, the, the user rights and privileges. Yeah.
If they want, they can grant access to particular users. Okay. Okay.
Cool. Thank you. But there is no direct hook into, um, HPE support in general?
Not at the time. Okay. Okay.
But that would be like a pretty good, uh, Thank you. Uh, you mentioned, uh, Morpheus. So I imagine that this is, this is possible to be, sorry, integrated inside Morpheus, Right?
Um, so, so right now, um, since the, the acquisition was relatively new, right? So we're in the process of building out, uh, what integrations make sense. Maybe there's an integration on the Morpheus side to pull in ops ramp or an ops ramp pull in like Morpheus, right?
So, but right now they're both standalone platforms. Um, Morpheus is not, uh, SaaS based. It's Yeah, but, but Ops ramp is SaaS based.
I Know, but it's not possible to be integrated. I mean, not yet. Not yet.
No. There's no native integration yet. But there, there will be, just like abstract has an integration of Ansible, the next logical step is build integration of Morpheus.
Morpheus is much more powerful. I mean, Morpheus even integrates with Ansible as well, or Terraform or whatever, Morpheus, integr almost everything. So how do you, how do you license the customers?
Like Figure out, oh, uh, the s shop platform. Um, so the s shop platform is licensed on a subscription basis. Mm-hmm.
Uh, depending on the number of resources that we are monitoring, right? And so what is a resource? A resource could be anything we monitor, um, a server, physical server, a vm, a container, a network device, a wireless access point, and even, you know, um, uh, public cloud resources, right?
But it's, um, so let's say if I wanted to monitor a server, one server is, it's a one-to-one for metering one server. That's one resource in ops app that we're monitoring. 'cause not everything is one-to-one.
We have a four to one ratio. So if you wanted to monitor, let's say a wireless access point, that's a four to one ratio, I can monitor four, um, access points, and that's only billed as one srap resource. So we typically give you, um, like a resource calculator, and we tell the customers how many of each of these resource types are you gonna monitor?
And we'll take into consideration the count, whether it's one to one or four to one, and out pops the number of how many, uh, resources off ramp resources you would need. Uh, you said a Container is a Resource. Uh, it could potentially a container R four to one.
I can monitor four containers. That's only a co. Um, and, uh, one SRA license, you, You don't charge for, uh, storage.
Um, so aside from the number of license, uh, resources, each resource you're allocated, um, to pull in 50 metrics series per resource. Okay? Mm.
So let's say you are licensed for a thousand resources. So a thousand times 50 is 50,000. I can pull in 50,000 metrics, okay.
Per month. Okay. That's an aggregate.
Some, some resources. You're only gonna be pulling two metrics. Like you're gonna be doing a ping only, right?
Ping only returns to metrics response time, average and packet loss. Some resources like a database might pull in like 50 or, you know, 60 metrics, right? Like, I think what he's getting at, like what about data retention?
If I wanna save my data for longer, that doesn't affect, Depends on the data type. Um, let me actually, like, I can refer you over to our documentation site. Let me jump into our, to my, uh, environment here.
com. And then from here, so the data retention, uh, from this website, you can just type in, uh, data retention. And this is our default data retention policy.
I'm gonna scroll down. So depending, I'm gonna scroll down. So depending on the data type, and we're gonna see the data type here.
So data type, let's say metrics. So for metrics, all those raw metrics that we're collecting, we're gonna store those for 12 months by default. Right?
Um, what about these alerts? Uh, if it, if it's an alert, it's open, it's gonna stay in there indefinitely, but if the alert has been closed or suppressed, it'll get aged out after 90 days. Okay.
Um, we can even do, let's say, um, network configuration backup. When we talk about our network observability, we can back up the network configurations of these network devices and we store those for, for one year. Okay.
These are our default data retention policies. If you need to extend it, we can do that. Okay.
Any chance you're gonna have like a, a free tier, you know, limited use for enthusiasts or Customer? I, I, I think we are in a pro. We had that in the past, and I think we will have that, like a kinda a demo environment where you can go in and, and um, you know, test it out onboard, monitor your infrastructure and it's got good for like 90 days, like a try before you buy, right?
Did that answer the question? Alright, so, um, So what happens to alerts and things that, that go beyond the, uh, thousand messages per month or metrics per month? Um, Let's say you have, like you said, 50 resources.
So you got 50,000 and 50,001 comes in, um, Oh, then, then you are charged for the overage. There's an overage if you go above the, the metrics allocated, but remember it's an aggregate again. Um, oh, I understand.
Yeah. I'm just trying to understand, you know, that that 50,001 might be a serious, serious alert. Yeah, I can throw it away.
Oh, no. But, but no, we, we don't throw away any data. So let's say you're, let's say you're, um, you know, you have resources to monitor a thousand, but you onboarded like, you know, 1200, we, we won't drop or anything.
We'll, just, we're gonna, we wanna, we're gonna rightsize you every month. Say, Hey, you're over by about 200 licenses. But we never, never drop data.
I.