Democratizing Observability in Complex IT Environments – Shahar Azulay, groundcover
Groundcover CEO Shahar Azulay dives into what it will take to make observability both more accessible and affordable for a wider range of organizations as IT environments become increasingly more complex.
Transcript
This is Textron tv. Hey guys, thanks for the throw. We're here with Shahar Azulay, who's CEO for ground cover.
And we're gonna be talking about observability and why it's not quite as easy as just as it sounds. It takes a lot of a work, b resources, and seeing money, and we're gonna try, if we can figure out how to make that easier for all concern. Shahar, welcome to show.
Thanks, Mike. Happy to be here. One of the challenges that people are running into, uh, on paper at least observability sounds like a good idea.
Predefined metrics aren't enough, and we got logs and traces and all this good stuff to look at, but putting it all together seems to be a challenge. Yeah, I mean, eventually the bottom line is kind of what occupies people's minds at the moment. 'cause, um, um, observability is like the second bill in line after the cloud providers, which, you know, by themselves is like the, the hugest or biggest bill any company pays, uh, for, you know, set running up its infrastructure in the cloud.
So eventually when things come up to about 10% or 15% of that, uh, cloud provider cost, uh, you see organizations start to wonder, are they getting the proper RRI from the vulnerability stack where they're just spreading money on, you know, logs, metrics and traces they don't need. Uh, that's kind of where we're at in a situation where people can't afford to actually pay these prices and starting to look for options And how do I be smarter about what I'm collecting? 'cause I think there is a tendency to just collect everything and then you wind up paying these bills and then somebody starts asking some difficult questions.
But is there a better way to go about doing all this? So I think the, there's kind of a deadlock in observability. 'cause uh, eventually part of the data verticals are you're forced to collect, right?
You, you're forced to collect logs for compliance and all, all sorts of stuff like that. And, uh, basically people can't revert back from data. They can only have more data and kind of consume more data.
And, uh, trying to have peace of mind with, uh, you know, more data to kind of cover you in case of an, uh, problem. And, uh, the reason we got there is because, uh, the structure of, of observability solutions on the other end, the vendors basically, they're kind of, uh, not on your side when it comes to, uh, how you should actually operate. 'cause they were built solutions like Datadog, new Relic were built to be to kind of hoard your data in a sense.
So the more data you push in, the more money they make, the more promises that they can throw in the air of, you know, you being covered in a case of, uh, of a malfunction or downtime, but you actually query like half a percent of this data maybe on a good day where you see the part on the keyboard and, you know, just punch into Datadog all day. So they're not there to help you. The more you push in, the more you're gonna pay and everybody's happy on the other end, the Datadog.
So I think, uh, that's, that's the main issue that we're currently facing. So nothing is gonna resolve that problem on its own. And you see that with companies, uh, complaining about, you know, Splunk and Datadog costs and stuff like that, kind of spiraling outta control.
They're not gonna support you in that. The, the, the, the only incentive they have is to convince you that more data makes more sense. So, and, and that's kind that kind of forces engineering teams, uh, to make hard decisions.
And mostly, uh, what we're seeing in the last, uh, couple of years is that people are dropping out out of the APM tier, which is basically, you know, we've been promised the tracing and application metrics is kind of the way to go with troubleshooting and reducing downtime and, you know, reducing MTTR. But eventually we see that, um, like 10 or 20% of organizations actually actually activate the steer. Most of them just rely on logs and infrastructure metrics.
So in a sense, even though like a decade has passed and, and you know, solutions like Datadog, new Relic and all the others are offering these solutions and even OpenTelemetry, right? You kind of solve tracing with solid application metrics, uh, yet like a fraction of organizations use that. Uh, and they're using sticks and stones with logs and metrics.
And it goes to show you that something is broken in the way these solutions are built, priced people are just not getting, uh, to a point where it makes sense for them to utilize these tiers, which they should. So what are the alternatives? I mean, are there other ways of thinking about this?
Can I, uh, store data more efficiently and rehydrate it when I need it? Or are there ways to think about observability that are just fundamentally different? Yeah, so part of the, the, the main problem is that eventually, uh, solutions that collect observability data are usually running as part of your code as part of an SDK instrumented into your code collecting application metrics and traces and all that.
Uh, first it's really hard work. I mean, uh, if, uh, if any engineer here in US or DevOps here in us, uh, have tried to instrumental telemetry in a real company, when you have like 15 teams, you, you know, working in different, uh, tech stacks and different languages, that takes a long time and a lot of effort. And once you do that, you have to become an expert of, you know, what's the sampling rate and how do you, uh, push in all the data.
Um, so there's a lot of hard work around that. But even when you do that, you compromise your applications because, uh, people are measuring, uh, new Relic and OpenTelemetry and Datadog, uh, impact on their application. And you see that, uh, these solutions eventually take up a lot of CPOA lot of memory and even, uh, can affect the response time of our applications in some cases due to the fact that they're running as part of your code.
So eBPF is one of the solutions that the industry is moving into. Eeb PF is a technology that ground cover is harnessing, and basically it allows you to monitor observability or to get observability data basically directly from the Linux kernel, which would be the operating system where your containers are running on top. Um, once you do that, basically you're running out of band from the application.
You're getting all this rich data about application metrics, traces and all that without actually running as part of, of the application code without requiring developers to instrument your SDK into their code. That's both safe, but also basically out of band from the application resources. So you can do more complex stuff like, uh, you know, sample the data differently, capture events that make sense, transform data like traces into metrics on the fly and not just, you know, later on at the, at, at data logs or new relics backend.
So you can be much more cost effective with data, uh, without compromising the app performance on one hand. And on the other hand, try and make sense of how, and, uh, how this basically data backend should be built. Like, so, uh, we see solutions like, uh, like modern solutions like Databricks and different solutions, basically storing data for you moving into a concept where the data plane resides at the customer end.
It doesn't make sense anymore to go full suss on everything, right? Not every logline has to go all the way to Datadog servers and me having to pay all of that, my egress costs, then my Datadog costs, 'cause they have to roll it out to AWS at the end. Like everybody has to roll their, their built database at the end.
Um, so we're also creating a different concept where that data is being crunched in your cloud environment as it passes to be more cost effective. But at the end, it doesn't leave the cloud environment if it if it doesn't have to, if you don't query it, if you don't look for this data. So that saves a lot of, uh, cost and allow you to control retention better.
So we can also price the data differently. And that's kind of the main, uh, main component or main pain point that people are currently suffering from. People hate the unpredictable pricing models that are basically volume-based, right?
If suddenly a developer pushes the pushes in a new log line, uh, you can get 30% more logs next month, you, you won't even know about it until you get the new relic or data. So, uh, moving into concepts where the data is in your control can also have kind of break that bond between volume of data, which you have no idea how to control. 'cause your developer teams are basically generating this data and, uh, the cost of this data.
So once that connection is broken, pushing in more observability data that makes sense suddenly, uh, becomes reasonable in budget and in effort. And that's kind of where, uh, ground cover is trying to, to help teams today, In theory, um, observability has always been a core tenant of DevOps, but I have to wonder, you know, what percentage of applications do you think are actually instrumented? Because it seems to me it's relatively low.
I it's very, very low. Um, I I I think that, uh, 10 or 20% of all organizations actually actually activate some APM layer. And when we meet those organizations, I think, uh, you, you'd be surprised by the variety of different, uh, situations we meet in these organizations.
Most of them have instrumented part of this, uh, uh, services just because it's super hard to go through all that. And eventually you have to pay the cost of all this data. So you see, you see them pick and choose and kind of instrument different stuff.
One team has done something the others haven't got into kind of doing that. Or they choose specific environment based on kind of, uh, verticals inside the company, specific products that have been instrumented versus others. So in a sense, even companies that we see that uses code instrumentation are, uh, far from her kind of hermetic in most cases.
Um, and I think part of the way you can measure it is kind of the adoption of OpenTelemetry. Everybody's talking about OpenTelemetry. Uh, it's a great concept towards kind of standardizing collection of, of, uh, observability data.
But in a sense, um, uh, the adoption in the the SDK part and actually instrumented the code is, uh, extremely low compared to what you would expect even so many years after, you know, this concept has been introduced. Um, so I think that that's part of the problem. I mean, these solutions eventually are engineering, correct?
They solve the problem, they can collect data for you, but they are hard to implement and collect so much data that eventually most teams just, you know, punch out. Um, with EVPF, we can help these teams get, uh, get the data with much less hard work. So there's no more pick and choose, right?
If you install an EVPF sensor on your Kubernetes cluster and you have 118 different microservices running in the cluster, which is by the way, according to a CONC survey, that's actually the number, uh, like the average organization have 180 different microservices. So imagine instrumenting all that. So if you, once you install the UPF sensor, you get to see all of them equally, like immediately, 60 seconds after.
So the, the thought of what to instrument and how to instrument is kind of taken away from you. And you can start focusing on what do you want to do with this data and how, how do you store it? And that's where we see the industry moving into.
I mean, eventually collecting, uh, all your data from an SDK instrumentation would make less sense after a while and you would have to use it in specific cases where you think you'll wanna get deeper or do some, do some stuff that maybe eBPF can can't do for you. But getting these 90% coverage, that should happen quickly and very fast. 'cause most teams need, you know, basic coverage.
What versions of Linux do I find EBBF in? Do I have to upgrade to that before I can take advantage of these functions? And are we waiting for that cycle to happen or is it happening now?
So I think that that question would, uh, make sense a few years ago. 'cause as, as, as you are completely right, uh, EVPF is not a new thing. I mean, it's been kind of boiling up for over 20 years, but in the last few years, kernel Virgin have become, uh, kind of the feature richness of UPF has become so usable that you can do like very complicated logic like e like a sensor freight for APM.
So around Linux kernel, uh, four or, uh, kind of layered than that, uh, which is already a few years back, we have enough feature richness in the eBPF uh, virtual machine to do practically what we need to get, uh, application metrics and traces. The fun thing about it is that if, if you combine that with Kubernetes, suddenly you don't have to worry about it at all. 'cause these, you know, managed providers like AWS Azure and GCP are pushing the newest Linux version for you.
I mean, eventually most teams deploying Kubernetes are deploying it in a managed way, like E-K-S-A-K-S, stuff like that. And when you spin up a new cluster, you would get out of the box, the newest Linux versions, uh, available could 'cause the vendor has the incentive to do it for you. So, uh, ground cover in most cases meets uh, uh, kind of the newest Linux versions, which are very feature rich for U pfs.
So we can do whatever we want. And in some cases, you know, on-prem or kind of more legacy companies, it'll still take a, a Linux kernel a few years back to kind of not be relevant to, to run, uh, a modern EVPF sensor. So we're definitely in a clear zone that, uh, like 90% of companies right now are super relevant to moving into EVPF.
Definitely if, if they're in the cloud and, you know, look ahead two or three years more, there's gonna be a certainty because some, some kernel versions, uh, already in in Linux four are getting kind of into end of life. So people will move just from compliance to the newest versions very soon and, you know, can enjoy EVPF, uh, in the cloud in on-prem or whatever they're using. So we're very close to that.
What's your take on how AI might be applied to all of this? 'cause it seems like there's a lot of effort to apply algorithms both predictive and generative to observability, or are we gonna see more of that? Yeah, we definitely will.
I think one of the major things that we feel in observability is that people basically don't know how to query data effectively. I mean, you have two approaches in the market, either the no query language approach, which is, uh, you know, uh, nice for some things, uh, uh, fancy UI with some, you know, uh, things you can let developers and DevOps play with and kind of, uh, filter by a few keys they see in the ui. That's, that's a good approach.
We also use that, but in some cases you have to go deeper and build dashboards that build kind of custom panels that to display specific metrics. And that's really hard to do. I mean, uh, in, in a a hundred people development team, there's gonna be one or two people that know how to do that, that knows prom QL from PROMEUS or SQL for different, uh, sequential, uh, databases that they're using.
Uh, there there's only a fraction of these people who feel comfortable with it. And if you've seen companies using custom dashboards, you would usually see, you know, dashboards hanging on for a while. No one, no one really wants to touch that too much.
And I think that's a downside because you have so much observability data being collected, but people don't really know how to cross-section it and query it properly. And they kind of, you know, um, fall into specific categories and and queries that the organization has been used to work with. So natural language querying or prompting on that will definitely be a very prominent, uh, aspect in the next couple of years.
'cause if I can write that s QL query, which I don't know how to write, uh, in natural language, that will make much sense, uh, in Kubernetes where people know kind of the, the superficial layer of what they're asking. I mean, can I see errors from my workload that may translate down, down the line to a very complex query, collecting different reasons and different metrics describing the problems you wanna see and creating multiple panels or tables for you that will be very, very cost effective and very easy to do for Mars developers. So I think specifically querying data with, uh, AI or generative AI is gonna be something we're gonna see a lot of.
Um, and you know, the, the, the other section of correlating events and getting to root cause with AI and correlations around kind of AI that's been going on for a few years already in the market. And I assume that's gonna be, um, you know, getting more and more, uh, significant. But I think the major jump will be letting people query their data in a different way.
Will observability replace monitoring tools or is it just another layer above? Because there's still some value in those dashboards. There's some predefined metrics that are easy to understand, but you know, what is gonna be the relationship between what we're calling monitoring today and observability tomorrow?
I think that eventually the, the two kind of combine, 'cause in a sense, the part of the monitoring eventually translates to your day to day. I mean, you don't come, you don't wake up every morning refreshing your observability UI to kind of check out what's going on. You have to build, uh, a very rich alerting mechanisms to, you know, be alerted about stuff that goes on.
I mean, uh, the, the main, uh, serving of your observability stack is in these downtime hours where you would just wanna know what happened really quickly and kind of start from there. Once you kind of finish that monitoring layer, you know, these dashboards and alerts that kind of help you figure out if something deviated from normal, then you know, then observability kind of kicks in with, uh, correlating data and kind of helping you get to the root cause better. And, uh, correlating a few data sources in, I guess more native and complex screens that allows you to get the experience you want easily.
I mean, eventually, uh, debugging the huge Kubernetes system from logs to metric to traces. You can't just do it in dashboards, it's not enough. Uh, you can get the alert of something going on, but eventually you would have to see data in different aspects in dependency maps and in logs and in, uh, traces that kind of, uh, move through a distributed trace to figure out where to pinpoint stuff and, and metrics of latency and CPU usage and, and, and a lot of stuff like that.
So putting that all together, that's kind of what observability is to me. Uh, but it doesn't replace eventually the need to be dependent on a strong layer of, of alerting and dashboarding to kind of start all that. So what's your best advice to folks about how to get started with observability?
I think there's even been a few folks that may have stopped and started and they're kind of still experimenting. So how do I get over the hump? I think we're in a situation right now where a lot of people are migrating to Kubernetes.
Uh, so, uh, in, in some cases that could be a, a good trigger to kind of examine a different way to monitor stuff. And I think that when, when it comes to Kubernetes, CVPF is definitely interesting. There are a lot of ways to kind of experiment with this technology, uh, through open source tools, which are more ad hoc that you can play with.
Um, and there are vendors like ground cover that offer free tier that you can try out. I think that seeing what this technology can do is, is definitely an interesting experience to whoever didn't experience that specifically if you're in Kubernetes. 'cause the amplifying uh, effect there is much more strong.
So I recommend people to try out, if they have Kubernetes workloads to try out EVPF and see how you can monitor this stuff in a different way than what you're used to. Uh, how, how you can onboard into a full observability stack differently than what you imagine, you know, no more of these hard weeks of work, but kind of seeing stuff out of the bed immediately with just a simple sensor. And once you get a sense of that, I think it kind of, um, yeah, you get the hundred to experiment more and try out to figure out how you can use it to your needs.
I guess everybody has a different needs and you have to tailor, uh, the solution to your needs. But, uh, just seeing all this data, I think what we see with, with a lot of the organization kind of opens up their thinking about what is missing right now in their stack and how how can they get better. Um, and since it takes only a few minutes to experiment, I would just recommend try it with your own eyes on your own clusters, uh, and see how it looks like.
Alright folks you heard in here using legacy technologies to capture data for a brand new way of thinking about managing it, you may something of a disconnect. So take a look at the whole thing, end to end and figure out what your best approach is gonna be from there. Shahar, thanks for being on the show.
Thank you so much Mike. And back to you guys in the studio.