Innovate or Perish IT Organizations Must Rethink Observability in the Cloud Age with Catchpoint
Catchpoint’s presentation at Cloud Field Day highlighted the critical dependence of modern organizations on digital experiences for business operations and employee productivity. The increasing frequency of outages and performance issues underscores the limitations of traditional application-centric monitoring. By shifting to proactive user experience monitoring, Catchpoint argues, organizations can prevent costly incidents, improve revenue generation, and boost employee satisfaction. This approach is vital in today’s complex digital landscape where users expect the flawless performance seen in companies like Google, Apple, and TikTok.
Mehdi Daoudi, CEO of Catchpoint, emphasized the growing complexity of modern applications, which often involve numerous third-party vendors, CDNs, and cloud providers, resulting in hundreds of requests per web page. This complexity necessitates a shift from reactive to proactive monitoring. Catchpoint’s solution offers an “internet-centric” approach, visualizing the entire internet stack and providing insights into the performance of every component, even those outside the direct control of the organization. This complements existing APM tools by providing an “outside-in” perspective, enabling faster identification and resolution of issues.
The presentation showcased how Catchpoint’s Sonar solution, combined with real-user and synthetic data, allows organizations to identify problems proactively, even those caused by external factors like fiber cuts. This data can empower manual or automated remediation efforts, for example, by dynamically rerouting traffic to alternative CDNs or cloud providers based on performance. Catchpoint also highlighted the growing trend toward AI-driven automation, using insights from monitoring data to create automated playbooks and runbooks for resolving performance issues, furthering the shift toward a proactive and autonomous IT management strategy.
Presented by Mehdi Daoudi, CEO & Founder, Catchpoint. Recorded live in Santa Clara, California on February 19, 2025 as part of Cloud Field Day 22. Watch the entire presentation at https://techfieldday.com/appearance/selector-ai-presents-at-cloud-field-day-22/, https://techfieldday.com/event/cfd22/ or visit https://www.catchpoint.com/learn-more for more information.
Transcript
So medic, co-founder and CEO of Catchpoint. Uh, and we're going to be talking about a little bit more details on, on, on this topic. So I don't think anybody refers to digital or analog anymore, right?
I think it's a given, right? Again, you booked a right to come here. You ordered some food.
I think I heard somebody was ordering, uh, Starbucks earlier. I don't know if it's coming and being delivered by Uber or somebody went to pick it up, but it was done over a phone, uh, which we're on Zoom. Uh, we're checking our emails.
I hope not everybody, right? We're not multitasking, paying attention, but I'm kidding. Uh, so, uh, submitting an expense report, Brendan, I saw him do that earlier.
And, uh, so it's all digital. And again, we are here to, things need to happen super fast. So performance is important.
The other thing that is, uh, really interesting is like our customers and our employees, uh, our are are influenced by some of companies that are amazing, like Slack, apple, of course, TikTok, Instagram. These are the companies that are setting the bar for performance, right? When you think about infinite scrolling of Pinterest, when Pinterest invented that a few years ago, I mean, you have no, you have no idea what stress it added on all the CDNs to be able to have that content continuously scroll and be available at any given time.
Uh, but that drove a lot of CDNs to do immediate Pershing, immediate load caching, et cetera. Um, Google, we talked about Google, but you know, Google set the bar when it comes to performance. I mean, blank page with a search book, kind of easy.
But then everybody compared themselves to that. Um, I have, uh, my kids that come home, they are capable today of telling me that dad, the wifi lags, like, where the hell did you learn that word? Not, uh, but Learned it from watching you dad.
Yeah, That's why I block YouTube. Um, but the expectations are rising everywhere, right? And it's stuck by some of these companies.
So it's not a question of if how outages and performance issues and things like that will happen. It's a question of when. Uh, but it's also a question of how bad it's going to be or how prepared our companies, how resilient our people, our companies, um, and who gets fired.
'cause at some point, somebody is responsible and accountable for, for outages, right? Um, if four out of five servers are done and nobody knew about it, I mean, that person wouldn't be working for me. I mean, maybe the second time.
Uh, anyway, so how outage will happen, we did an SRE survey. Uh, we do one every year. This was the last, uh, one, uh, for the seventh or eighth year in a row now.
And, uh, and this is, we asked, we asked, how many incidents have you responded to in the last 30 days? And this is insane, right? I mean, one to five, 40% of the people responded that they, they dealt with an incident in the last, uh, 30 days.
Now, an incident takes time. An incident makes you tired, an incident creates toil, an incident makes you fall behind some of your projects and some of the other things you're, you have to do. Or it might start, Hey, maybe it's time for me to update my LinkedIn resume or my resume because I'm done working in a company that has too many outages because I'm burned.
I'm, I'm tired. He agrees with me. So again, going back to the black box and, and, uh, uh, the modern applications are complex.
They're distributed. Uh, some of our customers use five seven CDNs. Some of them use five cloud vendors.
Um, their API, uh, is the, is finance of the API, as I call it, uh, multiple third party, uh, uh, vendors, uh, involved in delivering that user experience. That single webpage, that single mobile app that shows up instantaneously. There are hundreds of calls that happen.
Uh, and it's complex because in every one of them, there is a DNS lookup. There is a network happen, and it goes on and on and on. So this complexity is, is really explosive.
I think the average, when we started Catchpoint, there were maybe 30 requests on the page on average, according to web HGTP archive, there are over 200 HGTP requests on average on a single webpage these days. 200 HGTP connections, that's 200 connections, sometimes lookups, DNS lookups, IPV four, IPV six, the slowness of IPV six, et cetera, et cetera. So all these things, uh, uh, create complexity.
So we believe that the, the internet is your platform, and there is no more, well, it's not in my hands anymore. I can't control it, thank you very much. Right?
That doesn't, that doesn't exist. You're still responsible. Our customers are still responsible for delivering that user experience.
Uh, you can't go back to your CEO. I mean, I, I, I have, I, I can't, I can't see someone at Apple or whatever, say, eh, sorry, Tim. Uh, we're not responsible for the internet, so we can't deliver, uh, a web, uh, you know, web app somewhere, iCloud somewhere that doesn't work.
So the internet is your platform. It's complex, it's fragile. That's basically the key message.
And so we, we want to look at things in, um, in the way of internet stack. So what we started work looking at is like, uh, instead of thinking about monitoring the network or the, just the ISPs is like, how can we, how can we bring this concept of internet stack to our customers? How can we visualize it?
How can we layer it? How can we identify which part of this layer, uh, to the customer so they can triangulate and correlate as quickly as possible? So today, the tool allows customers to be able to do that, and Brandon will do a much better job at showing you that later.
But it starts with like the network monitoring, BGP monitoring, being able to detect what's happening on the internet, even the things that you don't control. We have this solution called sonar, which literally we, we harvest like all the billions of tests we do on a daily basis to basically be able to say, Hey, is there something going on? Or was there a fiber cut in Indonesia or in, in, in, in Egypt, in the Swiss canal that caused the problem on the internet?
Um, so this is the example I said earlier about the core banking system. And the, the key thing is for every request, this is an exa, this is a true example of this customer of ours that is in the financial industry when we mapped out their entire network and all the dependencies. How, how, how does a CIO an SRE leader or anybody will start this and say, okay, where do I start my monitoring?
How can I, where do I deploy my traditional monitoring? And then traditionally, if I go and put a, uh, an a PM solution or some kind of, uh, uh, tool like that, it's very limited what you can see, because you can't knock at the door of every one of your vendors say, Hey, can you install a an A PM collector? That's not going to happen.
Nobody's going to let you come in and install this, A PM collector, that a PM collector, et cetera, et cetera. So that doesn't work. So we want to be able to see it from the outside.
So the way we, we work with most of our customers have an a PM solution, and there are plenty great ones, Dynatrace, Datadog, new Relic ad, et cetera. So we coexist usually with, with those type of platforms. And, uh, and then we come in and we, we give the other vision of the, of, of how an application works, which is this internet centric, uh, from an internet perspective, right?
And we understand that outside in view, and we compliment and be able and, uh, and basically help customers see, uh, where their issues are. So when we think about monitoring, as I said, I've been doing this for a while, but in the nineties when I started my career, it was very reactive. Monitoring was very reactive.
We will, but it was also easy. You had the data center, you had your servers, you controlled everything. You controlled who had access to your data centers.
Uh, today, you don't. So then we became more proactive. We introduced the concept of observability.
Like, Hey, let's look at our logs. Let's look at our metrics. Let's look at synthetic data and whatnot, and let's put it all together.
But when we spend our time with our customers today, this internet centric, this automation, we talk about ai, it's, is really about being proactive, being intelligent, being autonomous, and also being intent based. But we have a lot of customers that literally have done amazing work in collecting all the data from all these monitoring vendors. They use and literally drive either, uh, automation or create AIOps model that say, Hey, if we see these five things, this is the playbook to run, or this is the right run book to run.
Right? How to again, shrink and make things faster. But the world where heading towards is automation.
Um, and, uh, and, and just like this whole intent based, I mean, Google is doing a lot of really good stuff on the GCP side with their intent based networking capability. Um, questions? You I have a question.
Yes. So, and we haven't really gotten into the meat of Yeah. How the Catchpoint solution.
Yes. Yeah. You did mention the fact that there have been internet outages before.
Yes. Due to fastly having a problem, or Akamai or one of those large scale providers, your solution would've certainly detected that was an issue. But would it provide any kind of path of remediation, um, so that you can avoid the resume generating event?
Yeah. Portion of things? Uh, no.
Um, so the, the, the actionability part is usually taken today by the customers themselves. So we have customers that, for example, have two three CDNs, for example, in, in the case of, uh, and if one of their CDNs is having a problem, they have now the ability to better signals, richer signals. You have confidence levels to say, I can take an action.
I can basically reroute the traffic from Akamai to Fastly or Fastly to CloudFlare or whatever, knowing that there isn't going to be an impact, because I'm continuously monitoring everybody, and I know that even if I shift to traffic somewhere else, I'm not going to send traffic into a d another black hole, right? So the automation is happening manually to the, when I say manually, it's like more in-home, in-house grown solutions rather than off the shelf solution. But, uh, I think this is where we're going to start seeing some really cool innovation, uh, where companies come in with automation flows when it comes to taking the monitoring data and driving automation.
Okay, that makes sense. So in a scenario where, say CloudFlare was having a, uh, a, not a full outage, but was having performance desegregation in Australia, For example, because we're picking on Australia recently, um, you could see that and then direct vastly to take over, correct? For that specific region, that would be a, a manual process you'd undertake if You'd have Well, it's automated to some degree.
So for example, we have a set of customers today that the combination of Catchpoint, NS one IBM bought NS one. Uh, so Catchpoint manipulating the DNS in real time, okay. And then saying, Hey, we're seeing real user data and, and synthetic data showing a problem in, in Australia for provider A shift the traffic to provider B.
So, but that has, that took about three, four years of confidence building mm-hmm. In the data. You can't just like, it's autopilot, right?
It's like, I dunno, how many of you have a Tesla or drive a car that has a auto autopilot, uh, it's a freaky thing in the first, the first time, right? It's, it's like, whoa, reassure on the 1 0 1. Do you really want to turn that on?
But then once you do it, you, you know, it's still sometimes a little bit weird, but it works. So, uh, that part is automated. So we've, we've done that automation, we have customer, we have like five or six of them doing that kind of, uh, high level automation today.
What we're seeing is like moving away from just the static content con piece, like CDS to saying, Hey, I have this, uh, I use now multi-cloud vendors. Um, I have, I have my, my con I have my application running on AWS on Azure on GCP, for example, and I have the shopping cart system, this microservice that is, uh, not doing so well on GCP. Hey, what is the Catchpoint data telling us?
Oh, it is bad. How is the, how is in real time that same service running on Azure and, uh, and, um, and AWS doing, or it's doing much better. So automatically do a DNS change, stop sending traffic there, send the traffic to, to only AWS and, and Azure.
So those, this, this is the kind of stuff we're seeing today, which is like, not only static CDN, but now application routing is happening that way. Hmm. That is the full part.
And I don't wanna preview what I think and I know is coming, but it's also, there's black box loss of understanding along the way because we don't understand what's going on in the network at Azure or whatever. You know, I said, I've got a literally this goofiest thing, it's called Cafe Bar now. I like cafe bars.
And so you go to Cafe Bar now and you press a button and it makes a cafe bar for you. Simple, right? Should be, there's nothing to it, but it's really slow because I have external calls to A CDN for my JavaScript library, an external call to my bootstrap library, and then I'm running on Heroku.
So I've got all these things and Heroku runs on top of Google Cloud, right? So Understand Everything is black box to most things, and I think this is what you're kind of going, it's like, what's the difference? What are you getting?
Because you're seeing data that's coming from the provider that Form. But we're also simulating, we're also continuously simulating, but we don't even have to wait sometimes for the data to come back. We're, because we're simulating that end user journey, we're seeing the data.
We can test every path from everywhere all the time. So that feedback mechanism, we don't have to wait for it. We're generating that feedback mechanism and then using it in real time.