The Internet Performance Mamagement (IPM) Platform with Catchpoint
Catchpoint’s Internet Performance Monitoring (IPM) platform proactively ensures the resilience and performance of digital experiences. The platform’s core function is to reduce mean time to repair (MTTR) by enabling detection, identification, escalation, and validation of issues. This is achieved through a multi-faceted approach that goes beyond simply identifying symptoms, such as a slow website, to pinpoint the root cause through triangulation of data from various sources.
Central to Catchpoint’s approach is its use of “mystery shoppers,” or synthetic monitoring agents, deployed globally across various ISPs and even in residential locations to provide a comprehensive view of performance from the end-user perspective. This is complemented by real user monitoring (RUM) using JavaScript beacons, BGP data collection for network-level insights, and integrations with tools like WebPageTest.org and Tundra for deeper application performance analysis. All this data feeds into a central platform for dashboards, alerting, and analysis.
The platform provides a holistic view of digital service health, focusing on accessibility, uptime, speed, and reliability. Key features include the StackMap visualization for a single-pane-of-glass view of performance, Internet Sonar for AI-driven issue identification, and end-to-end tracing using open telemetry. Catchpoint also offers advanced features like automated experiments to help web developers optimize performance and integrates with various alerting and ticketing systems for seamless workflow integration. The platform’s flexibility allows for customization to meet the specific needs of diverse organizations, from simple alert routing to sophisticated data integration with observability platforms.
Presented by Mehdi Daoudi, CEO & Founder, Catchpoint. Recorded live in Santa Clara, California on February 19, 2025 as part of Cloud Field Day 22. Watch the entire presentation at https://techfieldday.com/appearance/selector-ai-presents-at-cloud-field-day-22/, https://techfieldday.com/event/cfd22/ or visit https://www.catchpoint.com/guided-product-tourfor more information.
Transcript
Uh, so we're going to talk a little bit now about the Catchpoint platform. Uh, my name is Mary, I'm the co-founder and CEO of Catchpoint. I'm joined here by Brandon, who, uh, right after this is going to, to give us a cool demo of the product.
So we've talked a lot about who we are, what we do. This is just, again, a quick reminder. This is about reducing meantime to repair.
Like most of our customers and the practitioners we talk to, wake up every day thinking about how do I shrink this equation meantime to repair? And it starts with detect, identify, escalate, fix and validate. Catchpoint has not help you with the fixing, it's not yet.
But our job is to help you detect, help you identify, escalate to the right team, and then once you deploy the fix, sometimes it's also make sure that there is nothing left, right? So again, I want, but it's about protecting revenue, brand, customer experience, employee experience, morale, and innovation. Of course.
And so before we get into the how we do things, it's very important that in the, in the world of monitoring, uh, our job is to peel the onion. Whenever there is a problem, you all we see is the symptom. Ah, the website is slow.
That's the symptom. Oh, the application is not working. That's the symptom.
The root cause can be deep. And so our job is to, how do we identify that? So every time we, they said, I stole this from a customer of mine, actually.
So Ted, thank you very much. If you're watching and listening, uh, so we triangulate, we need to triangulate to find the root cause. So how we do it, so before we get into the, the, the, the, the guts of Catchpoint.
So what we believe in is we believe in a stereo vision in whatever we do. So we believe in synthetic monitoring or the term synthetic monitoring, which is the ability to create, uh, to have robots. Basically.
We tend to be a user to do a DNS request to do API request to do web request to basically have a robot, a a mystery shopper, right? Imagine, imagine if this hotel had mystery shoppers, they would've known that my toilet wasn't working right? Versus me coming in in the morning and saying, Hey, the toilet is not working.
So mystery shopping allows you to stand proactively to every room in a hotel to make sure is the bathroom clean, is the bed clean? And that's what Catchpoint does at, at the core, right? Is we are, we have, we have mystery shoppers, we have these agents that are everywhere around the world connected to every ISP around the world where most of the world and our job is to basically pretend to be a user doing something to catch problem.
We do that from the last mile as well, from people's homes. So we deploy a Catchpoint agent at people's houses, we pay them and we host that, uh, there. So we get that last mile perspective.
We have real users, uh, by putting a Java script beacon or little agent on my laptop that collects that data. We have BGP collectors that understand the internet. We understand who hijacked which BGP to send it to Pakistan right in the middle of the day.
Or is Russia sucking some of the traffic from an ISP in the US in the middle of the elections? Ah, that seems suspicious, right? Maybe they shouldn't be doing that.
org, which was the open source that is for the majority of the use cases, and that's to allow web developers to also understand their performance profiles anyway. And then we also acquired another company in Turkey called Tundra, uh, three years ago, that allows us to have open telemetry data and, and be able to dig a little bit deeper into why an application is bad. So as you can see here, we're starting to get into the whole end to end, uh, of what we want to do.
All of that gets into our, our, our, our, our core, what we call our core, which is in, in, in one of the greatest data data centers in the US in Las Vegas, uh, with the lot of redundancy and whatnot. And this is where catch our customers have access to all their dashboards, alerting, et cetera. And at the end of the, all of this comes out a bunch of things, a lot of data, a lot of telemetry, but it's really about being able to understand these four critical things.
Can I get to you? Are you up or down? Are you slow, fast?
And how reliable are you at delivering that user experience? So can I get to you? Can I medi from here in Santa Clara?
Get to data center, A, B, C. So my, is my network up once I get to you? Are you up?
'cause getting to you is one thing, but then maybe your website is down, right? Or your application is down. So are you up or down?
And if you're up, are you up at one second? Where you up at 25 seconds per page? Right?
That's not acceptable. So the speed and then the reliability is the ability to consistently deliver that user experience 24 7, not only at two o'clock in the morning when nobody's using your website, because you never know. So our platform has a bunch of components.
So it's one platform, but we think about the customer, we think about the workforce, we think about the application, the network and the website. And so different customers buy different things, but the goal is to be able to give them flexibility in what they want to require. Um, the most important thing is this is where we take the measurements from.
So we operate a fairly large infrastructure, uh, around 15,000 servers around the world today and growing. Uh, we keep adding more and more like last year alone, we added 168 new, uh, mystery shoppers. So agents as we call them, um, and, uh, these are data centers.
And this is not just cloud AWS whatever this is physical servers that we, we, we basically put around the world connected to the right ISPs because again, the, our job is to say, Hey, company A, B, C, you're having problem in Phila in, in Philadelphia on Comcast, or, and by the way, it's because AWS is having a network connectivity issue with Comcast. This is where the issue is, go and fix it. And so the output of all of that.
So we collect a lot of data, but the output is being able to help you understand how is your, how is the health of your digital service, right? Uh, this is a new thing actually. Last year when we came here, we presented version one of Stack Map.
Uh, it, it looked store 1990 compared to what Brenda's going to show you today. So we, we got it out. Customers gave us a ton of feedback.
It was unbelievable. And, uh, we, we, we changed, we tweaked it. But the, the concept is very simple, is I go in one place and I see my entire health.
I see my entire body, I see my entire scan. I know where, where the my heartbeat is. I know how much my blood levels are, I know where there is a blockage.
I know which, which third party is having a problem. And I can know exactly who to call to say, Hey, Akamai is having an issue, or CloudFlare is having an issue. Or, or actually we did switch to Akamai and performance is great.
Look. 'cause it's not always about negativity. It can be positive changes can be good.
Again, our customers, the, the IT ops, the SREs, the network engineers, being able to give them in one view this what we call the customer experience score, but also being able to understand all the elements, how they're connected to each other. Is it the network where the network, which ISP, was it packet loss or latency? Or maybe it was a BGP change that happened right there and there.
We can give all those answers. The other thing that we, last year when we came here, we had the prototype version of Internet Sonar that was released also after actually cloud field day. And um, and it has gotten a lot better, of course.
So literally here what we are applying machine learning, AI, whatever, uh, buzzword. So to avoid buzzwords today, but it's, it's machine learning. And so we're harnessing the entire data set that we collect plus some other third party data sources.
And we're able to say, yes, that performance issue in Australia, it is a real thing, but there is nothing you can do. You can go back to bed or at least you can notify your customers that, hey, we're experiencing an outage, we're experiencing slowness because VO in Australia is having an issue. There is nothing you can do about it.
Sometimes it's not like you can route or you can go and buy new ISP, uh, to, to those people in Australia or whatever, right? So there sometimes there is nothing you can do about it, but knowing is super important, having that level of understanding. So you can go tell the boss, get off my case.
There is nothing I can do about it. Right? My CEO at DoubleClick would say, Medi, go and fix the internet.
Literally true story. He would call me on Saturday. It's like the ads are slow.
It's like, Kevin, level three is having a problem. It's like, go fix it. Okay, sounds good.
Uh, at least you know now, and this is the level of, of sophistication we've gotten by, uh, being able to understand and monitor and, and, and map the network to be able to say, yes, there is a network outage with this ISP. And there is, here is exactly where, again, that laser focused thing is lifesaving because time, right? We have SREs that are very expensive, network operations, et cetera.
All these teams are very expensive. Uh, and you're better off having them do more interesting things. Um, very quickly on the tracing.
So we acquired this company, as I mentioned, Tundra. We, we spent the last year and a half integrating it into Catchpoint. Uh, we invested heavily on open telemetry, which I think is something that the world is moving to.
Uh, and uh, I'm super happy to see that because now we can end to end trace all the way from the end users, both synthetic and rum, all the way to where the customer's applications are and basically pinpoint exactly where the issues are. Nothing new here in the sense that this other companies have been doing that, but we're going at it from the end user first rather than from a server application perspective out, right? So flipping the coin, uh, on its head, uh, of course webpage test, huge fan.
Uh, the gold standard of web performance testing keeps getting better at Catchpoint. Uh, and we fully integrated and one of the things we did with AI is what we call experiments. So today, if I'm a web developer and, uh, somebody says, Hey, my website is slow, I have to go and figure out why it's slow.
And then I say, okay, maybe I have to try this. Did it work release? Try it.
No, it didn't work. Try the second thing with experiments, what we call experiments is the tool automatically tells you what are the 5, 6, 7, 10 things we found that could be bad? And you can literally add them to your cart, like shopping cart and say, can you run an experiment on these things?
So we, we launched these cloud edge workers on CloudFlare that basically go and experiment and tell you actually these three things, we're going to improve your performance. These are the two things. Don't waste your time.
There is no Yeah, I'm not a developer. Yeah, don't pretend to be. So the me either webpage test thing is new to me.
Okay. Um, that's not something that you developed a Catchpoint, correct? Correct.
So, uh, webpage test was developed by a good friend of mine, Patrick, uh, Menen. He was at a OL back in the days built webpage test run out of his garage. And then in 2000 I've been, I've been trying to get him to partner with us forever, but a few, about three, four years ago, we, we reached an agreement to acquire webpage test from him, and we kept webpage test org, which is still running in, in, in public.
It's free. And then there is a commercial version that is integrated in catch. Okay.
And what does that very briefly, what what does that include enable you to do? Can you just like point it at a website and be like, correct, go explore this website and then create synthetic tests? com while we were sitting here.
org, you put in a webpage and it gives you a complete report Yep. And it tells you what to fix what you look at. It's, it's an awesome tool.
Okay. org. I think a generation of people literally started their career, uh, on that, on that tool.
Okay. Where does the handoff, and it might change to different places, the handoff between what you identify and when you escalate or when you make the, uh, your customers aware. So when do they, when do you say, okay, hey, here's a, uh, right, an issue.
And also just be curious, when you talk about escalating to teams, I have to believe it varies which team is responsible and how much of that has sort of handcrafted figure out who it is, or you've already got that pre-wired. Yeah. If it's this team, we send it this way.
If it's that team, So in the tool, there is a lot of pre-wiring that happens, and then customers customize a lot of their stuff. So some people apply labels and based on the label, it, it stands one template of alerts to these guys. Some people will prefer their Slack channels, some others Microsoft teams, some side pager duty alerts, some others is just like old text messages.
We're developing a solution right now for a customer that wants us to integrate with Twilio to make phone calls and wake them up. So it really depends. The, the platform is fairly flexible to, to allow basically customization and allow you to talk to whatever team.
But a lot of these companies have already established some of those books rules, if you want. So between, like for example, uh, this networks and, you know, not if it's a network issue, notify the network engineering team, if it's everything else is the SRE, if it's an application performance issue, it's the application team. So companies that have already defined that we just plug through their workflows, to be honest, Or the platform, like you said, is intuitive enough that they can take the Absolutely.
So, so the, the most sophisticated customers we have, what they do is they suck all the data into a data lake of sort, right? So they come in, they either pull the data from us or we push the data to them. So we offer both, uh, part of the, the service.
And um, and, uh, one thing that we are seeing today a lot more is actually a lot more customers are pulling the, are putting a lot, pulling a lot of the data from observability systems into the Databricks of the world and snowflakes. Because what they're trying to do now is connect the dots at a whole different level mm-hmm. Business and IT metrics together kind of thing.
But, uh, yes, in general, it's just like the ability to open the fire hose, have the fire hose data or the fire hose of alerts go somewhere. Uh, we have customers that use, for example, Moogsoft, right? Uh, so a bunch of AIOps solutions that literally take care of the triaging or ServiceNow, um, where you open tickets automatically, right?
So some people have gotten to the, the very sophisticated, they say, okay, if it's this alert, if it's this threshold, blah, blah, blah, for example, open a ticket in ServiceNow to this team, and it's all templateized and everything. Thank you. Yeah, of course.
Um, of course real user monitoring. This is, this RO is cyclical. We, we've seen in the last 16 years.
Uh, so we want ro we don't want ro it's too much data. It's not enough data, but basically the concept of real user, uh, is instead of synthetic or you have robots that are actively going and checking things, you put a Java script, uh, on the webpage, and the Java script is capable of literally listening in to every single user interaction, being able to re record, record the session, uh, to be able to understand the funnel of the users. So it's like Google analytics kind of thing, but for SREs and DevOps, right?
So it's a little bit richer in terms of signal and data and uh, of course being able to understand frustration. Like one of the latest thing we've added was what we call rage clicks. Like how many are you, are you fighting with your mouse?
Are you track pad, right? So being able to understand that and correlating that to slowness, right? Oh, look, there is a rage click.
And it's like, oh, it's because of things gotten slow. And then we are taking rum now to the next level, but we've been working with this great customer of ours, and this is what I love. Part of my job is like literally sitting with some of these customers solving problems that other customers care about solving maybe two, three years from now.
But we've been working for the nine, nine, the last nine months with this, uh, um, company. And, uh, literally we're, we're building, we've built actually a, a mobile SDK real user monitoring solution for their mobile app, iOS and, and Android, but with open telemetry in mind. But this is the first, and we've been pushing the boundaries of open telemetry on the ROM side with them, with this particular customer.
And we're going live with this. It went live last week in the la la latest release of Catchpoint. And then I think I'm coming to an end here soon.
But like, even even on the monitoring side, we've partnered with some companies to do what is called ECN congestion monitoring. This is a very obscure protocol that was done in this 80, I think, on the network side, but suddenly it's getting a lot of popularity about like literally pushing the boundaries of, of TCP IP and just making things faster. Of course, a lot of people are switching to quick, uh, protocol because that gets rid of the overhead of TCP.
So being able to understand the, the network latency there. Um, yeah. So this is what we do.
We, we ensure resiliency. We ensure that you have all the data, you have all the metrics, you have all the, the, the signals, because this is a business that is not our business, but monitoring observability is data rich, signal rich. Uh, and the, the richer the signal is, the more you can act on it, you can learn on, you can learn and act on it.
So.