The IPM Platform in Action with Catchpoint
Catchpoint’s presentation at Cloud Field Day showcased their Internet Performance Monitoring (IPM) platform through a live demo and user scenario walkthrough. The goal was to demonstrate how organizations can proactively address digital disruptions and outages impacting both customers and employees. The core of the demo centered on a visual representation called the “Internet Stack Map,” which provides a detailed, real-time view of the entire internet stack for a given application—in this case, a movie ticket booking app—allowing users to see the performance of each layer, from DNS resolution and CDNs to backend services and databases. This visualization instantly highlights performance bottlenecks and potential points of failure.
The presentation explained how the Stack Map is constructed using a combination of user-defined synthetic monitors and Catchpoint’s auto-discovery feature. Users can define tests at various layers of the network stack (BGP, ping, trace route, etc.) and Catchpoint intelligently identifies and maps third-party services involved in the application’s operation. The platform also integrates with APM tools like Dynatrace and Datadog, allowing seamless transitions between the high-level overview provided by Catchpoint and the deeper code-level insights offered by APMs, providing a comprehensive view of application performance. This integration facilitates a collaborative approach to troubleshooting, enabling teams to quickly identify and resolve issues.
Finally, Catchpoint highlighted its Internet Sonar map, a proactive monitoring tool that tracks the performance of key internet services and infrastructure, providing advance warning of potential outages affecting an organization’s application. By monitoring a broad range of services and identifying correlated events, Catchpoint helps organizations anticipate and mitigate the impact of external disruptions. The company emphasized its commitment to customer success, offering dedicated support and value engineering services to ensure quick adoption and ongoing value from their platform, including proactive alerts and detailed incident reporting.
Presented by Brandon DeLap, Senior Solution Engineer, Catchpoint and Mehdi Daoudi, CEO & Founder, Catchpoint. Recorded live in Santa Clara, California on February 19, 2025 as part of Cloud Field Day 22. Watch the entire presentation at https://techfieldday.com/appearance/selector-ai-presents-at-cloud-field-day-22/, https://techfieldday.com/event/cfd22/ or visit https://www.catchpoint.com/guided-product-tourfor more information.
Transcript
My name is Brendan Dela, senior Solutions Engineer at Catchpoint. Uh, what I'll be doing for the next 33 minutes or so is walking through the Catchpoint portal here. Feel free to ask questions for the folks in the room here.
Would love to make it as interactive as possible when I'm clicking through. Um, I can literally go anywhere through the portal, um, whenever you want me to. So feel free to to stop me if you'd like to.
Um, but where I would like to start, um, is this Internet Stack Map visualization that you see on my screen here today. What we're looking at is actually the, the inner workings in the entire internet stack for a demo, movie ticket, booking app. Um, so think of it, you know, Fandango or, um, a CM and just going on and booking a movie ticket, right?
A, a Long Jersey or user flow, maybe it's, uh, you know, logging in, uh, finding a movie, you know, picking that right seat. You know, a lot of microservices, a lot of backend database calls occurring throughout that user transaction. However, as Medi mentioned, the user doesn't know, doesn't care, or doesn't understand any of that.
Um, so that's where Catchpoint comes into play and helps out SREs, DevOps, essentially anyone responsible for observability of their platform. Yep. How did we get here?
Uh, like how did we get all this stuff on the screen? Is an agent, are you just pointing to end points? What, what does that look like?
Yeah, so how do we get here? Um, so I can jump forward there. There's, there's a couple of components here that goes into this stack map, and, and the first component would be, well, let's monitor what matters.
So that's where you come into the control center here, and you have the ability to set up synthetic monitors, which are essentially just, you know, emulated, um, tests, um, different protocol types. If we wanna start at the, the lowest level here, so BGP for example, if you do manage your own networks, or if you wanna keep on your, keep an eye on your cloud provider, you can drop in the prefixes for those. Then we get into the more network side of things, right?
Can a user actually reach that service or this, uh, movie ticketing app from a trace route or a ping perspective, whether it's T-C-P-I-C-M-P or UDP or quick nowadays, was that created Manually that map? Or Did you Yeah, like how do we like up? Yeah, no worries.
So you, first you set up all these tests, and then within Catchpoint, we do have a sense of what we call auto discovery. So you just drop in a domain that you're testing, we'll go ahead and look through all the tests that we have or that you have configured in Catchpoint, and then we'll spit out some of these known third party services that either we're already monitoring for you and have a sense of, all right, this request is hosted by this provider, or, you know, this IP links back to Akamai, for example. So we're doing all that behind the scenes and then pulling out the test data that is coming from those tests that you have.
So in this case, we can see, uh, obviously this is the primary domain of a user coming to this application. We then have a DNS layer with several different DNF providers, each one of them having different service time, different availability metrics. We get into the front end later here, in this case, we're using CloudFront.
This is where we see a little more latency and some downtime and packet loss. There are different third party, uh, tracking pixels, different third party APIs that may be occurring on the initial, um, page load there, like double click as, as medi mentioned, you know, Optimizely tag Manager, right? How are those services impacting the overall experience?
And then we get into, you know, more of the, uh, potentially backend cloud infrastructure type layers. So origin or API gateway, right? Are you using AWS, are you leveraging their API gateway?
And then you can actually go into the more backend services. Are you using Lambda for your microservices? Is that where we're seeing impact of, uh, or increase of service time or latency, potentially even a third party authentication service, like auth zero?
And actually what we're seeing here in the stack map is we are seeing an issue with, with auth, auth zero, and you can follow that path all the way from where the user hits your application to the front end within AWS to auth zero, and then finally to my, my SQL database. So the stack map highlights that at this point in time, we are looking at the last three hours of data. You can go back in time, um, and then we'll also notify you and lets, you know, via alerts you may have set up.
So you can quickly drill into, uh, recent errors for that example. Um, so there is some auto discovery, but if there's obviously customization needed, some of these could be internal custom services that we obviously don't know about. So you can come in there and create those custom services as well.
Right? So stuff like the My SQL database, that's not necessarily discoverable by crawling the website, right? Correct.
Yep. That's a detail I would have to add, Correct? Yep.
Yep, exactly. And then some of these for, you know, those microservices as well, right? Uh, you can come in here and say, well, this is my seat service.
It's based off of this domain or this URL, any test that is essentially calling out to that, pull that into this specific service in this layer on the stack map. Okay. So We can see.
Does any of your integrations with the A PM tooling help with building out the internal infrastructure behind this? I'm sorry, can you repeat The question? Does any of your integrations with the AP PM tooling help with building out the internal components of this stack map?
So with the integrations with like an A PM tool, it'd be more like pushing the data to the A PM tool. Okay. Um, we do support, um, as medi mentioned, tracing and open telemetry.
And that will be coming here in the stack map. So we'll actually be introducing that, and that'll obviously automatically pull out, all right, what is that service that's called? What is the endpoint?
Right? Are there any, you know, exceptions or failures from that server perspective? To answer your question, we have integrations with, uh, um, companies like ServiceNow, where companies that have like that single source of truth being serviced now feeding automatically some of this stuff and automatically creating the test in Catchpoint.
Okay. And then that will surface into Stack Map. Yep.
One, the one thing I'm seeing here, and I'm guessing that the topology of the stack map will also then alter based on the time period you're looking at. So you are actually seeing how it was at that time, and then you can really see the diff, you know? Exactly.
Exactly. And then when issues do occur, right? Obviously highlighting and pinpointing, um, at a glance, you know, what could be going wrong at this point in time, in in which we realize that it's an issue with my MySQL database, it's caused by the authentication service, and obviously my front end users are seeing that when they're not able to, uh, to log in.
When you've got an, let's say you've got a Dynatrace or you know, you've got something else that's different sources, is there a way that you can kind of go from here? Is there any integration to those third parties where you can say like, kind of open this app in my A PM? Absolutely.
Yep. So we have deep integrations with all, most of them, where literally from here you can click and link directly to a trace in di inside Dynatrace or inside, uh, Datadog or whatnot. Yeah.
And we have some customers that have publicly, they went public. And so SAP for example, uh, big, big Catchpoint, Dynatrace integration, flawless integration, where it's like literally click and then land directly inside Dynatrace and be able to explore. So seamless for the, the user.
Yeah. Yeah, exactly. And then through Catchpoint here with, uh, open Telemetry, you can actually do the same.
So I'm just looking at some recent tracing errors here. So client aboard exception, you know, 4 0 1 errors, SQL exception. You can then come in here and see, all right, what are those traces?
Again, these could be synthetic or real user. Um, so not relying on your real users to, to have or cause that exception, you can actually use your synthetic tests and then your one click into, uh, a trace chart. For that example, You can act, act as an o upstream o hotel collector for any o hotel native inbound.
Exactly. So for example, here we can see, all right, this specific trace map, right? What failures occurred.
We have a couple of four oh ones, and then at the end of the day, database not responding for this specific trace. Not bad. Excellent.
So hopefully that answers the question of how this was put together, um, through your existing synthetic tests that you created. We have auto discovery, and then you can also come in here and, you know, customize your layers, customize your services, uh, according to you and, and internally how you see your, your internet stack. And People have multiple stack maps, right?
Yeah. So this is not just one, and you're done. Like, we have customers that have, like, for example, a shopping stack map, a, uh, uh, delivery, uh, a warehouse stack map, right?
So, uh, different services can have their own map own because different teams are in charge of different functions. What, what kind of visibility and capability inside things like versal, uh, Heroku, uh, and rep and such, where that SaaS layer's kind of obscured, but yeah, there's obviously something that could affect it, Right? So one of the best integrations we have done just because of relationship with GCPI would say is the Google, the, the Google Cloud integrations.
We've built custom just because of a lot of mutual customers. And, you know, sometimes we just have to follow where the customers are. And, and so we have a lot more customers, uh, on GCP that are, that have pushed us to do some really cool innovation where we can go really deep in GCP and bubble the data in, in Catchpoint.
Right. And vice versa, If they're using GCP, they almost are qualifying themselves as performance oriented. Mm-hmm.
Like they are, they're after the science Delivery all things off. Yeah. Is there an overarching map?
If I didn't create a stack map, is there an overarching one that's like a catchall where I'm, I'm not sure where I'm having the problem, but it shows me, here's the problem, and then I can like, yeah, We have default dashboards. I mean, think about Stack Map as a, as a, as an Uber, Like a zoom in onto it or something, which Yeah, Like, yeah. Yeah.
There's several different ways to visualize the data. Um, but so for example, you can look at geo maps, right? What I'm looking at here is actually synthetic and real user data side by side.
You can see the correlation here, right? With whether it's emulated or then seeing that from the real user perspective. So it all depends on do you need to see a geo map?
What, what specific metrics is metrics or services do you wanna keep an eye on? Um, so that's where you would kind of get your catchall when it comes to, you know, coming in at first and, and setting up your, your initial, uh, visualizations. I'm assuming you have some default dashboards already set up everything.
Yeah, Definitely. If I don't have the dashboard set up and I just come in there, how the heck do I know what I'm looking at? That, that would be the, the one.
But, uh, one of the things that we pride ourselves is with our, our service. So meaning that, again, the, we are software as a service. The word service at the ends mean that there is a very big white glove, not that it's complex or whatever, but we truly believe in having that aha moment as quickly as possible with our customers.
So you can have it through tools, but if, uh, but I, we also believe in making sure that hey, let's accompany them and get them to that aha moment as quickly as possible. So we have dedicated what we call value engineers, so they spend time with our customers post-sales in making sure this is part of the service, uh, where we spend time making sure that they get the value out of the product. And you help setting up all those tests and all those services and Yep, Yep.
Exactly. And some people said, leave us alone. We, we've used the product and or we know what we do.
Oh, that's fine. Yeah. Do you see a lot of mixed environments?
The first thing I think of is even if I'm running Dynatrace, if I've got a thousand applications only, but a hundred are actually gonna be instrumented by Dynatrace. Hmm. The rest are gonna be some homegrown sort of thing.
Uh, where like, this seems like you could map really well to those, because you can take it all in. What are the interactions between mixed APMs? Like, do you ever get duplicate data?
How do you manage where there's a lot of information that could be driving, you know, two APMs, seeing the underlay? How do you reconcile what's what? So about 90% of our customers have an a PM solution of sort, uh, the big ones of course, uh, uh, we see a lot of Dynatrace.
We see a lot of Datadog, right? So these are the two major ones we see day in, day out. Uh, then there is New Relic, and then there is APPLI kind of thing, but there's also a lot of new stuff coming in, right?
O open telemetry, um, Grafana, there are a lot of new things as well. Uh, the, the, it's not so much the overlaying, it's more like, like in the case of SAP for example, and you know, the, the gentleman was, uh, at, uh, the Dynatrace conference the other day talking about what they've done is actually they leveraging every tool to its maximum potential of cutting that meantime to repair, right? So Catchpoint is very good at telling you where the smoke is before people start screaming there is a fire, right?
So being able to detect smoke is very important, and they use us for that. Then it is like, okay, how do we take that data and put it into their own AIOps tools to then look at, at what Dynatrace is saying to say, Hey, Catchpoint is seeing smoke. What are you seeing here?
Oh, look, we're starting to see database latency or, you know, error crashes or some stuff like that. And then bubble that stuff up. So it's not so much like the, the, the union of the two products, but it's like, okay, I'm getting signal here.
Let me validate it somewhere else and keep on going until the triangulation example I give. And the, the most important use case there is like, what if your servers that aren't actually up and running, right? That's where the external synthetic will come into play and be like, Hey, there's an issue.
You may not be seeing this. You may not be actually seeing traces or events on your APM solution because no one can actually reach your application. So that's also where we come in and kind of, you know, can we be that, that we can find that smoke at the end of the day when it, before it becomes a huge fire, um, segwaying over into, so I kind of showed you like how we monitor what matters, right?
So it starts with breaking it down into different emulations, right? Different types of tests. You're not just looking at a 200, okay, and saying, all right, well my users can now book a ticket.
'cause unfortunately they can't, right? There's more than, uh, just reaching the website to actually, you know, check out for that, that movie ticket. Um, so taking that over into, uh, monitor from where it matters.
So Medi touched on this a little bit. We are close to 3000, um, intelligent agents for you to, to leverage from a global perspective here. Um, and how we do that is we break it down into different network types, getting closer to where your end users actually are.
Um, typically what you'll see today are, you know, providers allowing you to monitor from the cloud, whether that's AWS, Azure, you know, GCP, now that, that checks the box when it comes to, you know, availability monitoring, right? You can check from next door or potentially, you know, a couple server racks, um, down the row if you are also hosted from within that cloud provider. But it's really important, um, to move a little closer to where your end users are connected to.
So that's when we refer to, you know, the core of the internet. So backbone providers, providers like level three, um, N-T-T-A-T-T cogent. Um, and then we also get a little closer to the end users, so more consumer fiber broadband connections, uh, whether that's Spectrum, Comcast, Verizon, Fios, we do have starlink as well.
Um, so that's a pretty interesting use case, right? Um, T-Mobile's actually starting to offer that for free for everyone, so everyone can test that out right now. And, you know, if you are delivering services to those customers, you need to check to see whether or not they can book that movie ticket, um, from that specific provider.
Uh, we do support wireless as well. So 4G 5G obviously IPV six infrastructure here, a completely different, um, you know, problem or set of problems to deal with from a networking perspective. So we give you the ability to, to essentially mystery shop from those providers as well.
Um, and then last but not least, we also do a lot to test internally. So there are applications that may not be accessible from the external or from external users. They may just be internal users.
Um, so the ability to drop a Catchpoint agent into your, your office location, your data center, um, a cool use case there is actually monitor from within your data center and then monitor from our backbone, and you can actually see the delta there. And that's essentially the delta that you're trying to lower or decrease so that your users have a better experience. Um, so that's another use case that you can leverage there.
And we're getting a little smaller here where you can actually install on raspberry pies plugging into a router. Um, I have one actually sitting at home on my desk where we're running a quick net quick network tests, uh, every minute or so to Google, just making sure my internet home is accessible for, for my wife and I. So, um, just another use case there from, from a Do you have a learning and monitoring on that too?
So you can call up catch point support and say it's exactly, I wanna make sure I get alerted before my wife alerted Yeah. Or your kid, right? But one of, one of the use cases of the, what Brandon was talking about on the, on the Catchpoint minis or the, the lightweight agent is call centers.
Yeah. Uh, uh, um, shopping, uh, systems like points of sales, Factories, Factories, like we have a, uh, company that makes cars. And so behind every robot there is a Catchpoint note that literally test the API call to see, can I, can I get this?
Because, or checking even the NTP, right? The time protocol, because if things get out of sync, you get the exactly the wrong device or the wrong part in the wrong part of the cars, which is not good. Um, but What about remote locations or like secure locations?
Yes. Isolation, yes. Yeah.
So we have that deployed in some secure location. You wanna just close? I'm just kidding, just kidding.
Micro did not work here. Right? And, and why this is important, I'm gonna show you another dashboard here.
Um, this just goes to show you the difference between monitoring from the cloud and from the backbone and last mile. So the top left or the top, um, map that you see here, these are, you know, average response time and availability from cloud locations like AWS, Azure GCP. So for the most part, right?
Everything looks pretty green, especially in North America, Europe, within China here is where you start to see some issues, right? Great. Well, firewall China, uh, causes to, tends to do that to, to some, um, applications and services.
Now switching over to more of the end user point of view, right? We can see where, oh, wait, in Phoenix it's actually taking 12 seconds for this booking app to load compared to cloud where it took 50 milliseconds or so, right? So there's a huge discrepancy there between what your end users are actually seeing and, and what potentially, you know, your internal monitoring solutions are, are telling you at the end of the day.
So monitoring, we talk a lot about false positives when you have a signal telling you, you know, it's like that alarm in your house that beeps at two o'clock in the morning, right? But there is no fire, it's just a false positive. The battery is off, right?
That's what we talk a lot about is false positives. The danger is false. Negatives is you not knowing you have a problem, and then assuming that everything is okay and you go about your day and everything is okay until at the end of the day, you look at the cash register, it's like, Hey, how come nobody bought anything on our store today and nobody could get to it?
Well, how come I didn't know, oh, we were not monitoring or we didn't pay attention, or we were monitoring from the wrong location. So what, what what Brandon is showing is that false sense of safety and security that some companies have. It's like, oh, I am doing margining, but I'm monitoring from the cloud because my mom doesn't connect to the cloud.
She, her ISP is not Amazon. So not knowing is very important. So the false positives, but the dangers that false negativity not knowing that you have a problem.
And what sort of like, so when you're monitoring from this many different locations, yeah. What are some of the things that you set up with companies when they're getting started to actually track what's important to them in terms of understanding false positives? Yeah, what's a real problem?
I think if, if we start with what Medi started with, like what is, what's that main service of your business, right? Are you, do people buy cars from you? Do people buy tickets from you?
Right? So typically where we start is that transaction is probably the most important, right? So at, at the end of the day, I want to end up scripting that user flow using playwright or puppeteer Selenium to ensure that, you know, every minute, every five minutes, every 15 minutes from these points of presence, our user can, can complete that task.
And then from there we start peeling that, that onion, and we say, all right, well what goes into that? Obviously there's a DNS resolution, there's an HCTB object, um, that you need to get, it needs to return to 200 K. You know, SSL is also important, making sure that, you know, your certificate's not expired or there's no performance, uh, degradation because of your SSL provider.
Um, so it's, that's how we go into it and we, we start picking apart their application or their service and set up dedicated monitors, uh, for them. Yeah. And then obviously, you know, where are your users coming from?
If you don't have users in China, obviously don't monitor from China, they probably probably can't reach it anyways. So, um, it, those are the kind of scenarios that we walk through with customers. And then the last piece is, that piece is obviously get though to get the data or the alerts in the right hands as quickly as possible.
So where do we need to route these, these issues to? Is it if it's a DNS issue, right? Do you have an SRE team or do you have a team dedicated just for, for the DNS side of things?
Um, and then obviously developers as well. If it's a JavaScript issue, pass that along as quickly as possible to, to the dev team. Yep.
Gotcha. I think you said it there, sorry, but how are the, how are the tests written? What tooling?
Yeah, So if we look at the user flow kind of scenario, um, playwright, puppeteer, and then, um, selenium is, is what we started at, but that's obviously it's going away in the industry, so we're moving over towards the playwright and the puppeteer. Oops. Yep.
And then the other test is just drop A URL in, drop a domain, drop an ip. Yep. And we'll just simply run a trace or DNS resolution to that.
Yep. And is that something that your team helped with in writing those tests if needed? Yes.
Yeah. Yep. Definitely.
So my team specifically, uh, the solution engineering team, we help with getting those initial set, uh, tests set up, um, during A POC or an onboarding phase. And then as medi mentioned as well, we have dedicated value engineers and customer success teams afterwards that are with you throughout the entire journey. 'cause you know, what you have configured today for your application might be different, you know, a couple of days from now.
Or maybe internet stack map is telling you, Hey, you actually have a dependency you might not know about. Let's go ahead and set up a test against that as well. Something that's slightly adjacent, but maybe relevant.
Do you guys do any performance and load testing services? Right. Okay.
Yeah. No. Okay.
We do a great job of letting you like, you know, if you are doing load testing, Catchpoint will tell you, Hey, there's, there's an issue starting to creep up here. I was just thinking some, some of those tests in those frameworks of relevant when you're trying to run through load testing and performance testing. Right?
So we Have a lot of customers in the e-commerce, uh, vertical. And so they spend a lot of time, we, we have special programs during the holiday. We have this thing called Black Friday monitoring, where we offer like some really dedicated high, high-end monitoring capabilities with a full wide glove service.
And usually we start working with those customers in the summer, uh, with their, while they're doing their load testing, they're using us in parallel to say, okay, did we detect the right things? There is another group that we work with is the security team, because we are usually the first one to see the smoke. So many times we're the first one to see, we, we are the first one to see a DNS hijack, a BGP hijack, a DDoS attack.
Mm-hmm. Because we're testing so frequently these, these properties and this systems that the first one to say, Hey, there is something abnormal. There is, there is a slide that, uh, that I didn't share, but, uh, this very large company, we detected DDoS attacks before any of their systems were able to catch anything.
Because typically what an attacker or somebody doing a DDoS does is they send these Trojan things to test you, right? They test your perimeter, so they're going to do a little burst, see how you react, did you move the left foot, the right foot, et cetera. Then they basically test your, your, your defense systems, right?
And then, oh, they didn't react, so let me try another time. Oh, they didn't and that's when they come at you, right? So we have some customers, uh, again, the data is fed into these data lakes, and so they fine tune their stuff to say, Hey, we're starting to treat the Catchpoint data as a security signal as well.
So if DNS goes up X amount of time, blah, blah, blah, then we have all have the security team take a look at it because maybe it's a, it's a DDoS in about to happen kind of thing. Uh, SSL Brandon said the, the, the challenge with SSL is, uh, we have this customer, the reason we did, we built SSL monitoring into the product is I, I went to see this customer, this guy was in the room and his job was literally to click on 40,000. There was a spreadsheet for, this is true story guys, but 40,000 domains that he had to click on every day to check, oh my God, make sure that SSL was working.
And I, I, I, I got, so, I had so much PTS that we have to help this. So we added SSL monitoring into the tool because again, many times people forget the, the basic functionality. And that's SSL, right?
And oh, by the way, when you have an SSL monitor, we have an SSL, you don't have it just for yourself. You give it to your cloud fair, you give it to your fast, you give it to your Apigee, you give it to, and you have a whole daisy chain of SSL now that is in the wild west. And if one of them expires, it takes you two days to get that back online.
Yeah. Imagine a big Fortune 500 that relies on s on e-commerce, for example, being done for two days. That's impossible.
So we, so we added that and then somebody said, Hey, it would be really cool if you could alert us if our domain is going to expire like seven days before. Mm-hmm. And by the way, the data shows that the right time to alert for s SSL is about two and a half days, not two weeks, because people, ah, I'll get, I'll get to it two weeks from now, congrat Later, But it's literally between two and a half and three days.
That's when people get that sense of, That sounds like me. Yeah. I think Eric can use that on his 48 domains.
Yeah. Yeah. Meanwhile, that guy no longer has a job.
Yeah. I was gonna say 40,000 URL guy. Yeah.
Hopefully He has a better, just a better job. So do you actually see a lot of real like internal intra service and intra ops like MTLS? 'cause almost every time I go into place, they've got all the stuff to be ready for MTLS, but never get it interservice.
They just kinda like trust the boundary because like as a performance monitor internally, that's also one of the problems. Like just the overhead of actually encapsulating the traffic from end to end. They're not ready for it.
So they usually back off. No, most of the comp, most of the customers we work with have that nailed down as well. The internal monitoring external monitor, they see it as one.
Yeah. There is no, this is my, there is no field. It's like, this is my territory.
This is so SSL monitoring internally, externally, the DNS, the internal DNS systems, because when you think about like running containers, right? There's a lot of DNS that is involved internally. You guys have the info blocks earlier, right?
But we do a lot of work where companies have deployed, for example, Infoblox or other IPM tools internally for, for their, for their DNS, uh, network connectivity between data centers, offices, multi-cloud vendors. Uh, so again, the sophistication with these customers is like they, they understood that everything is connected. It's about when there is smoke, just you don't know where it's coming from and you can't ignore it, right?
So they've learned that the hard way. So if I'm, uh, if I'm the receiver of, uh, an alert SRE team or whoever it might be mm-hmm. Am I typically gonna come to Catchpoint first, have a look from that point of view, and then if needed, go through the a PM tooling from there, if we've got it deployed to dig deeper into the issue if it hasn't been surfaced already, is that the kind of normal workflow that I might do from Yeah, it Definitely depends on the maturity and, and the scale of, of the, the team and the organization that we're working with.
But the, the ideal point is you get an alert from Catchpoint and it tells you what the problem is, right? What, what's failing? Is it DNS?
Is it network? Is it from a specific location? And then you route that to the, the team, right?
There's no middleman having to jump in and, and and click around in a portal to tell you what's wrong, right? The alerts are, uh, you know, sophisticated enough to tell you, Hey, problem X and region Y, and then you can actually build in a runbook in the alert to say, you know, send this through this channel or through this team. Okay.
Yeah. But there's definitely, um, you know, some organizations where, you know, they get in the tool, maybe they get an alert about packet loss or destination failures. You open up a waterfall and it shows you, oh, okay, traffic stops at this pier on the network, right?
Let me go reach out to them and see if we can fix that, that issue that's going on between the, those two peers. Okay. Um, the last thing here that I wanna leave you with would be, um, our internet sonar map.
Um, so Medi touched on this a little bit. This is, um, Catchpoint essentially doing the monitoring for you for a lot of these critical third party services out there. Whether it's A-C-D-N-A-D-N-S, um, you know, a SaaS solution, um, e-commerce cloud infrastructure, an ISP.
So we're, we're leveraging that global agent infrastructure that we have and all of these different monitor types that we have. And we're testing different applications, whether it's alter DNS, you know, specific ass, uh, weather service. Um, here I have an example of a, a generic DNS, um, cloud example.
You can then click into that incident that we're detecting and we'll tell you what potential tests are your, of yours are actually affected by this, or are there no tests affected by it? Um, where are these issues being seen? So we can see there's 10 different cities, uh, throughout the United States here where we're seeing the issue.
What are the domains impacted? So if this is a CDN or an ISP issue, you'll see all of your domains and potentially your competitor's domains or some other industries domains that are also impacted by that outage. And then we'll also go into, are there any correlated events to this specific incident?
So is it a, you know, Azure West US issue that's then causing an issue for Azure front door? Uh, we'll be able to correlate that for you as well. Um, and you're, you have the ability to come in here and, you know, filter by specific region specific type of service.
You can actually say, show me my stack map and only apply, you know, the current state for my services on that stack map. So we can see NS one and generic cloud having issues, but all these other incident or services do not have incidents at this point in time. So the best way to think about this is like, uh, in the good old days when you had the network operation center, you had also two TV stations running at any given time.
You had like a weather type of station, and then you had either CNN or whatever your, your flavor is, right? Because people wanted to also understand what was happening geopolitically, weather-wise, all the things that could impact. I remember Monica Lewinski for those that were young enough or old enough to remember, but we, we literally had the peak of traffic at DoubleClick.
We didn't understand what was going on, but all the websites were on fire. And then our, we were pumping ads, we run out of ad inventory that day, right? We literally did not have enough ads to show, and we were trying to understand what the heck was going on.
Well, because there was the president admitting something. And then basically the news, all the news organizations went, went ballistic. So having that, this is that CNN, that Fox, that whatever TV station you look at in your No, to try to understand what else is happening on the internet, that could explain why I'm seeing certain, How often a day do you expect or want somebody to be in your platform?
Like, it's one of those odd things where if you're doing things really well, they shouldn't be in there much. But how do you retain the like understanding and commitment of that team with like, 'cause you're feeding so much information capability, how do you make sure you see? So We, we have about 50,000 people on the platform, uh, registered, right?
Uh, so U users, I think at any given day we have like about 10,000 people that log in. Mm-hmm. And then the rest is all APIs, right?
You asked about alerts. The majority of our customers consume this data via APIs. So I, I have a hard time tracking, but they're about 50,000.
And then with saml, it's even harder to to track how many people, because for privacy reasons, sometimes we don't even see them in, in the, in the thing. They log through their SAML authentication, especially European customers. Um, but, uh, but the majority of folks is API consumption data is consumed somewhere else.
I guess that's on the other side too. You, APIs are now the new websites, right? Like Correct.
Correct. So you have to monitor your API, correct? Yes.
Like I would wanna monitor my API as much. Yeah. And, and, uh, our API consumption is through the roof.
Uh, we had to, we are constantly beefing that up just because, I mean, we have customers that are pumping millions of, of, of, uh, of API requests a day, uh, consuming, uh, you know, the sonar and all the other data. And with the other thing that we have that we, we have, I wouldn't say visibility into, but one of our data stream is what we call the fi the actual fire hose, where our notes are sending data, streaming data in real time to our customers. But the data doesn't even come to us.
It goes directly to the cus I mean, it comes to us as well, but they're consuming it in real time. They get the data at the same time as we get it from each node as soon as the test runs, and they're acting on that data. So the people that have built automation, they didn't build automation on like, let's get the, the Catchpoint data from their APIs?
No, they, they built automation. They're getting it directly from the source. Exactly.
But you're doing a dual pipeline to them. Correct. Exactly.
And that's where the, I saw the most successful automation because they're acting at the source. That's the level of Trust, correct? Well, it's built over years.
Yeah. The average, the average customer with Catchpoint is seven years. Right.
So we have customers that have been us with us for eight years, nine years, 10 years. And a lot of them, it took them three, four years to get to the level of comfort to say, I'm going to build automation based on their data. Yeah.
How old is the company? 16 years. Okay.
Wow.