Focused on Monitoring Internet User Experience Introducing Catchpoint
Catchpoint, founded in 2008 by Mehdi Daoudi and two others, focuses monitoring on the user experience rather than solely on the application stack. Daoudi’s experience at DoubleClick, where he was responsible for resolving a major outage, solidified his belief that performance monitoring is not just an IT concern, but a critical business imperative impacting revenue, brand reputation, and employee morale. Catchpoint works with large global organizations, emphasizing the importance of trust and reliability in maintaining customer relationships and preventing costly outages.
The company advocates for an end-user-centric approach to monitoring, arguing that traditional infrastructure-up methods are insufficient in today’s complex, distributed environments. They highlight the inefficiency of waiting for massive error thresholds to trigger alerts, emphasizing the need to detect problems early, before they impact customers and cause brand damage. Catchpoint uses real-world examples of how their platform has helped companies identify and resolve performance issues much more quickly than using traditional methods, saving time, money, and reputational harm.
Daoudi stressed that many organizations overcomplicate their monitoring strategies, accumulating excessive data that hinders rather than helps problem solving. He advocates for a simpler, more focused approach where monitoring is designed to answer specific business questions related to key user actions and business outcomes. He draws on examples from Google, highlighting their performance-driven culture and the importance of designing for performance from the outset. Ultimately, Catchpoint aims to help businesses move beyond reactive troubleshooting to proactive, user-experience-driven monitoring to ensure that their digital services are consistently reliable and meet user expectations.
Presented by Mehdi Daoudi, CEO & Founder, Catchpoint. Recorded live in Santa Clara, California on February 19, 2025 as part of Cloud Field Day 22. Watch the entire presentation at https://techfieldday.com/appearance/selector-ai-presents-at-cloud-field-day-22/, https://techfieldday.com/event/cfd22/ or visit https://www.catchpoint.com/learn-more for more information.
Transcript
Good afternoon. Good, good morning, good evening, uh, everyone. So my name is Mary.
I'm the co-founder and CEO of Catchpoint. I'm joined here with Brandon. He will present later on to give us a demo of, uh, of the platform.
Um, so we're going to go over a brief intro of, of the, of the company, what we do, how, what problems do we solve, and, uh, literally what we want to talk about today is like how organizations must really rethink their monitoring observability, uh, within the companies, right? It's not just an IT thing, and this is the message is monitoring is just not an IT thing. It's a business thing.
It's a important for business. So, uh, again, Brandon and, and I will be your, uh, your, your presenters today. Um, so who is Catchpoint and what, what, what do we do for a living?
So at the end of the day, we've been around, uh, for 16 years. Uh, we started the company, uh, three co-founders and I in 2008. Great year to start a business.
Highly recommend was fantastic. And, uh, and, uh, we came out of Google, uh, the double click acquisition of Google. And, uh, my job there was with my team was to basically create the create build by deploy monitoring tools to deliver 40 billion ads a day.
At the time was a pretty big, uh, pretty big set of numbers to do 40 billion transactions a day. And, uh, performance was important. Ad serving was important.
Uh, how many of us have been perturbed by by how slow ads can, can literally damage, uh, website experience. Uh, actually I did take double click down in making a mistake in operations. And so I was responsible for an outage that lasted three hours.
And for the rest of my life, I like to joke that I've been, uh, sent to purgatory and I'll be doing monitoring as a punishment. So here I am. Uh, but I really love monitoring.
And the reason why is because what I saw firsthand is when we created a culture of performance, when we brought better monitoring, we stopped the finger pointing within it, engineering, qa, operations, network engineering. So with better data, we allowed the company to run better, and we, and we kicked their competitions behind, right? 2 billion.
So we, we were the better ad serving technology. So, uh, today we support about 400 customers globally, mostly major organizations, fortune 2000, 3000 companies, um, that literally have understood, or they have, they know firsthand that performance matters. They know firsthand that an outage, um, hurts, hurts, their bottom line, hurts their brand.
And I'll get to that in a second. We have about 300 employees worldwide as well. We're headquartered in New York, and we're fully remote since covid.
So, uh, why do people buy our solution? Or why do we engage with some of our customers? And at the end of the day, it's not an IT thing, it's not a a metric thing.
It's trust. It's about trust, trust, medi. What the hell are you talking about?
Trust? Well, at the end of the day, we buy from people we trust. You buy from Amazon because you trust Amazon.
You buy from Amazon because you know they're reliable. You know that your, your package is going to be arrive on time. So companies that believe in trust and, and having that relationship with their customers, then realize that they can't waste their p their customers time.
And so good monitoring, good availability, good performance, good reliability is good for revenue, it's good for brand. Nobody wants to see their brand trashed on Twitter or X or whatever, uh, saying, Hey, you know, this, uh, Southwest is not on time. Their website is slow.
So the, the world is filled with this example. So bad, bad reliability, bad resiliency is not good for the brand. And I have yet to meet a CIO that wakes up every day and says, today's going to be an awesome day.
Let's shoot for 50% availability. Right? That doesn't exist.
So these are the three pillars that we engage with our customers on on a daily basis. It's about improving their revenue, protecting and helping them increase their revenue, protecting their brand, because nobody, again likes to see their brand trashed on a social platform. And for it, uh, our, the teams we interact with, the site reliability engineers want to be part of an elite team.
They want to be working in companies that value reliability because people are tired of waking up at three o'clock every day, uh, and dealing with the crisis after crisis. That's not good for morale, the toil increase, et cetera. So this is the business value of Catchpoint.
Uh, so when I started my career in two thou in 1997 in monitoring, we used to look at monitoring from an infrastructure up, right? We monitor our data centers, our servers, the network, uh, the leftover was a little bit on the applications, and then the left, left leftover budget was spent on Gomez and Keynote, and whatever tools were available back then, we do end user monitoring. If I had to run the same team I run back, then I would focus on the end user first, because that's what matters, right?
So it's a, it's an end user centric approach to monitoring that we preach. Why? Because at the end of the day, if you want to protect your revenue, if you want to protect your brand, you have to monitor from where it matters and where it matters is the end user.
And so that lesson was in, was learned one day when I walked into the no back in my previous life, and everybody was sitting chilling. It's like, everything okay? It's like, yeah, absolutely everything is okay.
It's like, yeah, look, Mary, all the screens are green, the database is green, the networks green. We were not delivering ads, but from within those walls of our data centers, everything was a okay, except that my, our users were not living in the data center. So today they're not living in the cloud or inside AWS.
So monitoring from where it matters is important. Also, back in the days, people did not outsource willy-nilly their DNS to dine or, or no longer dine, but you know, NS one or Route 53, the same way nobody had multi CDNs because that was already expensive to have one. CDN today have customers that have seven CDNs around the world, or nobody is, nobody was putting stuff on edge compute.
So today it's about the end users, about the certain services, internet services that are important, but also the cloud providers and all the other things that you rely on. So it's an end user, uh, approach because cloud compute is giving you all the telemetry already on your infrastructure, right? So that stuff is already available, you just need to pump it into a graph or something else, and then you have it.
So again, our philosophy from trust to an end user based approach to monitoring is extremely important. So this is a true story. This is me visiting a customer or prospect.
They never became a customer. Uh, and my conversation with them translates into this slide. Hey, do you guys need Catchpoint?
No, we're fine. We, we, we have an a PM tool, again, is very important. I'm not saying one versus the other, but we have an a PM tool and we're totally fine.
Great. Oh, but by the way, medi, how do you solve this thing of like, Hey, today we have to wait for a hundred thousand API calls to fail, for us to declare an emergency. It's like, really?
You guys have to wait to see a hundred thousand errors in your logs to say that maybe we have a problem. Wow. So this was pre pandemic.
And the example I used vividly was, so hold on. So if you were running the hospital in Seattle, you have to wait for a hundred thousand people to die before you say, we have an emergency. I should have written a book right there.
But that, but anyway, so basically what happens is this, it's always a change. There is always a change. Something causes an outage.
It's usually a change. Either human led change or like the internet change or an ISP or a provider or your vendor made a change, right? Amazon updated something and suddenly your staff failed.
So then what happens is your customers are going to notice, then they're going to go on down detector, they're going to go on Twitter, we're going to start seeing sentiments that, hey, somebo, something is going on. Comcast is having issues on the East coast. What, what happened?
The loyalty start getting impacted. People start complaining, and then the voice of the customer start getting bubbling and, and everything. Then you finally reach the a hundred thousand API error threshold.
Oh, oopsie, we have a problem. Then you escalate, you create a crisis bridge. A bunch of people jump on the call.
Is it you? Is it you? Is it this?
Did you do this? Did you do that? Finally, we find the, the resolution, somebody maybe deploys a fix, and then the problem goes away.
This is how typically it works. Literally, every company we work with, this is the typical scenario. There is a change, something changed, something broke.
They have all kind of monitoring, but you have to wait for a threshold. You, you need to wait for something really, really big to happen, otherwise you have false positive, you have too much noise. Nobody wants to be alerted if it's only a hundred API alerts or whatever microservices does.
So where we come in is like, Hey, would it be better if you simulated the end users and detected that as soon as the first errors happened, you could have caught that problem right there and then you would, you would've saved that entire problem that entire time to having to detect the issue and whatnot. So this is why we come in, and our value proposition is about saving time. It's about detecting things faster sooner.
So you don't have to wait for things to get escalated to, to CEO of Delta saying, Hey, nobody can book a flight. Maybe you could have picked that up earlier, because what we, what we deal with is the unknown unknown. Most of the problems our customers see they have never seen before, and they don't know they're happening is their famous Rams Feld quadrant, if you guys remember that, uh, that part of, uh, our history.
So they are No, no, no, you, you, you don't know you have a problem. You're finding out on Twitter, you're finding out on down detector, you're finding out on some kind of social media platform that your systems are having a problem. And it's like, Hey guys, do we know what we're, what we're dealing with?
Is there a playbook, a runbook on this? Oops, no, this is the first time we don't know how to deal with. And again, it's about compression, right?
That meantime to repair is about shrinking that. I'll give you two examples. This is, these are true examples.
This is a famous sneaker company that shall remain nameless. We detected the problem. This is Catchpoint data.
So we detected the problem, it took them a month to solve, uh, real problem happening in the, in the APAC region was happening in, in New Zealand, Australia. 5 seconds to seven, eight seconds. Um, by the way, in our business, we're not in the bus.
I mean, nobody should be in the business of quadrupling their performance metrics. This is bad, right? This is not what people should be doing.
It should be going down, up is bad enough in our world. So then basically it's not the network. The connect time was fine.
It was the time to first bite, which is a metric that shows infrastructure related metric performance issues. Uh, it happened from AWS, it happened from tele try. It happened from Verizon.
So various networks. So it wasn't one particular ISP that was causing the issue. So you can triangulate and, and keep, like eliminating.
Is it the network? Is it the infrastructure? This, that, and because these guys were using Akamai, uh, A CDN, it is like we saw that across all of the Akamai ip.
So it wasn't Akamai per se, it was the origin server that was having a run. But you, you were saying about how your wife has, uh, doesn't understand the black box of the internet. This is the black box.
This is the stuff that goes behind, but the user shopping on this, uh, website doesn't care. They care about the great user experience. They care about you not wasting their time.
I'm here to book a flight. I'm here to buy a ticket. I'm here to order food.
I'm here to join a Zoom call. I'm here to do something. Don't waste my time.
You're supposed to make this black box work for me. That's the whole concept of the internet, right? So this is a one example, and the other example is this one.
Uh, this is the performance, uh, of a another famous company. And, uh, there is a pattern here, right? And this is reduc before ai, uh, the pattern is, hey, interesting, uh, something happened here, and then there is a cohort of metrics that stayed flat and another cohort of metrics that just boom jumped.
Um, so I know you we're not going to play games to try to figure out what caused this. So I'll give you the, the answer very quickly. Uh, two data centers.
Uh, one release, uh, one release went into a data center that had all the hardware, the older hardware had lower memory, uh, specs on the servers. And, uh, there was a, a, a Java framework, uh, really a version that was a little bit older, some memory leaks. And then as you can see here, the response time increased because there was a memory problem on those servers on the data center.
B, right? Again, it took from four 15 to 5 0 2, I mean, earth was created, quote unquote in seven days, right? I mean, we're talking about like 15 days of troubleshooting of people even not knowing there was a problem.
These are, by the way, billions of dollars impacted trillions of dollars of brand impact kind of thing, right? So what was in the slide before you said that there was quite a bit of lag between when it was detected and when they solved the issue. The human, this One Right here, the reason is the human element, literally, um, the customer is notified.
The employee is notified. There is this typical attitude towards, no, it's not a problem. We're not seeing it in other tools, right?
This, this, this I, I'll give you a personal example. So I dunno if I shared it last time. I was here, um, uh, a a few, uh, a few years ago in 2021, um, uh, a, a routine x-ray, my doctor detected a a dot on my lung, right?
Turned out to be cancer. Guess what my first reaction was? It's a faulty equipment.
Can we do a second thing? Literally, right? So this is a human element, a human thing that we have in our brain.
It's like discarded. It cannot be true. You, your stuff must be wrong.
Give me, gimme another data set. So we went back and forth for a month arguing that there is a problem. I was like, oh, yes, there is a problem.
It took a month. And then once the problem was detected again, origin servers that was badly configured, they had like five servers. This particular company had five servers behind the VI that was doing the origin to Akamai.
Out of those five servers, four of them were dead. One was available. This, Would this be the typical example that people culturally, people don't want to accept it or they're feel that they're the ones blame?
That's one. This an outlier. This is, this is not an outlier.
I wish it was an outlier. This is very common, unfortunately. And you, but yeah, go Ahead.
You would think that having four out of your five servers non-functional would have raised the alarm on some other monitoring solution, correct. They had, Correct. So again, the, uh, the, the interesting thing here, uh, and this is like a debate around like maybe just the human aspect to how we, we perceive monitoring and observability, we tend to overcomplicate things.
We want to go and buy the latest shiny thing or the, we want to deploy the latest technology and we don't look first at what we currently have. Are we monitoring the current environment? You know, instead of, Hey, let's add some open telemetry here and do this and do that.
Are we monitoring the, the infrastructure correctly? Do we have the right monitoring in place? Are we capturing the right telemetry versus like, oh, let's capture everything.
And then it becomes, so there is so much data that nobody can look at stuff. Mm. That's the other problem we see is too much data, and then it's already hard to find a needle in a haystack.
Now it becomes like a needle in haystacks of data. And it's just like so much people give up. It's like, I, I don't know.
Maybe I'll deal with another alert that might be easier to deal with. I wanna, I wanna add one thing. Yes, Sir.
Number One, thank you for the way that you've been setting this conversation up. Um, the, the point that I wanna make is you had a tagline, you buzzwords, that's RE Yes. And I'm gonna latch onto that.
Yes, sir. And then you talked about user experience and you talked about finding the right things that are important to the end user. And then you added, let's not capture everything.
So let's backtrack on traditional, I'm not even gonna say products. Yeah. Or network teams.
Lemme just say that. Yeah. System teams are just as guilty.
Let's capture everything the old traditional way you just said it. And then when we get to the user experience, we're like, what do you mean? All the data's there tell us what's wrong.
Correct. Rather than the business impact from an SLO and getting all the way down into sli. So thank You for this.
And Larry, I think what, what you, what you said is something important. One of the, one of my best customers that taught me a lot, I went to their, uh, I went to their, uh, in a meeting with their SRE teams. And these are people that, again, some of the most advanced, uh, companies in the world.
And uh, and they have this dashboard and said, for us, we keep it simple. And basically all the telemetry that we collect must answer a simple question. Can people download a movie?
That's question mark. Can people buy a a device? Can people purchase a song?
Can people search for a song in French versus a song in English? And then because these are the business problems that they are held accountable to by their SLOs, and then they go and fetch the telemetry and they figure out the monitoring that is needed to answer that question. Not let's capture gazillions of data.
Let's spend a hundred million dollars on storing data and, and logs and whatever. And then when stuff hits the fan, which is a technical term, then it's like, I don't know what to look for. Right?
So they have these dashboards. Seriously, it's like, and it has helped us and Brandon will show you some of those things, but they've been on, they've been helping us get, get you those. How can we answer a business question as quickly as possible?
And basically this, right? It's like if, if I'm Amazon, can somebody buy a book on Amazon? Yes.
No. What is the real user data telling us? What is the synthetic data telling us?
What, what is, what is the network telemetry telling us? You should be able to answer those questions, uh, from the business perspective. Well, and just to go even further, you can actually, you should be designing towards what will test.
Well, like the Google Gmail team is a famous story of Yes. As you filled in your username and they move to the password field, they preload your inbox Yes. On the server side.
Correct. So that when you type in your password super fast, it means it's already there. Yep.
But if you didn't do that, there'd be lag. Yeah. And sometimes watching the totality of the experience is when you would see where those gaps are that you wouldn't see unless you can see everything.
Absolutely. But that's also when teams are organized to build features right. Towards performance.
Yeah. Right. Google, I mean, is built around that.
Everything is performance driven. You performance is not an afterthought. Performance is a feature.
It's the most important feature monitoring at Google was, was a, it's as a code, right? You don't, you didn't think about monitoring after the fact. Actually, that was my biggest sticker shock after the Google acquisition of DoubleClick.
When, when the Google SRE team said, no, no, we, we don't have ops people and network people and this and that. You are supposed to do everything. And that was a learning experience because we had to become experts in, in SRE, which is you have to own the entire stack from A to Z Iron.
Is that like Google is paint log right? Fails its own core web vitals. They don't, can't ask core red vitals on their own platform.
It's a, it's a real problem. Um, so, uh, the type of customers we work with, uh, are here. These are just examples, but literally there is, we, we, we, we are very lucky to work with obviously all these companies, but they are really, they care about performance.
They care about their user experience, they care about the reliability. So.