Christoph Pfister, Kentik | KubeCon + CloudNativeCon Europe 2023
What is the role of network observability when deploying Kubernetes applications in a hybrid-cloud environment? Learn how tracking and mapping network interconnections can help troubleshoot K8s deployments faster, ensure that network policy is acting as designed, and secure against threats.
Transcript
This is texturing TV. Hello and welcome back to kubecon plus Cloud nativecon Europe and we're here with Christoph Pfister who's from kentig and we're talking about Network observability. And you may very well ask.
What is that? So let's ask the man here. Oh, that's a an easy opener, isn't it?
So, thank you for having me. First of all. And so we're an observability company, but we do things a little bit differently because we go.
Deeper on the network than you know, most anybody at least I know looking around the show. And so we take in all kinds of network Telemetry flow VPC flow SNMP, of course streaming Telemetry and we stitch all of this bgp, and we stitch all of this together to give our customers a very deep view inside their networks and when I say networks, it's really hybrid networks. So it includes the cloud stuff kubernetes, which we're gonna probably gonna talk about in a little bit as well as they're on-prem environments.
So, how do I collect that Telemetry? Do I drop an agent everywhere or is there some way to streamline this or what what it started? It's very simple actually so obviously flow net flow is a very pervasive thing and you know, you just tell your router to send netflow and they send it to us and we ingest it and then, you know make the magic happen.
On the cloud side. It's VPC flow a little bit the same principle. One of the things that we do very definitely is that we enrich all this flow so we add additional.
Metadata around geolocation around things like, you know VPC VPC names, you know container names things like that. So you get a much more complete picture of your environment then if you just would take you know, the traditional flow that makes sense. So you're here at a cloud native show is networking in the kubernetes context different than traditional networking.
Are there different things to track and observer is a very similar. Well, it's very different in that. Of course, you know kubernetes is a very ephemeral environment, right?
And so you can't just look at your notes as an example and you know think that you'll get something out of that. By the way, it's a very highly Network intensive type environment, you know, I've heard a few talks here at the show that people talk about IP exhaustion as an example. And so you know, what we do in the kubernetes environment is we have an agent based on eBPF.
Which is the latest and greatest and so we basically have that agent deployed as a demon said which then extracts flow from the kubernetes environment enriches that data same thing with container names and stuff like that. And so any level some very unique use cases that again. I haven't seen here at the show and maybe you know, you can let me know if you have but so for example what our customers are interested in is of course, you know policy type information compliance type information, but also the Telemetry so in the policy side as one example, we allow our customers to Basically figure out who their pods or their containers are actually talking to not just within let's say an AWS environment, you know eks whatnot, but also out to the internet and so, you know, one of the questions we ask our customers to draw them into the booth is like would you like to know better your pots talk to an embargo country?
And that's one of the things that people go. Oh really? You can do that and you know the uniqueness of candidates that because we have all this Telemetry including, you know bgp information from the internet.
We allow customers to really not just look into into their clusters and into kubernetes, but what's going on among clusters and Beyond, you know into the into the wild basically I'm observable any implies some ability to query that data. So how do I kind of get at it or ask some questions or interrogated? Because and frankly what should I be asking that now because when I talked a lot of people about any observability solution, they're like this stuff is great, but I have no idea what to do with it.
Okay. So we talk about two things one is guided exploration. And the other one is open exploration.
So guided exploration is we help our customers with dashboards with maps with kind of the traditional, you know set of visualizations. If you want to call it that way, you know time series, you know stuff alerts and so on but then we also offer what we call open exploration. This is a concept that we have that is pretty unique.
It's called the data Explorer and it basically allows customers to query the data in real time. And basically ask any question which is you know, is there any egress in? in Frankfurt that I don't want or how much is the cost of my you know AWS egress and so we allow a pretty free form and open-ended capability to basically ask these questions.
Now what's behind disability is this very scalable back-end that we have so versat solution. We have our own database which we run and it's a massively scalable back-end basically, so to give you one data point we last year ingested about 270 trillion. Records flow records over over the years.
So that's about six million plus flows per second. And so You know one of the unique this unique pieces of kentig is that because we have this massively scalable back end we can take in very high cardinality data like netflow. And so, you know, if you look around and look at some of the other observability solutions, they are, you know, maxing out pretty quickly if you look at the networking stuff.
And so that's why we think some of the other companies are not going as deep into the network as as we can. So massively scalable sasses part of the secret sauce. They saying that network is the ultimate source of the truth.
So how far off the stack can you see from where you are and into applications and what kind of behaviors going on up at that level? Yeah, so we do so, let me maybe take a step back so part of the reason. We're successful is that you know, if you look at the traditional APM observability players, you know, if it's not the app, you know, if the traces and the metrics show, it's not the app.
What do they do? Well, they go to their networking team say must be the network. It's always the case.
Skype 100 memes around that stuff. But and so then you know, what? What do you do?
Well, you look at you know some of the data and the analytics we provide. And we tie that and mashed it together with application context. So for example based on the traffic we see in an environment we can say well you're doing 50 gigs of Netflix in addition to you know, all your Business traffic and so we can tie Ott application data back to the flows because it's enriched we can look at you know, custom applications based on port numbers and protocols.
And so we allow our networking teams to have the application context in addition to just looking at well routers and switches and you know CPU and that kind of stuff because you're right application context is super important for figuring out what's really going on number one, but number two, what are the applications that are actually impacted by, you know, a network outage or slow down or latency or whatnot. So ultimately do you think network operations will be folded into more of the devops kind of workflows or is it still going to be the land of a specialist? You know, our customers are still pretty siled, especially in the Enterprise side.
And so we have netops teams. We have devops teams. And you know, they talk with each other but really they're using different tooling and I think part of the reason is that the devops tools today.
They can't go as deep into network and infrastructure as the net work tooling like kentik can do and so it's important to you know, give both teams the right context. And so the context from our end is that we allow them to look at kind of the applications that might be impacted by networking issue. And so that's how they you know build the bridge is now You know one of our customers actually said it best the other day when I visited with them, which said you know, there are so very multi-cloud.
Type capability. What does that mean? Well, we look in deeply into AWS and Azure and gcp and so on and you know, there's not like workloads that kind of span these clouds but there's always like a business unit that says, well, I want to go with AWS and then another business unit or an acquisition is I'm gonna go with Azure and so they have these different clouds that are running in a company but the networking team really has to take care of all of them.
And so the quote from the customers, you know, we have to keep up with them developers. And this is a you know, networking Guy saying that which we allow them to look into and Azure and data IBS ngcp and some other clouds to help them figure out what's going on. All right, folks, there's an old joke about what's the one thing a server admin and a developer can agree on the answer is it's the network guys fault turns out though.
It's not always the case. Hey Kristoff. Thanks for being a thank you so much.





