Navigating the Observability Landscape: Logs, Metrics, and Traces Demystified at Cloud Native Now 2024
Transcript
Hi, everyone. I'm super excited to be presenting at the, uh, cloud, uh, native, uh, now conference. Um, uh, 2024.
Um, today we are gonna explore a critical aspect of modern cloud infrastructure, uh, which is observability. Uh, as systems grow and more complex, uh, and distributed, uh, understanding the health and performance of our applications, uh, becomes increasingly, uh, vital. Uh, so observability in the cloud and encompasses are three main pillars, uh, logs, metrics, and traces.
Um, so in this presentation, we'll, uh, demystify these, uh, components, uh, showing how they, uh, work to provide a comprehensive view of, um, the, uh, systems health, uh, discuss best practices, uh, tools, and our real world examples, uh, to help you navigate the, uh, observability, uh, landscape, uh, effectively. Uh, so a little bit about me. Um, my name is, uh, Neha, and I am a technical program manager at LinkedIn.
Um, so here at LinkedIn, I've been part of the team, uh, that, uh, develops and maintains, uh, critical observability infrastructure, uh, including logging metrics and, uh, tracing solutions. Um, my mission here is to ensure that our systems are, are robust, uh, scalable, uh, and they provide an, uh, excellent user, uh, experience, uh, by leveraging advanced observability practices. Um, so in today's, uh, agenda will be, uh, going over, uh, some of the, uh, guiding principles, uh, why observability is, uh, critical for us here at LinkedIn are some of the, uh, goals.
And I will, uh, deep dive into the, uh, logs, metrics and our tracing. Um, so our, um, main, um, mission here is to enable, um, LinkedIn, uh, to effectively monitor and, uh, optimize the, uh, cloud systems using, uh, logs, metrics, traces, and also at all times ensuring there is high performance and reliability. Um, our vision is to ensure that, uh, LinkedIn can, uh, seamlessly navigate, uh, its observability landscape, uh, achieving, uh, deeper, uh, insights and, uh, proactively, um, issue, uh, resolve, uh, any, uh, issues, uh, in the, uh, cloud in infrastructure.
Um, so our core, uh, values, uh, we promote open and, uh, clear communication regarding the, uh, system health and performance. Uh, we foster a culture of visibility where data-driven decisions are made. Uh, we highly encourage a proactive monitoring and, uh, issue resolution, uh, to prevent downtime and inefficiencies.
Uh, we also, uh, emphasize the importance of, um, anticipating and addressing potential, uh, problems, uh, before they impact our, uh, customers and, uh, end users, um, the advocate for a strong collaboration between the, uh, development operations and business teams. And, uh, we finally, uh, support a unified approach, uh, to observability, uh, that, uh, spans across all orgs, uh, within LinkedIn. Um, so why is observability, uh, so critical for us?
And, uh, why does it play, uh, such a big role in, uh, LinkedIn's operation? Um, so one, uh, it ensures optimal, uh, user experience. Um, observability, uh, ensures that, uh, LinkedIn remains, uh, fast, reliable and available.
Uh, so it provides a seamless experience, uh, for millions. Um, actually, uh, we've hit 1 billion users. Um, uh, so it also provides continuous monitoring, um, that helps detect and our resolve issues, uh, quickly, uh, preventing, uh, any negative impact on, uh, user satisfaction.
Uh, second, uh, it, uh, it's, uh, proactive in, uh, issue, uh, resolution. Uh, so observability, uh, allows for early detection of anomalies and potential problems often before they become very critical. So, detailed logs, metrics and traces, uh, enable a rapid diagnosis and, uh, resolution, uh, and it minimizes our downtime and, uh, maintains, uh, platform reliability.
And finally, uh, performance optimization. So, uh, metrics and traces help pinpoint, uh, performance bottlenecks, sorry, uh, performance bottlenecks, um, enabling our targeted, uh, optimizations. Um, so observability, uh, data reveals inefficiencies in our resource utilization, um, allowing LinkedIn to optimize infrastructure usage and our reduces costs.
So, some of our goals, uh, in respect to observability at LinkedIn, um, are that we ensure high performance and our reliability. Um, at LinkedIn, we maintain a seamless and our responsive user experience, uh, for the billions of, uh, users, uh, worldwide. Uh, we approach this by, uh, utilizing metrics and traces to monitor the, uh, system performance, uh, identifying bottlenecks and, uh, optimize our resource usage, uh, to ensure there's high availability at all times.
Um, we are proactive in the, uh, issue, uh, detection. Um, so, um, the way we approaches approach this is by implementing a comprehensive logging and monitoring system to detect the, uh, anomalies and patterns, um, enabling us with our response and mitigation times. Um, so our approach to enhance, uh, security and compliance is that the, uh, monitor, uh, logs for security events and anomalies.
Uh, we maintain audit, uh, trails and ensure adherence to secure security policies and, uh, compliance, uh, standards. Uh, we also facilitate, uh, efficient, uh, incident management and, uh, clear deep dives. Um, so here, uh, we minimize the, uh, downtime and reduce the impact of incidents.
So we approach this by, uh, leveraging the, uh, detailed logs that we have, the, uh, metrics traces, uh, to, uh, quickly, uh, diagnose these, uh, issues and, uh, perform root cause analysis to, uh, prevent them from occurring in the future. Um, we also approach, uh, um, uh, resource utilization, uh, by using the observability data to analyze, uh, resource usage patterns, uh, scale the, uh, infrastructure appropriately, uh, and perform, um, uh, any, uh, cost efficiency, uh, related, uh, targets, uh, for supporting, uh, continuous improvement. Uh, we, uh, gather and analyze observability data to, uh, inform decisions on, uh, system improvements, uh, any feature enhancements, and also, uh, use this data, uh, for performance optimizations.
Um, so let's, uh, dive into, uh, the observability, uh, in distributed systems. Um, so what is a distributed system? Um, a distributed system is a collection of, uh, independent, uh, computers that work together, uh, to appear, um, as a single coherent system, uh, to end users.
So, uh, these computers, uh, communicate and, uh, coordinate their actions by, uh, passing messages to one another over a network. So the, uh, key, uh, characteristics, uh, here include, um, that they have, um, they have multiple, uh, components and, uh, network communication. Uh, so distributed systems are composed of multiple autonomous components.
Um, they're often referred to as, uh, nodes or services, and each component performs a specific, uh, function and can operate independently, but, uh, they work together to achieve a common goal. Uh, for example, uh, in, uh, microservices architecture, uh, different services handle, uh, different aspects of an application, uh, like, uh, user authentication, uh, payment processing and inventory, uh, management. So in network communication, the, uh, components, uh, in a distributed system, uh, communicate and coordinate their actions by exchanging messages over a network.
Um, so this communication can happen over local networks, uh, wide area networks, or even the, uh, intranet. Uh, the, uh, network plays a crucial, uh, role in ensuring that the, uh, components can share, uh, data and work together efficiently. So, some of the challenges here, uh, with, um, you know, distributed system, uh, is that, um, you know, there's complexity around managing and coordinating multiple components.
Um, the developers and operators need to handle issues like network latency, uh, data inconsistency, and fall tolerance. Um, so ensuring that all of these components, um, need to have like a consistent, uh, view of the system's state can be challenging, um, especially when dealing with distributed, uh, databases, uh, or, uh, transactions. Um, so with security as well, um, you know, a distributed system, uh, involves, uh, protecting, uh, communication channels, uh, ensuring authentication and authorization and, uh, safeguarding, um, data across, uh, multiple, uh, nodes.
So these are some of the, uh, complexities, uh, within our distributed systems. Um, so, um, to, uh, dive a little, uh, deeper into the, uh, key concepts or key concept of this, uh, presentation, uh, which is, uh, logged metrics and, uh, tracing. Um, so the underlying, uh, technology is Kubernetes here.
Um, so Kubernetes is a complex, uh, system, uh, due to its distributed and dynamic nature, um, managing a large number of containers and, uh, services across multiple nodes, uh, which makes it difficult to, uh, manually, uh, monitor and manage. So this is where observability, uh, comes in by, uh, providing insight into the, uh, health, uh, performance and behavior of the application and, uh, clusters. So what is observability by, uh, definition?
Uh, observability is about understanding the, uh, internal, uh, state, uh, of a system, uh, by examining the, uh, data it, uh, produces. So think of it, uh, like a doctor, uh, diagnosing a patient by looking at their symptoms. Um, observability allows us to infer what's happening inside the system, uh, from the outside.
Um, so although observability and, uh, monitoring are, uh, often mentioned together, uh, they serve different purposes. So, um, monitoring, uh, involves, uh, collecting and analyzing, uh, data and metrics, uh, from the Kubernetes cluster to ensure it's performing as expected. Um, so this can include, um, things like, uh, checking, uh, CPU usage, uh, memory usage, and, uh, network traffic of the cluster.
And, uh, it's components such as, uh, pods and nodes. Uh, for observability, uh, we can say that it is, uh, built on top of monitoring, uh, observability, uh, includes, um, monitoring, but, uh, it actually goes beyond it to analyze, uh, logs, uh, traces, and, uh, other techniques, uh, that can provide insights into the, uh, behavior of the, uh, application as a whole. So, the, uh, four, uh, pillars of observability are, uh, logs, metrics, uh, traces, and, uh, the newly added is a profiling.
And, uh, let's look into, uh, what these exactly mean. Um, so logs, um, by, uh, definition refer to the, uh, logging as the, uh, practice of, uh, collecting and storing data about events and activities that occur, uh, within, uh, your cluster, uh, in, uh, in order to monitor and diagnose issues, uh, with applications and infrastructure. Uh, in Kubernetes, uh, we use, uh, logging to keep, uh, track of what's happening, uh, in the, uh, cluster, um, or within our application.
Uh, for example, if there is an error, uh, we can look at the, uh, logs to, uh, figure out what went wrong. Uh, so, uh, metrics, uh, by definition, uh, refers to, uh, the, uh, metrics as the, uh, data which is collected from, uh, different components of the, uh, cluster to, uh, monitor and measure the, uh, health and performance of the, uh, system. Um, so metrics can provide insights into, uh, key aspects of the cluster, such as superior utilization, uh, memory usage, network traffic, and other performance related data.
So, tools, uh, such as, uh, Prometheus, uh, Dynatrace, uh, and Datadog, um, help us to, uh, fetch these, uh, metrics. Um, Grafana, uh, provides a better view of these metrics, uh, using, uh, dashboards. So, uh, profiling has been the, uh, latest addition, uh, to the, uh, pillars of observability.
Uh, it refers to the, uh, practice of, uh, analyzing, uh, the, uh, performance of, uh, your application and cluster, uh, in order to, uh, identify areas of inefficiency, uh, that may be impacting, uh, the, uh, performance. Uh, so profiling, uh, can help you understand how your applications are using, uh, resources like CPU and memory, and, uh, identify where optimization may be, uh, needed, uh, to improve the, uh, performance. So we'll dive, uh, into, um, the implementation, uh, strategies of, uh, observability, uh, in large, uh, scale distributed systems.
Um, so, uh, we basically, uh, uh, you know, uh, instrument, uh, the, uh, best practices here. Um, so, uh, standardize it by, uh, using, uh, consistent libraries and, uh, frameworks for, uh, collecting logs, metrics and traces. And we ensure that all the, uh, services adhere to a common, uh, standard, uh, for observability data.
Um, this is also, uh, done by integrating Absorbability, uh, into the, uh, CI cd, which is the, uh, continuous, uh, integration and continuous deployment, uh, of the, uh, pipeline to ensure the, uh, new code is automatically, uh, instrumented. And also use the automated tools, uh, to add instrumentation to a code basis wherever possible. Um, so the, uh, data, uh, collection, uh, techniques, uh, include again, the logging metrics, the logging metrics, and tracing.
Uh, we can implement a structured logging, uh, to ensure logs are easily parable and searchable. Um, you can use, uh, log levels like debug, info warn error, uh, appropriately, uh, to control our velocity. And you can centralize logs using tools like, uh, logs, dash and, uh, elastic search.
Uh, so for metrics, uh, you can define the, uh, key, uh, performance indicators, your KPIs and, uh, service level objectives, your SLOs, uh, for monitoring. Uh, you may use libraries, uh, like Prometheus, uh, client, uh, libraries to expose these metrics. And, um, you can also collect both, uh, system level metrics, uh, CPO and memory, uh, and application level metrics like, uh, request rates, error rates and such, uh, for tracing, uh, uh, instrument applications, um, uh, to generate, uh, distributed traces, uh, distributed traces, uh, using tools, uh, like a zipkin, uh, you can ensure tracing, uh, context is, uh, propagated across services, um, to maintain, uh, trace, uh, continuity.
So, uh, for real time, uh, processing and, uh, analysis. Uh, we have the, uh, stream batching and, uh, batch processing. Um, so, uh, uh, you may use a stream, uh, processing techniques, uh, like Apache Kafka and, uh, Flink, uh, to, uh, process observability data in a real time.
Um, for historical analysis, uh, you may use a batch, uh, processing jobs, uh, to generate reports and insights from a stored observability data, um, for effective, uh, visualization and dashboarding. Uh, we create intuitive and, uh, actionable dashboards, uh, using Grafana. Uh, we include key metrics, trends, and alerts and dashboards for, uh, quick insights.
Um, for the, uh, custom, uh, views. Uh, tailored dashboards, uh, are used for our different audiences, uh, for example, for developers, for operations, for managers, um, in order to ensure, uh, that it's relevant to these, uh, target audiences. Uh, we use templating and, uh, variables and Grafana, uh, to create our flexible and our dynamic dashboards.
Um, so alerting is also key here. And we embed, uh, alerts and notifications within the dashboards to, uh, provide context and facilitate our quick responses. Um, another key thing here is alerting.
So, um, we set up, uh, alert, uh, thresholds based on KPIs and SLOs, uh, ensuring they're, uh, neither too sensitive, uh, not too, uh, lenient. And we use, uh, rate limiting and, uh, d uh, duplication to, uh, prevent, uh, alert fatigue. So, alert routing is also a key, uh, concept, and we, uh, use tools like, uh, Prometheus, uh, to route alerts to the appropriate teams and individuals.
Um, and we've also implemented escalation policies to ensure critical issues are addressed promptly. So incident management is key here. Um, so, uh, we've, uh, used, uh, platforms, uh, like PagerDuty, uh, for tracking and a resolution, and we do, uh, internally maintain runbooks and playbooks, uh, to guide response efforts and ensure that there is consisting, uh, hand, uh, consistent handling of, uh, the incidents.
Security and compliance, again, is like key here. Um, so we do this by ensuring absorbability data is, uh, encrypted in, uh, transit and at rest. Um, so implementation of the access controls to a restrict access to sensitive absorbability data is how, uh, we also, um, you know, uh, implement security, uh, for compliance, uh, monitoring of the, uh, logs, uh, for any compliance related events and anomalies are in place.
Uh, we retain, uh, logs and observability data as, uh, required by our regulatory, uh, standards and policies, um, continuous improvement and feedback loops, um, during, uh, incidents, uh, post-incident, uh, review, and, um, you know, know feedback integration is sort of key here as well. Um, so we conduct, uh, a post-incident reviews to identify the root cause and, uh, prevent any, um, and take any, uh, preventive measures. Um, so the use of observability data to inform and, uh, improve, uh, future incident response is also key here.
Um, so for, uh, feedback integration, uh, we incorporate, uh, feedback from users and stakeholders to ensure observability, uh, practices, practices, and also to enhance it. Um, and we continuously iterate on, uh, instrumentation and monitoring strategies based on the real world issues and incidents that we come across. Um, so training our operations, uh, team is also, uh, one of the, uh, strategy here to, uh, implement, uh, observability.
Uh, so we provide, uh, training sessions, um, for like, uh, the development and operation teams on observability tools and our best, uh, practices. Um, so the, uh, key, uh, benefits of observability, like I mentioned, uh, is, uh, improved, uh, system reliability, uh, faster, uh, issue resolution. Um, it also, uh, promotes for proactive performance management, um, and it enhances, uh, user experience, and it allows us to, uh, better collaborate, uh, and communicate, um, between the, uh, development operations and business teams.
Uh, so sharing, uh, visibility into our systems, uh, health and performance, uh, helps us, uh, align priorities and, uh, improves our communication. Um, so looking into the, uh, future of observability here, it's all things ai. Um, so AI can, uh, automatically, uh, detect unusual, uh, patterns in, uh, logs, metrics and traces, uh, identifying potential issues, uh, before they escalate.
Uh, for example, uh, ai, all algorithms, um, that we are trying to, uh, build, uh, can, uh, flag, uh, anomaly, uh, spikes in CPU usage and error rates, uh, enabling, uh, faster response times. Uh, so AI can analyze vast amounts of data to identify the, uh, root causes of issues, uh, more quickly and, and accurately. Um, this addresses the, uh, mean time to resolution, also known as MTTR by, uh, pinpointing the, uh, exact source of the, uh, problem, um, such as specific service or network component.
Um, AI can predict when components are likely, uh, to fail based on historical, uh, data and trends. So this allows for, uh, proactive maintenance, uh, reducing the, uh, downtime and improving, uh, system health. Uh, for instance, uh, AI can forecast hardware, uh, failures or, uh, software degradations, uh, promptly, uh, for, uh, timely maintenance actions.
Um, so ai, uh, driven systems, uh, can make, uh, predefined, uh, actions, uh, to remediate, uh, issues automatically. Uh, for example, if an anomaly is detected, uh, the, uh, system might, uh, scale up resources, uh, restart resources, or even roll back, uh, deployments, uh, without any human intervention. And, uh, this, uh, helps reduce, uh, the, uh, downtime.
Uh, this marks the, uh, end of my, uh, presentation. Uh, if you have any, uh, questions, please, uh, feel free to reach out to me via LinkedIn. Uh, I hope, uh, everybody, uh, was able to get some, uh, good insights into, uh, observability here.
And I thank you so much for your time.