Why OpenTelemetry Is Paving the Way for the Rise of the Observability Warehouse
Eric Tschetter, chief architect at Imply and creator of Apache Druid, explains how the rapid adoption of open source OpenTelemetry for instrumenting applications is reshaping modern observability architectures. As telemetry data volumes surge, organizations are moving toward an “observability warehouse” model that unifies logs, metrics and traces into a scalable analytics foundation capable of delivering real-time operational intelligence.
Transcript
Hey guys, thanks for the throw. We're here with Eric Cheddar, who's the chief architect for Imply, and we're having a little chat about, well, the rise of observability warehouses. It seems like we're actually pulling in more data than ever.
We've instrumented applications, but now what Eric, welcome to show. Yeah, thank you. Thanks for having me.
I'm happy to be here. Alright, So walk us through how this is all evolving. I mean it, for a while there we couldn't instrument many things 'cause it cost a lot of money and it was just hard.
And now we have OpenTelemetry and well, it costs less and maybe it's not quite as easy as we want it to be, but apparently we are pulling all this data in these days. So where are we on this journey? Yeah, that absolutely we, we've got OpenTelemetry.
It's pulling in data. I mean, we've also got AI now coming, generating a whole bunch more data and, and, uh, everything and it's all kind of dumping into things. But, um, kind of taking a step back on, on where we are.
Uh, we at, at imply we've been around for about 10 years. We've worked with, uh, open source column oriented kind of database for a while. And we saw across our customers that they fit kind of two profiles.
One in the business intelligence world and one in the observability world. And we realized that in the business intelligence world, there's this clean separation between, uh, layers. You've got your visualization layer with Tableau and Looker on top of your data layer with your SQL data warehouses, snowflake, Databricks, all that, and your data acquisition, ETL tools with, uh, um, five Tran, DBT, uh, Informatica thing, things over there.
But then when we looked at the observability world, we realized that it's all kind of verticalized stacks. It's, uh, you've got Kibana on top of elastic, on top of Log Stash or um, Grafana and Loki with OpenTelemetry. And uh, we realized that, you know, looking back at the history of business intelligence, it got to where it is through a natural market evolution.
And in the observability world, in this observability and security data world, really, it's just the beginning of that market evolution. OpenTelemetry really is it, it kicks it off. It's starting to decouple the data acquisition from the other parts of the stack.
But if we look at business intelligence, we see that there's this interaction layer that decouples from the data layer as well, which is where the observability warehouse comes in. Mm-hmm. I got you.
Um, are we gonna have a unified observability warehouse because, uh, you know, if I look across all of it, everybody seems to be collecting telemetry data. Is it different data or is it all the same data and we just need to figure out how to share it better? Like we were taught in kindergarten?
Well, I mean that I, that's fundamentally the same question that led to the evolution of the BI space is you had this business data, it came from various different business units. You wanna do corporate level recording, you want each business unit to look at what they're doing. You want each of the analysts to see stuff and you're wondering like, how do we share this?
How do we have the data? How do we make sure that it's in alignment with what we want it to look like? And then share it so that the organization can generally, uh, kind of get value from it.
And the answer there was the SQL database. And by putting everything in sql, by abstracting everything by a query language, it became possible to have the data in one place and let each individual group kind of interact with it, with using whatever tool they prefer. That same evolution is what will come to the observability world.
And so will there only be one place to store the data? I mean, that's not how markets work. If there's only one place that's a monopoly.
And, and now, uh, there will always be new people that come in, but people will be looking more and more for I've got my data here, don't make me move it. Just let me query it. I want to use this tool.
Great. Use that tool, query it. You wanna use that tool, great.
Use that tool, query it, use what you wanna use on top of it. There's no reason that just because the data's in one place, it can't be used by multiple tools. The only the that, the kind of separation there is really the query language.
And that's the fundamental difference between the observability and security world and the business intelligence world where business intelligence is kind of, they have a heavy adoption into sql, where the observability and security world has a number of different query languages that people have adopted over time. Well, AI kind of flatten that out a little bit 'cause the AI agent theoretically will be able to speak multiple programming languages and they'll be able to query data where it happens to be found and pull it back into something that looks like, I don't know, a unified dashboard possible. Y yes, but no.
But yes, I, I I don't, I don't know, like, it, it's, uh, like I, I agree with the hope and the vision that AI can just take away all of this knowledge of language and it's just like I ask AI a question that gives me the answer. Um, so far, at least when I've been interacting with ai, I, I, I, I liken it. Um, so I'm not a sculptor.
I don't know anything about sculpting, but I can make up metaphors about it. And I liken it to, if I wanna, if I wanna like chip wood or I wanna do a wood sculpture and I have a block of wood, um, it's really nice and easy to use a machine to like make the general shape. But when you want that to look like a real sculpture, you've gotta get in there yourself and, and like do this stuff.
And I found AI to be the same way. Yes, you can use AI for kind of broad strokes things you can use it to, to start a first draft of a letter. You can use it to even write some code initially, but when you actually look at what it did, you're gonna find all sorts of things that need to be adjusted and changed and things that are updated.
And sure, AI over time will get better, but I'm not sure it's gonna like completely displace the need to understand what it's actually doing to look at it, to evaluate it, to supervise it. And so I, I don't know how much it's truly going to eliminate the need for an understanding of the language versus just reduce the population that needs to have a strong understanding of it, if that makes sense. It does indeed.
Um, I think people are struggling with storage of telemetry data, at least some I talked to. And, um, you know, theoretically at least maybe we've got too much of a good thing now. So how do I figure out what data to store?
So I, when I needed to observe it. 'cause you know, I'll talk to some folks who are like data hoarders and they just store everything. Yep.
Then I got other folks who are barely storing anything. And then of course when something bad happens, they don't have the right data. Exactly.
Exactly. And that, and that's the, um, I think that's been like, you, you touch on a really key point. People in this space have been faced with, um, well, I've got a bunch of data, the unit costs to store the data is higher than I wish it was.
So now what do I do with it? Do I throw some of it away so that the economics work out? Or do I find a different place to put it with better unit economics, but maybe a worse interaction pattern or it's slower to access or I have to take extra steps in order to make it so that I can actually use it.
And, um, these two options have kind of traditionally been given to people. Uh, more and more we see people, there is some amount of pruning and throwing away data, but as you mentioned now when you need it, you don't have it. And especially when you're doing like a security incident investigation or something like that, that can be key to not knowing what actually happened.
Um, so people tend to be leaning towards, well, with the introduction of public cloud, you've got your object stores, it's, it's pretty cheap storage, it's relatively easy to get access to. So people end up funneling, funneling it off into the object stores, but then they have the challenge of how do I make it actually interactable? I have people who use one tool, but now they have to learn yet another tool in order to interact with this data.
Or do I take that data and I transfer it back into this tool when we actually need it or, or what's going on. And, um, talking a little bit about, uh, our product, um, implies product itself. What we've built is we've said, you know, as people are leaning into this cloud storage, what they care about is that unit cost.
They want the best compression and, but they want it, they don't wanna give up the ability to actually query it and interact with it from the tool that they have. And, uh, we've found that when you take compression and you make it kind of domain specific, so for us that's logs. When you compress specifically for logs, you can do a really good job in order to reduce the overall bite count, which fundamentally then turns into a lower unit cost.
And you can also use the latest and greatest indexing techniques to make that compressed form still queryable. And that's kind of the, the product that we have. And that's kind of what we like to call this observability warehouse, where once you can take those logs, put them in a queryable and compressed format, you can minimize the unit costs and then you build the connections, you work with all the query languages of the different tools that people are using.
And now you can store the data once and interact with it from multiple different tools. Mm-hmm. Um, Will we ever get to the point where I just come in in the morning and there's, you know, a memo from some AI tool somewhere that just says, you know, here's the three things we found in the observability warehouse that are likely to get you fired and here's our recommendations for doing something about it.
Um, yes. Now are those three things actually legitimate? I think that's the fundamental question.
I I, I think we're already at the point where you can get a list of three things that, that, uh, AI says you should look at the, the question is the signal to noise ratio. And I'm, I'm sure there's, uh, di different things there, but, um, in general, like I, I, uh, talking about AI and, and my own personal framing of how AI is gonna impact the data world in general is I like to think, I like to think that history always repeats itself. And so like this whole business intelligence to observability warehouse thing, that's just a repetition of history talking about ai, I like to think of it in terms of manufacturing automation.
And so way back in the day there was the assembly line with people standing on the assembly line, doing things, putting stuff together. I wasn't there at the time, so I don't, I'm, I, I'm, I am, uh, perhaps making stuff up. But, uh, there were people on an assembly line doing stuff.
Then we started come out with robots to kind of automate the assembly line. And I'm sure when that was coming out it's like, oh man, humans are not gonna have jobs anymore. There's nothing gonna that we're gonna be able to do.
We can't do anything we that all of this. But what has actually happened is the assembly line still exists there. It's more automated, there's more robots on it.
There are supervisors and humans who watch it. There's people who design it and lay it out. There's people who manage it.
But that assembly line still exists. And I liken the assembly line to the data platform no matter what, how, no matter what robots you've got, what, no matter what AI you have, it has to interact with some data. So you need some system that that has the data.
But the other thing that automating the manufacturing has done is by increasing the throughput of one assembly line, it's actually had a ton of knock on effects on the supply chain because now you need more materials to get to one location in order to actually populate that assembly line, generate the production, and push everything out. And to me, I see that, uh, the metaphor over there of AI and AI consuming data is going to be querying the data platforms. It's going to be interacting with the data and the data platforms, but that's gonna have a knock on effect of being able to consume, aggregate, manage, and deal with so much more data that there's gonna be a lot less hurdles to people being like, oh yeah, but I can't do anything with that data anyway.
Nothing can look at it. No, nobody has time to look at it. And so it's gonna actually increase the pressure and increase the demand for more data coming in, which is going to require the data platforms to really scale out and truly optimize on that unit cost and make it available to the AI agents to, um, to tell you the three things you need to look at that then you look at.
And, uh, maybe they're right, maybe they're not, but mm-hmm. Yeah. Do you think It will get easier to instrument these applications and all these data sources?
'cause I think we solved the problem of the cost, but I'm not quite clear that the, uh, OpenTelemetry agents that are out there are easy to install and maintain just yet. So what do we need to do there? So My first answer is that my focus is entirely on the data platform and the actual agent tree to get it in.
I'm like, I'm unopinionated about it. I'll take data from anywhere. I don't care.
I just wanna store it as in the best way possible. On the flip side, I think like actually instrumenting things, connecting it, working with agents, um, I don't know. I I, I see it as a bit of a last mile problem.
It's always gonna exist. People are always develop, so like developers are always gonna do new things. There's always gonna be some hot new framework out there that like won't have a thing integrated with it.
When you actually get to the content of the data, which is another part. There's like, can I get the data from point A to point B? But then there's, can I actually understand what the content of it is so that I can make use of it on the other side?
And like, especially when it comes, comes to logs, the content of a log line was something a developer just thought. They're just like, oh, I'll type this because that means something to me. Doesn't mean anything to anybody else, who knows?
But it's all extremely schemaless. There's like, not any specific schema. There's usually commonalities between it, but it, it, it's all varied and, and different.
And I don't know that like, there's kind of two approaches you can take with this sort of thing. You can try to ratchet things down and provide people strict definitions that they're supposed to fit inside of and, and expect everyone to do that. The humans don't do that.
I don't know, at least I, I've not found it easy to get humans to do that. Um, I, I, I think, uh, humans are much more likely to just kind of go off and do a thing, and then you're left figuring out, okay, you did a thing. It actually has nice outcomes, but it's not aligned with these other things.
So now how do I figure out how to align them and smash them together? And I, I think that's actually the most important thing for the tooling and the, the platforms in this space to focus on is don't try and change people. Try and make it so that when they go and do things, you can still fit it together and get value.
Mm-hmm. You know, I'd love to get your opinion on this, but you know, as far as I can tell, there's gonna be a bunch of AI agents that are trained for different tasks that are gonna be tapping into this observability warehouse for data to figure out what they're gonna recommend. But at the end of the day, won't they just argue with each other?
Like humans do. I, I mean, yes. But, but their arguments are gonna be extremely positive.
They're always gonna start with, you are so smart. That was the greatest idea I've ever heard. Let, let, let, let me go do, let me go again.
No, I, I don't know, like, I don't know if you've experienced this, but every time I put something into ai, it tells me I'm the smartest person in the world. I'm like, wow. It, it, it took me a while to get over it.
At first, I started thinking, oh, I am smart. But then I was like, no. Yeah.
Anyway, who, who, Who, who knew they could program sy pants, right? Yes, exactly. Oh.
Uh, but yeah, I mean, yes, and it, it, you can think of us humans as the ultimate AI is, and in which case, I mean, are, are we gonna produce, i, is an AI gonna be able to, um, exceed what we are? Um, I don't know. Uh, uh, we'll, we'll see.
I think it'll be different, but, um, fundamentally, history always repeats itself. All right, folks, I think you're it here. We're gonna have better observability for sure.
And that's a good thing. Exactly. Who's gonna be observing what we don't know yet.
Eric, thanks for being on the show. Yeah, absolutely. Thank you.
All right. And back to you guys in the studio.