Revolutionizing Log Management with Logs 2.0 with Alok Bhide | Open Source Summit NA 2025
Logs 2.0 from Chronosphere introduces a new logging solution that enhances customer control over log data and provides valuable insights into data usage. It addresses challenges like log hoarding and data volume issues, offering insights based on engineers’ usage patterns. AI plays a crucial role in improving log management efficiency, supporting DevOps teams, and potentially simplifying tasks for SREs and DevOps professionals.
Transcript
Hey everybody. We're back at the Open Source summit in Denver, and we're talking with Aloc from Chronosphere, who's head of product Innovation. And we're gonna be talking about a thing called Logs of Palooza.
And I believe this has something to do with logs, and we'll jump in from there. But explain, you guys rolled this out at the show. It's brand new.
As I understand it, we're gonna make logs easier to manage and maybe less noisy, Right? 0, which is our, uh, second logging offering. 0 is our ability to offer the customers control over their log data.
That's the big new thing. So if you, if you, you know, if you, our, our entire platform, one of the differentiators about our platform is the ability to control data volumes through insights about what data is used, what data is not used into making informed decisions about what you should therefore collect and how you should collect it. Or we didn't have, have that for logs till today, and now we have that for logs as well, which is the big new thing.
Are people log hoarding or are they like saving so much of this stuff that they don't have to save and then the cost of storage goes up and they get yelled at by their CFO for how much storage they're consuming. But what's the core problem? I mean, I think you hit the nail on its head.
I wouldn't say log hoarding, but log hoarding sounds like they're doing this by, by choice. I think it happens by accident. Uh, it's difficult to always make good choices about what to collect, what to not.
Uh, administration teams are responsible for logging tools, not those that are using it. The tool people that are responsible for the tools always have two sort of things they're trying to optimize for, help the engineers get the data they need to solve an instant, but also keeps cost, keep costs down in that they tend to rotate towards trying to keep more things and throw them away. And as data volumes go up, this problem gets more and more acute, more painful, uh, for them over time.
Is there some way to identify which logs are more important? Is there, do I put a classifier on them, or will the logs tell me somehow that this one matters more than that one? As a, as a user or a customer of our product, you don't have to do anything to pre prior to using the data.
You don't have to do anything to your data to determine what's more useful or not. So the ma the nice thing about our product is it gives you that insight based on your engineer's actual day-to-day usage of the product use, usage of the data. So if your engineers run queries or run que, is it just ad hoc searches or, uh, build dashboards or have alerts based off of those logs, we will know which logs are being used and which ones are not.
And we can generate information or insights that give this information back to these admins that we talked about, admins, tools, owners, whatever you wanna, whoever, whatever the title of those folks are in the company, to give them quantified information about what's useful or not based on usage patterns of their teams. But they don't, as admins or tools owners have to do anything to, to get that information prior to get the, getting the log. Zig no tagging, no marking, nothing like that.
And that provides the added benefit, therefore, of streamlining the signal to noise ratio because I can, the, the platform itself is helping me do that Exactly. So you can, with that information, you can make pretty, uh, quick decisions as to what you need and don't need. Some of it's very obvious where you've got very low utility scores where you really, obviously no one's looking through this data.
And you can look at historical time periods, go back 1530, perhaps longer, uh, long, uh, uh, further, further back in time and determine was this data ever used? Um, and then that's easy data that you can just drop. But there's other data that falls into sort of this gray area.
You need it, but you don't need it in its full fidelity. You don't need all the errors in full fidelity. You don't, every error that comes in that looks the same, but you need to know the number of those errors that happened so that data can start getting treated a bit differently.
You can calculate the number of errors but not keep all the logs. Those kinds of things are possible within our product with those insights, but also the ability to actually act in this particular fashion. You can make those changes to this data coming into, into the product.
Yeah. Um, there's always that one log or two or three or however many it is that wasn't significant at the time, but then, I don't know, there's a cybersecurity incident six months from now and some regulator wants to look at that log, can I go find that and rehydrate that in some way that makes that still relevant? Or how do I kind of deal with those kind of one-off anomaly cases?
Yeah, I think that particular example that you mentioned, like a compliance use case where you have to keep something, I, I would split into, into two types of use cases. Something, a compliance use case where you need to go back to an old log. When I say drop sample or get rid of logs, I mean to say you don't have to keep it in your performant, more expensive logging store.
You may still choose as a user or a vendor to store it in an S3 bucket or some sort of object store. And yes, there's definitely a way to then rehydrate that or pull that back into your performance logging store or just query it straight into, into, into that bucket. Uh, so that can be solved generally through that kind of a strategy.
Um, however, there's situations where, you know, you might need a log where it's really necessary for incident resolution, but you maybe you didn't collect it. Those are the trickier situations, and that's, again, there is no silver bullet for that one. Uh, you've gotta be careful about how you use a tool like chronosphere or any tool to choose to keep things, not keep things.
So I don't want to claim that, uh, you know, solves those kinds of obscure problems. Uh, I think there's, you have to take care using any tool when you're making decisions about what collect, what not to collect. Hmm, Of course these days you can't walk down the street without somebody telling you about their great new AI thing.
So am I gonna apply AI to log management and what's that gonna look like? So I'll talk to you about where, you know, what we're thinking about ai. Um, so chronosphere is primary goal is all about making developers a lot more efficient, and that's how AI will continue to be surfaced in chronosphere.
And when it comes to logs is a prime example, is being able to do things like log summarization or getting meaning outta the logs. That's a very obvious place where you'll start seeing things in Chronosphere where it gives you insights and more insights into what these things mean, especially when they're very complicated, difficult to munch through for anyone human being. But we won't be AI washing our product, uh, and stamping AI in every AI feature every time we improve something within, within the ui Are where do AI agents fit in the DevOps team and the workflow?
Are they gonna be essentially members of the team or are they gonna like, augment DevOps engineers or a little bit of both? It seems people are trying to figure out where that entity sits, and is it, you know, uh, a thing I give a task to and then it's autonomous and comes back? Or is it something that I kind of use that I'm invoking alongside my activities myself?
So I'll, I'll give the, uh, I'll give my answer. Keeping keeping in mind that this, uh, space is evolving very rapidly as we speak. So what I say today may be, you know, again, evolution is pretty fast in this space, but generally what we've seen from customers, I'll speak to what we see from our customers on the whole, uh, it's more the augmentation of the developers or the DevOps engineers.
So a gentech ai, you know, ai, uh, from the perspective of instant resolution or SREs, is about helping them become more efficient. Uh, so in that sense, you could say they're a member of the team, but for very dedicated tasks. There isn't one particular AI that does, uh, agent that does it all.
There is, Hey, we want to be integrating with, uh, this particular tool. I want to be able to do this particular task. So you'll start seeing those things pop up.
And in some sense, that's becoming quite necessary for teams because there's a lot more work to be done, uh, at all times. So this is actually making teams more productive. Uh, but that's what we're seeing from people.
It's not, what we're not seeing from our customers is AI doing it all. Uh, we've not yet seen full fledged end-to-end troubleshooting root cause analysis. Here you go, and then rollbacks, and then that's a, it's a bit beyond what we see right now.
People are now starting to go a place where they see AI as a way to augment, uh, augment their work, make themselves more productive. And that's how I see agents as well. It doesn't sound like SREs are gonna be replaced anytime soon, but I wonder if the opposite might be true, where the number of companies that can afford to invest in higher SREs has been somewhat limited compared to the total population that could.
So might the overall number of companies that can invest in DevOps and DevOps engineering increase in the age of ai, as the cost become more accessible and the technology becomes more accessible. Uh, very likely. Very likely.
I, I do see AI as like with any technology, as it becomes more and more mature, it becomes democratized in some sense. So what was otherwise only possible for certain large organizations are well funded organizations, not just ai. This is not a response just to ai, you'll start noticing that companies start adopting what a big company does over time, because now it's cheaper or things like that.
And same thing I see with AI today. There's many, many organizations that are trying to make a, you know, jump to the cloud. Many have not yet.
And in that world, they need to operate very differently than they do today. And in that world, the advancements in AI will probably make this a lot easier for them, um, because now it's accessible to them. And it hasn't been, and it, you're right, like SREs are not that easy to find.
They're not that easy to, you know, hire. It is not that many of them out there all the time. So that lack of, uh, lack, you know, the, just the lack of personnel out there available just to, just to plug in and the costs will definitely become something that can be more manageable with ai, um, making them more productive.
We've been talking about the DevOps infinite loop as long as I can remember, but it always seemed like there was a gap between the, I can observe something, but that didn't mean I could fix it. 'cause I'd have to go get a different set of tools and there'd be be a little, there's a disconnect. Will we kind of close that gap more and where if I observe something, I can, within a set of policies automatically fix it?
Versus today I've observed something, but then I gotta go, you know, still set up a war room and have a discussion about who's gonna do it. Are you saying with ai could, could we do that? Yeah, I mean, I I'll say this, that it still makes engineers a bit queasy.
Um, I cannot, can, will the world get there In reality at some point it'll get there to be very honest. But will that happen, uh, in a very foreseeable future? I think that's really dependent on how people accept some of this change.
Less so what the technology can do. Uh, it becomes more of a human question, are we okay with this? Are we okay with this on many levels?
Uh, but primarily, do I trust this? Right? Do I trust that this change won't be terrible for us?
And I think we're not in a place to make that. I mean, I wouldn't be in a place right now to make that call this second. Uh, but knowing technology one day it'll happen.
Is it today? Is it 10 years from now and 15 years from now to five years from now, two years from now? It's very difficult to tell, but I can tell you that, uh, most of the customers we speak to and several engineers, I won't name names in my life, do not feel comfortable with this final step.
Um, to that end, I don't think that will be closed by AI yet. Uh, but it's hard to say that one, It seems like there's a push pull here. One thing is, you know, on one side is it's one thing to be wrong.
It's another thing to be wrong at scale. So if I let the AI do it and it goes wrong, it could be catastrophic, right? On The other end, the systems are getting so complex that me as a human, I can't keep up with what's going on anyway.
So, you know, between those two diametrically opposed things, which one wins the day? Uh, in the long run? In the very long run.
And I don't know what that timeframe is. I think what wins the day is, uh, uh, we have to get there. So let's get there.
Which means I think at some point AI will be accept, everyone will accept AI making those changes. Or also, I wouldn't even say ai, some automated system making those changes as opposed to a human being most likely ai. But I think is a question, is a question of fool when, like, is is it it, the challenge is not whether it will happen, it'll probably happen.
Uh, it's always a question of does it happen within my lifetime? And that's a very difficult question for me to answer, but I would agree with you that that tension exists. I think some, they'll probably start seeing some companies starting to accept that change and others following over time.
Just like with the cloud, everyone is yet not on the cloud and it's gonna be similar. So last question, short of all those wonderful outcomes in the meantime, might we just see some lowering of the overall amount of stress that SREs and DevOps teams currently experience? Because, and we don't talk enough about this, but those teams burn out pretty quickly because there is so many things that can go wrong, you don't know when it's gonna go wrong.
So there's always this constant sense of, you know, is it working? Is it working? Is it working?
Is it again, will AI make? Yeah. And, and will that help?
You're getting into a place of very philosophical questions about human beings. I'll say, uh, I don't think human stress ever really goes down. Uh, it's a choice we've made as a people, uh, to, uh, however I'll say I think the fundamentals of the technology allow for the stress to go down.
I think will the stress go down is dependent on how many more things we make complicated for ourselves in the future. But AI has the capacity to make things simpler in the sense that complex systems, complex actions can now be performed by something else. However, I'm making no qualifications as to whether the human stress will go down.
I think that's a totally different question in some ways. All right. Hey folks, you heard it here.
It's gonna get easier to manage logs and we'll make more sense of them for sure in the age of ai, whether or not we're gonna experience less stress. Well, that's on you. Hey, aog, thanks.
Thank you.