Agent Observability Redefines AI Trust at SplunkConf 2026
Agent Observability Expands the AI Trust Conversation
Agent observability is becoming essential as enterprises move AI agents from experiments into production workflows. In this Techstrong TV interview from SplunkConf 2026, Alan Shimel talks with Vikram Chatterji about the new requirements for monitoring, measuring and controlling agent behavior.
Chatterji explains that agentic software is different from traditional applications. It is probabilistic, adaptive and often made up of multiple agents, tools, memory layers and handoffs. That means observability can no longer begin only after production issues appear. It must support the full agent development life cycle.
From Telemetry to Agent Behavior
The conversation explores why OpenTelemetry matters for AI agents. Enterprises need a common language for tracing agent activity across internal systems, third-party agents and complex workflows. Without that standard layer, teams struggle to understand what happened and why it happened.
Agent observability builds on traditional monitoring, but it also adds a different product experience. Teams need to see tool calls, agent-to-agent handoffs, memory use, datasets, evaluations and outcomes. They also need insight into whether an agent is creating risk, wasting resources or drifting from its intended purpose.
Signals, Supervisors and Humans at the Helm
Chatterji describes an insights-driven approach that helps teams find unknown problems inside large volumes of agent data. Instead of asking people to scan endless rows of logs, a supervisory layer can surface suspicious behavior, inefficient tool use, privacy concerns and other signals that deserve attention.
That does not remove people from the process. It changes their role. Human teams stay at the helm, while AI-powered supervisors help them focus on the traces and decisions that matter most.
Tokenomics Brings Cost Control Into Focus
The interview also highlights SplunkConf 2026 announcements around Cisco Data Fabric, Splunk Agent Observability and tokenomics. Chatterji says enterprises are adopting AI agents quickly, even when budgets are under pressure. The challenge is not simply whether to use agents. It is whether the value justifies the cost.
For technology leaders, the takeaway is clear. Agent observability gives teams a way to connect trust, performance, cost and governance. As AI agents scale across the enterprise, that visibility becomes critical for building systems that are reliable, accountable and worth the investment.
Transcript
Hey everyone. Welcome to our coverage of Splunk Conf 2026 here from the Mile High City of Denver. It's great to be back at Splunk Conf.
I haven't been maybe at the last two, at least I can think of, but always a great event, great conference, great information, great people. I want to introduce you to one of those people right here. His name is Vikram Chatterjee.
Yes. Did I get that right? You got it perfect.
Yes. Thank you, Vikram. And Vikram is fairly new to Splunk Cisco.
That's right. Is it Cisco Splunk or Splunk Cisco? I don't even know.
We think of it as just generally one Cisco. It's one Cisco. There you go.
Cisco is a very, very big part of it. Absolutely. So first of all, Vikram, welcome to Techstrong TV.
It's great to have you on here. Thank you so much. Good to be here.
My pleasure. Let's talk a little bit about you first. Sure.
I mentioned you're relatively new to Cisco. Yes. But you've been in tech, you've been on this journey now a while.
Tell us a little bit about your journey. Sure. So without boring the audience too much, the really short version of this is, I started a company called Galileo.
We ran it for five years, was doing really well. It was the leader in the world of agent observability. Galileo has always had its roots in the world of language models.
Before that, I was leading product management at Google AI. Okay. So we worked very closely with language models, even at Google- Sure ...
because Google has always been very early in the game in the world of AI. And so very quickly, we were seeing that there's this big, big problem where anybody who's building any kind of software with language models as the core operating system is having this problem of, "I don't know what it's going to do," which is vastly different than traditional software where you have very deterministic outcomes. It's supposed to do X and not X.
It's deterministic versus probabilistic. That's exactly right. That's what gets you.
And so we're the leading provider of agentic observability, so that you can just get complete monitoring, but also have the measurement of the agents within that, but also now be able to tell about what the value of those agents are overall. So two and three months ago, I lead product management here as well, and generally this whole area of agent observability and agent trust. Got it.
So it's been a bit of a ride then. You've started five years ago. Yes.
Sold the company to Cisco. Yes. Now working as part of the Cisco and the Splunk unit.
Yep. All around agent observability. That's right.
Yeah. And of course, observability is such a loaded term, right? I personally have seen the kind of definition and the playing field of what we consider observability and what it does and what it's used for- Yeah ...
change so much, right? Yeah. But let me ask you.
Someone says agent observability. Vikram, what does that mean? Yeah.
What do you tell them? I think of it as it's at its core, it's all about monitoring and measuring and controlling agent behavior and agent outcomes. That's the core.
That's the overall really tough space to navigate because when you think of observability in general, historically, it's been very much focused on the engineering group and the group which rolls up typically into the engineers or the CIO. Right. But what we're seeing now is for the world of agents, it's different because these are all closed loop systems.
It keeps learning, it improves, it has feedback loops, et cetera. Mm-hmm. So observability now is more to do with the entire agent development life cycle.
It's the offline side pre-production. It's also post-production, right? Sure.
Because if you look at most observability platforms before, it's like, we literally go into production, and I'm going to show you a little dashboard so when things go wrong, you can see a little blip there and get an alert. That's basically it, and then you'll hook it up with PagerDuty or something. Right.
Donezo. That's it. But now it's not like that at all.
It's kind of like before you go into production- Well, there's a shift left of- Big shift left ... observability. Exactly.
Huge shift left there. But the other piece is if once you know what the failure modes are of what you're trying to build out, you have to proactively control that and stop these agent misbehaviors in the tracks, which is again, a very big difference. So that means there's now the audience in the room who cares about this problem of the risk of agentic software is not just the CIO anymore, it's also the security groups.
So the big, big shift that I see is agent observability is not just a traditional observability problem, it's also becoming a security problem. Right. That meld is very interesting to see play out in the enterprise.
So I've been tracking observability probably since about 2015. Mm-hmm. I remember back in the day when application perform, APM- Yeah ...
application performance management and so forth. Yeah. And companies like New Relic- New Relic.
Yep ... and some of the older folks. AppDynamics- AppDynamics ...
which became part of Cisco. Cisco. Yep, that's right.
Right? Let's call it toying with the idea of observability, because you really couldn't gather the data you really, really needed to get that- Mm-hmm ... picture that we get now.
That's right. We didn't have OTel. We didn't have OTel.
The guys who sold to Service Metric, what was it, LightSpeed or something like that? Yeah. Really brought OTel and a lot of that in Prometheus and so forth.
But observability started to take on its own world. Mm-hmm. Including that shift left.
Yeah. To go into pre-production and so forth. That's right.
I don't know if we were thinking about agents at the time, though. Yeah, absolutely. I think that's the other key piece of this, right?
Yep. Is now it's not just the observability that I knew- Yep ... from 2015 on that we've been tracking.
Yep. Now we need agent observability. Yep.
Let's talk a little bit about what the differences are. If any. Yeah.
Maybe it's the same observability is observability. Right. What are the key pieces there?
There are a few things which are different, but then there are also some lessons from the past that we have to take into the future. Right. What do I mean by that?
So as we were thinking about agent observability for the last two or three years, the ask from customers was exactly this. " Right. "I want to see everything across all my systems.
" Yes. Codex and Claude and whatnot. They're all agents.
How do you even know what they are? You don't know. Exactly.
We don't know what's going on inside of them. We don't get full access, but I want to know what's happening. So what that meant was we needed a single language system that we could all speak the same language.
And you saw two, three years ago, we actually saw a couple of these open source telemetry languages that came about, which started doing the rounds. But eventually what happened is OTel basically stood up, and we actually contributed- Yeah, contributed a lot. Exactly.
We contributed to a lot of that, frankly, to make sure that we could have an AI and agent version of OTel. Oh, okay. With the right kind of...
It spans and traces and sessions and stuff like that. And so that was very important. That has been important.
It's still an ongoing thing, but what that allowed us to do is no matter what the agent is, we all speak the same language. So now we have simple adapters so that we can actually trace data coming from every single agent. Using the same OTel that we use.
Same old OTel, yes. Which is great. Great.
So the need for the industry basically quickly realized that across the board, the need for standardization is extremely important. If you see this in any big industry shift, there's everyone trying to do their own thing and telling you that this, especially because of that reason. " And then the wave of that just takes you with it.
Also see this with MCP servers, for instance, right? Yes. Absolutely.
There was competition to it for a brief hot second. Day and all of that stuff. Yep.
Exactly. But then eventually people said, "No, you know what? " That was almost lightning like, though- Yes ...
because the move to standardize around MCP- Yes ... was like I haven't seen anything like that. Very quickly.
Within three months it was done. Exactly. Everybody was moving on.
Exactly. Was it secure? We'll figure it out.
Yeah. But that's another story. Yeah.
So that makes sense to me. So we didn't have to sort of reinvent the observability wheel just because- Yeah ... we wanted to observe agents.
You didn't have to reinvent the telemetry wheel. Right. What you do have to reinvent is what happens after that.
Because the workflow for observability used to be very much about get the data from telemetry and then show it in the rows and rows of data, the dashboard. Mm-hmm. That's the simple UI, and then it gets into remediation and stuff like that.
But in the case of agents, it's different because these agents, they're systems in themselves, right? Right. The model is just one very tiny part of the system.
It's a very important part of the system. But you have tool calls, you have memory, you have the harness itself that you're using. There are so many different aspects of this.
As with binding. Exactly, and now what's happening is if you want to perform a single task, let's say it's a contact center AI, you're calling the agent and it's doing this entire conversation with you. Guess what?
It's not just one agent anymore. Right. It's many, many small agents talking to each other for a different specialized- Sub-agents and everything else.
Exactly. And so now you have agent-to-agent handoffs happening, and so the entire system overall has to be looked at in its holistic totality. And so in that sense, the downstream observability product experience is actually very different than what you had before.
The traditional. Traditional. You have to see where the data is.
You have to have experiments. So when you look at the product experience of agent observability products, it's different. You have to have the offline, online.
Datasets become important. It's just different. Absolutely.
Vikram, I've been involved in startup world for 35-plus years. Yeah. A mantra that I've learned early on is that everything can be measured, and we measure everything.
Yep. I remember when I first encountered Splunk, probably 2005, '06, something like that. Mm-hmm.
No company kind of captured that better than Splunk. Mm-hmm. It could measure- Anything ...
anything. Yeah. It could keep anything.
That's right, yeah. Infinite scale and, yeah. Coming home, I remember from the first conference where I saw Splunk and I spoke to the VC, who was the backer of the company I was doing then in security.
" That's right, yeah. They capture anything. Anything, yeah.
Everything. Yeah. " Yeah.
Like my children when they were young, they wanted everything. Yeah. You find out that having everything isn't always the best thing.
Mm-hmm. Because with everything, it gets lost in a sea of opportunity- Mm-hmm ... if you will, the real nuggets of what you need.
Mm-hmm. And then sifting through everything- Mm-hmm ... to find out what's really actionable, what's really important- That's right ...
becomes harder- That's right ... the bigger the data lake gets. That's very right.
And then frankly, the cost- Yeah. That's right ... of just maintaining everything.
Mm-hmm. When we talk about agents and agent ecosystems- Mm-hmm ... swarms of agents- Mm-hmm ...
sub-agents. This is repeating that experience all over again. Yep.
How do we be efficient? Yes. That's a very good question, actually, because what's also happening in the world of agents is the amount of machine data that's being generated is increasing at an exponential rate.
Crazy. Right. So if you remember, I was talking about how you have rows and rows of data in observability systems.
Mm-hmm. " Yeah. It's very, very tough.
So, something which is, again, generally available today is within the agent observability product, we had worked on this notion of how do you take all the telemetry, and how do you take all the different metrics and all the different guardians that have been built in and throw that into a really, really powerful insights engine, which is also powered by a very large language model, right? Yep. And that engine's sole responsibility is to figure out what the unknown unknowns are.
It kind of knows what the intent of that agent is. It's kind of like thinking of it as a supervisor for- Okay ... that very specific worker.
And it knows exactly what that agent's trying to do. It knows what the outcome's supposed to be. It knows what the data is coming in from the users.
It knows what it's doing over there, and it knows the way it's making its decisions. And the job of the supervisor is to write down, like, "Ah, it made this tool call over here. That was kind of inefficient.
It could've done this other thing instead. It set this thing over here that has PII data in it. That doesn't sound right.
" So, that's basically a feature- Sure ... which we call the signals feature, which automatically is always on. It's constantly giving the user this thing.
And so from there, they can go into the trace, which matters. So in essence, you kind of flip the root cause analysis- Flip it on its head. Yep ...
right on its head completely. " That's right, yeah. But as you said, the amount of data we're dealing with- Yeah ...
the scale- Yeah ... is beyond a human in the loop at that point. Yeah.
So we need this agent in the loop- That's right ... this agent supervisor in the loop. Yeah.
But you still want humans in this. Yes. I call it humans at the helm.
That's right. That's very good. That's a very good way to put it.
Which moves you up one layer. Yes. One app check layer up.
Yes. And so they're kind of supervising the supervisor, if you will. That's right, yeah.
The way you think of this, Alan, is you're supposed to give these humans, where they shouldn't and can't go away, give them superpowers. When you have all of this data, how will you make sure that you can make them do their jobs more efficiently? And at the end of the day, just help them sleep at night.
Yeah. So that's the main goal. So that's what all these different features are trying to do.
They are the helm, but then they can't go through the logs. That's a horrible job. Right.
Make sure that it becomes easier. That's not anyone's idea of a fun Saturday night. Yeah.
I'd like to turn to Splunk Conf. Yes. We're here at Splunk Conf.
Yes. There's been a lot of announcements, a lot of news, ton of new functionality, and obviously a lot of it with AI and agentics and so forth. Yeah.
If you can highlight maybe some of the bigger things around agentic observability- Yeah ... from Splunk Conf this year. Yeah.
Well, this is my first one. Welcome. Thank you.
And it's been absolutely spectacular. And there's been, obviously, as you saw during the keynotes as well, there's been a lot of focus on Splunk being the system of record for the AI era. Yes.
And there's been lots of really, really cool innovations that we made there. We announced the Cisco Data Fabric and a bunch of other pieces there which are built for the AI era. Which is, to your point before, that you can actually throw as much data in it, and it can also measure whatever you want.
That is actually what is very, very important to the enterprises right now, as they're dealing with petabytes and petabytes of data. Yep. It's like, where do I put all of this?
If only there was one singular system where I could put this, which is built for AI, that's exactly what we've announced today. That's one of the big things. But the layer on top of that is observability, which is the Splunk agent observability product, which is also generally available today.
And another thing which comes up quite a lot when I talk to customers is, all this is great, but it's very expensive. Agents are very expensive. Running them at scale is expensive, but sometimes it's okay to pay for it as long as there is value.
So how do you do this cost value trade-off? And we call that tokenomics, and that's something which we also made generally available today as a part of this agent observability product. Tokenomics.
Yeah. I love it. Yeah.
I love it. It was token maxing. Yes.
That's right. Now we have tokenomics. That's true.
That sounds like a name for a book. Yeah. But it's also an evolution, right?
" There are lots of organizations that said that. Now no one's going to say that. Most people won't say that.
But then the question became, all right, great, let's adopt it like crazy. Let's go all the way, token maxing. Yep.
" Now it's in this maturity curve of how do I make sure that I adopt it, but with the right kind of throttles in place. We do Techstrong Gang every day, and today's Gang, we actually spoke about two surveys around agentic adoption at enterprise levels, and it's really crazy. First of all, 48% have exceeded their budget already.
Mm-hmm. Yep. But here was the interesting thing.
Out of the 48% that have exceeded their budget, something like 6% are saying, "Well, we don't have budget. " Ah. 94%- Yeah ...
" Yeah. "You know what? " Just go, yep.
Exactly. Right. We're going to ask for more budget- Yeah ...
or we'll take it from other budgets. That's right. We're not stopping.
That's right, yeah. And I think that's an important piece here- That's right ... of this, is that organizations, especially at the enterprise, value what's happening with agentics now.
Yeah. It's so mission critical. Yeah.
It's so imperative. Yes. It busts out of the budget.
Yes, it does. It's busting budgets. It does.
But that's where we meet them where they're at- Mm-hmm ... which is we are spending like crazy. " Right.
But then also, you should be able to do something about that. So you can actually take control, so you can take actions. You can actually stop that agent centrally through the system.
That's very important. You have to give them that power. Absolutely.
Look, some things don't change. Like I told you, I learned 35 years ago. Yeah.
Everything can be measured- Yeah ... and we measure everything. That's right.
Yeah. That's a good place to end this. That's right.
Vikram, welcome. Thank you so much. Thank you for coming here on Techstrong TV and attending your first Splunk Conf.
Of course. Really appreciate it, Alan. Thank you so much.
Thank you. Hey, we're going to have more coverage here from Splunk Conf. You're watching Techstrong TV.