Agentic AI Observability and the Data Behind Trust
Agentic AI observability is now a first class production concern. Yanbing Li, Chief Product Officer at Datadog, joins Alan Shimel on Techstrong TV. Furthermore, she frames how AI agents force observability to grow up as a governance layer.
About Yanbing Li
Yanbing is a lifelong engineer who spent a decade at VMware. In addition, she led observability at Google Cloud and engineering for L4 self driving trucks at Aurora. Consequently, she brings a mission critical AI lens to product strategy. Furthermore, that background shapes how Datadog thinks about agentic AI observability today.
Inside agentic AI observability
Datadog frames agentic AI observability as more than a health check. As a result, the platform tracks the behavior of AI agents, not just uptime and latency. Meanwhile, agents act like distributed systems that touch data, APIs and models on every request.
Yanbing explains that observability now runs as one unified surface. Furthermore, that same surface powers autonomous detection, investigation and remediation. Consequently, the tool that watches AI is itself becoming an AI driven operator.
Why trust and data quality matter
Meanwhile, trust is the eminent question for production AI. Therefore, Datadog watches sensitive data, freshness, quality and lineage across the pipeline. In addition, a silent schema change upstream can make a healthy looking agent return wrong answers.
Yanbing also reports an inflection point at real customers. Furthermore, agent traffic on Datadog grew 30 times in the past 12 months. As a result, one enterprise cut incident response from hours down to about four minutes. Meanwhile, AI agents now handle roughly half of that customer’s incidents on their own.
Explore more artificial intelligence coverage and the latest Techstrong TV interviews. In addition, Yanbing points engineers to hands on trials as the best way to feel how agentic AI observability works.
For more information please visit datadog.com
Transcript
Hey everyone. Welcome back here to techstrong TV. I'm so happy to have my next guest on here.
It's her first time on techstrong TV with us, but it's a company we've been covering, as I was telling her off camera. com 13, 14 years now. I think we've been doing Datadog 12 of those years, or something like that.
Literally since around the time they launched. And they've had a tremendous run, and the Datadog you see today is not the Datadog I first became accustomed to all those years ago. Let me introduce you to Yanbang Li.
Yanbang is the CPO, chief product officer, over at Datadog. Yanbang, welcome to techstrong TV. I'm sorry it took you so long to get on here.
I apologize, but now that you're here, you're going to have to come on again and again. Well, thank you, Alan. Certainly, I've been watching your shows.
And thank you for covering Datadog. Absolutely. My pleasure to be here.
Has it been 12 years? Ten years, at least? Well, Datadog has been around for 16 years now.
Okay. So we've been a public company for now since 2019. Right.
And we also became a S&P 500 last year. Congratulations. So, since you started covering it, we certainly have grown quite a bit.
Yes. Growing into a significant business. And I joined the company about two years ago.
And we continue to see the expansion of the Datadog platform, and especially the accelerations of O Sing related to AI. So, I'd be happy to take it in any direction that's interesting for our audience. Absolutely.
Well, we're going to talk about all the above. But before we do- Yeah ... you mentioned you joined in 2024.
Give people a sense of your background. Yeah. I've been a lifelong engineer, so I've been practicing software engineering for my entire career.
So, my journey to Datadog, actually, I spent lots of time in enterprise and cloud. At VMware, when VMware was the king of the hill for enterprise software, I spent a decade there leading various part of VMware's product and capabilities. And I was at Google Cloud in 2019.
This is actually when I first get exposed to Datadog, because- Okay ... when I was at Google Cloud, I was leading the observability capability that we built to monitor Google, as well as offering them as part of the product set as part of Google Cloud. So that was my introduction to modern observability and certainly started to pay attention to Datadog.
Love it. After Google, I spent three years building self-driving trucks. Really?
So, that was pre-LLM, pre-ChatGPT, so self-driving was the most prominent AI use case- Yeah ... at the time, and still a prominent AI use case today. I think that was my introductions of really thinking about how you actually take AI from an idea, a demo, to a non-compromisable production use case.
So I was working on L4 completely driverless trucks, commercial trucks. Actually, the only company that's operating this without a driver on US highways today, public roads today, so- Really? Yeah.
I remember, a lot of the autonomous vehicle stuff came out of Carnegie Mellon University in Pittsburgh, and then I think it was Google. Was it Google or Uber? Basically bought the entire department, the professors, the students, and everything, and took that on.
And it's a fascinating thing. We're seeing a lot of stuff going on, on autonomous. Yeah.
But Yanbang, I want to bring us back to observability and AI observability. Absolutely. And I've been in technology a long time myself, right?
One of the things you always see when whatever the new hot thing is- Mm-hmm ... they'll slap that new hot thing on front of what they're doing. So we're doing observability, now we're doing AI observability or something like that, right?
Yeah. And there's no shortage of companies that are slapping AI on everything they do, not just- Yeah ... observability.
But there's a difference between slapping a title on and actually doing it. Mm-hmm. And when we look at something like observability, AI is pulling it, stretching it, morphing it, changing it- Mm-hmm ...
because observability, in many ways, can become the governance layer for AI, right? My friend Mitch Ashley, Mitchell and I have been partners for 25 years. Mitch is the analyst from Futurum Group- Mm-hmm ...
because we're part of Futurum Group, and he's written extensively about this. Talk to me about how Datadog looks at AI observability and how AI is changing observability. Yeah.
That's a great way to frame it. Certainly. Let's start with how we're thinking about observing AI.
So- So observing AI requires us to understand what are the different things that AI brings. When we observe traditional workload, you typically pay attention to how is the functionality, is it up and running, and how well it's running, how fast it's running. So that's how observability started.
Now, with AI, it introduced a whole new dimension of not only the basic health of the services and the system we're observing, but fundamentally understanding what it is doing, the behavior. Is it doing the right thing? So this introduced a whole new dimension of observing AI.
So you could take the approach of just slapping it on top of observability, you observe this new dimension. The Datadog approach is, for observability, you can't really see AI in this isolation, because AI is really part of the broader ecosystem, and all the enterprises are already running. And so the way we think about AI observability is, yes, we do need to deeply understand the behaviors of the AI agents in order to achieve that governance objective.
But most importantly, we need to do that as part of that unified system all of our customers are running, because it's never in isolation. Understood. So we can certainly talk about more.
So that's what I would say AI observability. You also touch upon that AI is fundamentally also changing of how we build observability. Our observability is becoming a much more of a autonomous action platform from just simply observing things.
So we're building autonomous operations from autonomous detection issues, investigating issues, and remediating issue, all powered by AI. So that's another exciting angle of bringing AI into transforming observability. Absolutely.
It's funny, it's both the catalyst- Mm-hmm ... for changing as well as the tool to change it with. So, it plays on both sides- That's it ...
of the equation, yeah. Yeah. But you know another thing I see, and again, this isn't just confined to observability, but we see it up and down the stack- Mm-hmm ...
where AI gets involved, and it's what I call this issue of AI scale. Mm-hmm. What worked when the AI wasn't in the equation doesn't work when AI is in the equation.
Yeah. Right? It brings the scale of the information.
Mm-hmm. But then how do I keep my data reliability high, right? Without being overwhelmed.
Yeah. AI almost has to be in the equation to try to not be overwhelmed. The whole scale, speed- Yeah ...
depth, width, everything. It's changing the game. Talk to us about how Datadog is dealing with that, what you're doing.
Yeah. So, speaking of scale and speed, certainly AI is driving us into an entirely new level. So when we first started to think about observing AI, we started observing LLM because at the time, the AI models or intelligence layer is the new novel thing.
Then we extended to observing the entire agent, because if you think about an AI agent, it's really a distributed system of its own, it's just powered by intelligence because the AI agent have access to your data, have access to APIs, and is making intelligence decisions powered by the AI models. And I want to call out the particular importance of how data play a role in making successful AI, because AI is only as good as the data powering it. And so, certainly understanding data has always been an important problem, even before AI.
But the importance of this is being elevated, and also we need to move at a much faster speed now at the speed of AI. So we see what we call data observability is really kind of the other side of your hand, as opposed to AI observability. Again, AI and data comes together to produce the outcome that- Yes ...
we're looking for from those agentic system. Agreed. Talking about equations, another important component of the equation is trust.
Yeah. The people using the product have to trust, because they've got to rely on this, right? Mm-hmm.
So we're talking about, man, the scale is really increased. We've got to trust the data reliability aspect of it. Mm-hmm.
We've got to trust that we're capturing this at scale- Yeah ... in real time. If we're going to use this observability as actionable intelligence, and we're going to act on it, not only that, like with Datadog and other tools, we're going to automate actions.
Mm-hmm. We've got to trust it. Because we can't afford, right, to do things and then wind up with egg on our face.
And- How? Are you confident? I'm assuming you're not going to let any product go out there that you're not confident works as intended.
You trust it. How do you instill that trust in your customers? That's a great question.
Certainly at Datadog, we build products used by engineering teams and operations teams and security teams that help develop and operate and secure their products, and we use lots of those products ourselves, and we use those practice. I do think for AI to be deployed into the real world, trust is eminent. I've learned that lesson from the days of building self-driving trucks.
Obviously, the consequence is life and death. Right. Here, for mission-critical or business-critical AI, there is definitely the same expectation for trust.
And data is a big part of that trust. If I think about how we trust the data, there is a few dimensions we'd look at. First of all, is there any sensitive data that could get leaked through my AI agents or through my entire set of services?
And so there is always a security angle. Then there is also the quality of my data, the freshness of my data, where my data come from, and what is the downstream implications of those data. So we call those data quality or data lineage.
That's also very important because if your upstream application team is changing a data schema, and you don't understand it when you're using your data, your data may already be out of date, therefore producing wrong results. And it's very hard to detect because your AI agent maybe appears to be functioning correctly, but it's just pulling the wrong data and producing the wrong result. So there's definitely the quality, the freshness of the data that's very critical for the house of AI.
And from a Datadog point of view, yeah, we both from observing the data itself to understand those quality and freshness and lineage, as well the security angle, what are sensitive information has been compromised or leaked through my data pipeline. Love it. Yanbing, we're almost out of time.
There's one more area I want to hit on, and it's almost my own curiosity of is asking it. You sit in a really great seat to observe, no pun intended, but to observe how far along companies are in operationalizing their AI initiatives. Yeah.
I do think while it's still early, I feel we have seen an inflection point of certainly Datadog and a lot of our customers are putting AI agents into real production. And also, the sophistications of those AI agents also start to change from simply maybe a basic chatbot to now much more action-taking agentic type of complicated use cases, workflows, and we're very excited to see that. And the fact that we play a observability role allow us to have that front-row seat.
For example, we can see in a data in our own system, our AI observability in terms of the agent traffic that's being sent to this area has increased by 30 X in the past 12 months. Yeah. 30 X.
So it's really at this explosive inflection point. We're definitely seeing that exponential growth in data. We're also seeing that those applications become much more business-critical, therefore they demand a lot more sophisticated observability so that they have that trust, they have that confidence it's serving their customers well, and it's producing the business outcome.
It's actually really an exciting time. But I also think we're still in such early innings of seeing this explosive growth. As part of the AI agents we build for our customers, and I was just talking to one enterprise customer.
We're in the business of keeping our customers' production system healthy, and they were telling me using our AI agents, they were able to transform their human-driven incident response process to a much more autonomous process, and they're starting to have about half of those incident confidently handled by AI agents, and that has reduced their response time from an hour, sometimes to days, to an average of four minutes. Wow. I think we're seeing this transformative power happening in real customers running their critical workloads and seeing how they can keep those workload healthy and seeing that transformation brought in to their operating practice by AI.
I love it. Yanbing, we're over time. I apologize for keeping you longer than I promised, but I was curious.
com, the main site, but is there a subsection or anything they should focus on? Yeah. There's lots of good information on our website.
There's the videos on our YouTube channel and also various social media outlet. But I would say, as an engineer, the best way, just give it a trial. Play with it.
Give it a try. I remember when I was interviewing with Datadog, the first thing is to go hands-on and touch the product. I think that's always the best way to experience and best way to learn.
That's exactly what our audiences tell us, too, right? Yeah. They don't want talking heads and slides.
That's true. They want hands-on workshops. Let me get my fingers dirty, get it under my nails here, and- Yeah ...
really get it. That is the best way to learn, AI or not. Yanbing, thank you so much for coming on here.
As I said in the beginning, now that you've been here, you know where we are. I expect to see you regularly. I would love to come back.
It's been my pleasure talking to you, Alan. Thank you. Thank you.
Yanbing Li. Thank you. Yanbing Li, CPO at Datadog, here on Techstrong TV.
We're going to take a break. We got a lot more Techstrong coming at you.