82. NetAIOps Has Its Challenges – Tech Field Day Podcast
The industry has embraced AI for every possible problem. Operations will eventually embrace it as well but questions remain about how it will be implemented. In this episode, Tom Hollingsworth sits down with Pete Welcher, Rita Younger, and Jonathan Davis to discuss the issues that remain with implementing AI into an operations workflow. They discuss licensing and procurement, the need for institutional knowledge, and how this will all work in a multivendor world. They wrap up with some guidance about how to approach your next big AIOps project.
Transcript
You've just been put in charge of deploying net AI ops in your organization and you don't know where to start. Do you just take the statement of work and bill of materials that you were sent? Or do you need to get a little bit more information before you're ready to spend a lot of money?
In this episode of the Tech Field Day podcast, net ai ops has its challenges. Welcome to the Tech Field Day podcast, where we bring together a group of IT technical experts to discuss a single idea about key concepts in the industry. This podcast features a variety of perspectives from members of the Tech Field Day delegate community, and is often associated with one of our events.
Tech Field Day is a part of the Futurum Group. This podcast is also published on our sister company site at Techstrong tv. On this episode, as we head into networking field day, we will be discussing the role of net AI ops in your enterprise.
But before we get into that discussion, let's meet who's on the panel today, starting with Rita. Hi, I'm Rita Younger. I am practice lead for Data Center networking at WW t.
I'm Pete Welcher. I'm am a longtime networker fan of technology, start early CCIE and taught various subjects and, uh, just expanded my knowledge base from there. Hey folks.
Jonathan Davis for JD here. Uh, I am a, uh, senior Wireless and wired, uh, network engineer with Eastern Datacom. Alright, and of course, I'm Tom Hollingsworth.
I'm an event lead here at Tech Field Day specializing in all things networking. Let's jump into the premise for today's episode. AI is the solution to all of our problems, whether it's one we actually have or one we are not sure that we have, and it's gonna make operations so much easier, just install it into your network and everything goes so smashingly Well, but there are challenges associated with net AI ops, and if you're not asking the right questions, you probably won't understand what roadblocks you're about to run into.
In this episode of the tech field, a podcast net AI ops has its challenges. I kind of wanna open this up to all of our guests today to kind of maybe throw out one idea that you have that creates challenges for companies that are looking to deploy net AI ops. Uh, because we've seen this before, right?
With SecOps and DevOps and DevSecOps and ops devs, sec, and every combination of words that you can think of there. Uh, we, we've, we've taken some of the other words and ops and net and AI and dev and all those things, but maybe Pete, I'll start with you. What, what's one of the challenges that you see when it comes to throwing net AI ops into the real world today?
Uh, it's similar to the automation space. You're talking big software, big software takes installation. It has to be integrated, it has to get data sources, and it has to be maintained.
So there's a lot of knowledge there. And all of that translates into cost. And that's where, where maybe organizations fit into different, uh, cost tiers.
What can you afford? Can you actually afford this? Whether it's somebody in-house, uh, doing it with, uh, open source or a paid product.
That, that's honestly true. One of the things that I've noticed recently is the number of companies that are offering multiple licensing tiers based on the amount of data that the data lake has to ingest in order to be able to do operations on it. Uh, anyone who's gone out and bought something like, I don't know, Cisco, Splunk knows that there's workload licensing and then there's ingestion licensing.
And the question is, um, are you more worried about how many workloads you've got, or are you about to send this thing gigabytes worth of traffic on a monthly basis that you're maybe not prepared to pay for? I asked Ken Tech about their pricing model, by the way, and that came up Well, and the thing is, is it's, in a lot of cases, it's not a matter of whether or not the company is trying to charge you to ingest the data. It's that the company has to pay to store that data too, especially if it's a service or organization, because Amazon cloud storage is not free, and in a lot of cases it is cheap enough, but when you're looking at the scale of gigs or terabytes or even higher than that, it, it can be a huge component to the cost of the service that you're offering.
And some of the power of net AI ops is gonna be the application side where you're getting application performance statistics. Although deploying that one application at a time for a thousand applications could be a, uh, ongoing fund consulting task. Oh, yeah.
This is not just a group of raw sis logs anymore. This is heavyweight data that needs to be munged and crunched and and dealt with. And, and when you create the analysis on top of that, you're just adding more data onto the pile.
And again, that has to be stored somewhere. And if you want it to be retrieved in a reasonable amount of time, that requires expensive storage tiers and, and that somebody's gotta pay for that somewhere. Right.
Well, and I think the other aspect here that we, we kind of have to address while we're thinking about this, uh, this amount of data that we're uploading and and we're analyzing to make these decisions ideally, is it's ultimately based on this idea of an idealized network. And most of us don't have that idealized network. We have problems, we have issues within the network.
And, and in a perfect world, you know, net AIOps helps us find those and helps us resolve those. But if we are running a less idealized network, if we're running a network that barely meets maybe the, the definition for a, a designed network, much, much less a well-designed network, the, the amount of noise, just like the, just like the issues we had with, you know, slog or SNMP or anything else, the noise floor comes up. And then once again, we're running into, now we're paying a lot for the, for, for all of that data that we're ingesting and analyzing, but two, the quality of the data coming out, uh, you know, starts to decrease as well.
Right. Pete? I think, uh, I love the, um, you, you put out a, a a, a set of posts that basically were around, um, how do we takey log data?
How do we, um, ingest that locally before we then upload that to our, you know, to our cloud solution? But again, one of the benefits that comes out of that is, is now we are, uh, decreasing that noise floor to ensure that the, the data we get out is more usable. I think, I think without doing that, and without having that appropriately designed network, you're, you're essentially paying for a lot of, uh, for a lot of confusion.
Some of the, some of the customers I deal with, particularly in the financial vertical as well as, um, public sector, they don't want to use the cloud or cannot use the cloud for various reasons, and therefore they have to have everything OnPrem, these huge data lakes, uh, the capabilities all OnPrem, that is incredibly expensive. Yet at the same time, it's also a very exciting time when we hear about, uh, what Juniper's doing with Marvis and Cisco with AI canvas. So some of these kind of packaged solutions, rather than build your own, I think are gonna be very appealing.
It's just where are we really at? Like AI canvas is currently alpha, when are we actually gonna get to these solutions? Uh, but eventually I expect to see self-healing and even self-driving data center networks in particular.
You just triggered a thought, um, which is, uh, uh, let's say AI assisted deployment because what I'm sitting here thinking, uh, some large customers that will remain nameless, uh, with thousands of routers and a tech staff that was pretty darn sloppy about documenting anything. We did a project and kept finding devices that were ancient and they thought they'd modernized everything, for example. And so that's a case where I'm not sure anybody's really thinking about it, but part of the software might actually be de detecting cleanup tasks.
How do you make your network more robust? Yeah, I think one of the challenges that Rita pointed out is the fact that a lot of these solutions are kind of in progress right now, and they need our data in order to work for us. So the question that becomes, I I, I think of this as like the box problem, right?
Like, I have the perfect box that I'm saving over here in the corner because one day I'm going to need to use that box, but until that day comes, that box is taking up space and it's not useful at all. And you know, in the case of like, what if it was in a storage unit where I'm paying to have that box and that storage unit until the day when I need it, is there a point where I have to cut my losses and say, I'm just gonna stop storing all of this extra data because eventually will come a day, but we are not anywhere close to that anytime soon. So I need to reclaim what I can because, you know, there's an opportunity cost for everything.
Well, what Pete mentioned about, um, customers don't know what they have, what they're using, they're typically very horrible at documentation for the most part. So being able to have these AI ops solutions handle some of that documentation, maybe we can reclaim some storage space. Maybe we can reclaim some space or know what needs to be refreshed.
But with that said, any of those older devices may not be able to, uh, have the telemetry available that we would need for AI up. So it's kind of a catch 22 there, Right? And, you know, uh, we've heard time and again, how we can use Syslog or we could use NetFlow, but there are better options out there.
And I think of the people who are like, yeah, but I'm still driving a, uh, an 89 Firebird, and I'd love to have all of these cool new options in my car, but I can't because I'm still driving an old car and I'm not gonna upgrade anytime soon. So a lot of this is kind of founded on the idea that you're gonna need to upgrade. And I feel like some companies are using that as kind of a, a lever to push an upgrade cycle is Oh yeah, we'd love to give you all of this additional information through software, but not with a hardware that you've got.
You, you just need to go out and write a much bigger check and the world will be a better place when you do. Did you say you managed to say all that without saying the word telemetry? You're right.
Because who knows if telemetry is gonna be the thing that that takes off in this? I mean, we've talked about it for years, right? But I, I still don't know that, that everyone has sold on that.
I mean, people are still configuring an SNMP on things, right? Well, well, SNMP has one advantage, which is, it's kind of standardized sort of with extensions, whereas telemetry, it's every vendor for themselves. So you buy into a new vendor, you've got a whole this adventure, uh, to look forward to.
Yeah, exactly. Um, jd, I maybe I'll turn to you. What's a challenge in net AI ops that you think is, you know, something that people might not understand or, or maybe they're not prepared for?
Um, well, outside of just the, the quality of the data, um, that's being, um, that, that's being uploaded, I, I think it's, there's also a level of experience that needs, uh, to be present to understand when the model, uh, is hallucinating, right? Um, because there's absolutely been those situations where, um, maybe, maybe the, the data and the, the model is saying, Hey, this is the problem. And of course, the issue with that is, is we, we make early decisions with a implicit trust, right?
We go, okay, if that's the problem, I'm going to go solve that problem rather than let me confirm that that's the problem. Let me, let me verify that the data supports, that's the problem. And, um, and we can spend a, spend a lot of time, um, essentially, you know, chase, uh, in the wild, in the proverbial wild goose chase, um, only to come back around later on and, and realize either a, maybe that's just noise in the data, maybe that, that there isn't a problem that exists or maybe that problem exists somewhere else and there's another solution that might only take a few minutes to, to solve.
Uh, but we never actually get there because we've, we've kind of taken for, you know, taken the, the answer that's been offered to us because it's easy. I think that's the, that's the concern I have. It's something I've seen, um, you know, in, in the real world a a few times.
Uh, and, and it goes back to, um, you know, just like, just like using, you know, any LLM or anything, anywhere else within ai, you do have to be aware that hallucination exists and, and, and we, that what is offered to you may not be the right answer. So there's, so I think this idea of, we can let you know low, uh, uh, level, level one people make level three decisions with the help of ai, that, that, that's not, that's, that's not the reality. So if AI says it's not DNS, we should not believe it.
It's always dns, but it's always DNS. But I think that the, the challenge there is, I like to call it the Web MD problem, right? Is you plug all your symptoms into WebMD and it comes back and it's like, oh, you have cancer, you have pancreatitis, you have this extremely rare disease that's only ever been seen in four people, but your symptoms match perfectly.
And me, as a person would look at it and go, well, obviously it's smarter than me, so it must know when I go to the doctor, the doctor's like, yeah, there's no chance that you have that. It's actually this other thing over here that WebMD always misdiagnoses or something like that. Um, there, there does need to a little be a little bit of institutional knowledge involved there, but the problem is, is that the way that people are selling me on ai, AI has all the institutional knowledge that I could ever need, so I can just get rid of those people and, and that wouldn't affect a cloud provider having DNS deployment problems, right?
As we record this at the end of October, after we're watching two major cloud outages in two weeks. I, I don't think that's really the issue, Tom, the way I look at it is when we're looking at, and I'm talking about vendor specific solutions, so that would be, you know, nexus Dashboard insights right now and in future AI canvas, you know, um, Marvis, um, when we look at those solutions, those are based upon information that those particular OEMs have had and collected over the years, and it's specific to their infrastructure. And they also have the ability to do pre change analysis and post change analysis.
So you're not just trusting, I mean, you still need somebody who's a, a highly experienced engineer to be reviewing this. Um, but I think there's strong possibility for really providing a better day to experience when you consider that, you know, the purchase of the hardware is one thing, but that is nothing compared to day two running that infrastructure over the next five plus years. Uh, that's the real expense.
You know, I think I, to, to your point, Rita, I think, I think there is something to be said for that. Um, I, you know, uh, pattern recognition, we, we know pattern recognition happens to be one of those things that, that AI generally does very, very well, right? And in my previous role, you know, one of the, the stories that, uh, that I both heard a lot and, and and and talked about a lot was, was this idea of, uh, a very large organization with hundreds of thousands of, of, of, uh, devices, um, that you, they essentially turned on that solution and pretty much the very next day it came back and said, Hey, by the way, here's, you know, 10, I forget the exact number, but here's 10 misconfigured links, right?
Um, because it is very good at pattern recognition, and it was very easy for it to say, based on what we see across hundreds of thousands of these links, here's 10 that are outliers, right? And the, and, and ultimately that does solve a problem ultimately that that fixed a problem for that organization, that it would've taken hundreds of hours, um, to, to ultimately find that if they just simply said, we have to go do this thing, right? Um, so I think there's value in it.
Um, but I think, I think what I think maybe some of the negativity that we're, we're kind of presenting here is there's often this overpromise and under deliver where I think if we could flip that where we underpromise and over deliver, I think, I think I think that would be the better way to handle this. But of course, that's not what, that's not what pays for marketing and, and r and d and all of the other things that go, you know, that, that go into building these products. Um, but I think that, but I think, I think there is something to be said for that.
There are a lot of things that we can identify these patterns that we can identify and that are useful in the environments. Um, and so, and I think often we over, we wanna talk about the really high level stuff and we don't wanna talk about those simple time saving things that make impacts. Well, there's some products out there that do some marvelous things.
I'd like to see a checklist from each vendor of the things that their AI can detect and help you troubleshoot, um, thinking missed here, or Cisco's, uh, some of Cisco's efforts at troubleshoot or AI assisted troubleshooting, but I think no vendor's gonna publish that list because it gives their competitor vital information. That's absolutely right. June, 2025 is when AI Canvas was announced.
I have major companies looking for it, and guess what? It's alpha, so it's gonna be a little bit, so I have to agree with JD that, you know, we're over promising, we're talking about products that are still in the early development cycles, and customers get really excited. I have another customer, uh, very, very large financial, that gave me a laundry list of everything they wanted a product to do.
And I'm like, well, we can look at, 'cause they are a Cisco customer. I'm like, well, you can have Nexus dashboard can do 80%, it can't do the other 20%. Um, but they're talking about growing their own, building their own well, but would that be the better route to go better or should we wait for AI canvas and see if it does the other 20%?
So knowing what is coming and having that roadmap, uh, far out is very important. You triggered me on, I was helping a large company evaluate network management products that had way too many, and it wasn't clear that, uh, those beginning with the letters HP were actually solving any real problems. So I tried to get them to list what was important to them and, uh, couldn't get any participation to speak of how are you gonna evaluate products if you don't even know what's important to you?
Yeah, yeah. That, that brings up a whole nother kind of slightly off topic, the whole HP Juniper merge. Where's that gonna win out when it comes to the net AI ops?
Is it Marvis? Is there, you know, what, what's it gonna be Blue Lake? Yeah, I, I, I, I'm would never, um, uh, make a guess on that one.
Um, uh, I do know that there is one product that's shipping with AI and has been shipping for many years, and then there's another product that's been promising to ship for, um, well, this version for three years and still hasn't shipped, uh, and still can't be deployed to customers. So I, I, I would make a guess there, but it would purely be a guess. But I do think, I, I think that that really does resonate on the issue that we're talking about here, where if we get, if we start, if we start making promises that aren't deliverable within a reasonable amount of window, um, from a time perspective, by the time we, by the time the, the, the, the promise is delivered, uh, customers and and prospective customers have an assumption that, well, that's been talked about for so long.
Everyone's, everyone's gotta be doing that, right? And it's really hard to begin it. It's hard one to go in and have that conversation with them and go, this is why it's important, um, and this is why it matters.
And oh, by the way, yeah, it's been talked about for 18 months, but yeah, no one's actually doing this yet. It's had to hard to have that conversation, but it's also hard for the customer to know what is, what, what is out there in that matrix that they can check the box on. What is, what are, what's in the matrix that goes, okay, in, in this, in this list of product features, these are the things that matter to me because these are the things I can buy today.
And, and I think most, I think most customers don't actually have a, a good handle on that. So I wanna give Rita a chance to bring up one of her challenges with net AI ops. Having heard, you know, what we've seen so far, Rita, what, what's something you think people should be concerned about?
People could, should be concerned about? What is really available today, as we've kind of discussed? Um, there are multi-vendor environments out there.
So consider, you know, is a OEM specific solution going to provide for both campus as well as network and wireless, um, so data center wireless and networking in a multi-vendor environment. Um, that's, I'm not seeing that as of yet. So I think that's a big concern.
Uh, I, and you're absolutely right, but look at what the lead is on that. Because if you look on the other side of the fence in the AI market, um, you've got companies like Nvidia who have a completely vertically integrated stack, right? You order NVIDIA servers running Nvidia Spectrum X or InfiniBand and NVLink and everything just works because it's all one company.
And why would you ever want to integrate anything else with this? Because ours is so much better. And then that becomes the blueprint, right?
Well, we have ultra ethernet out there as an option. Um, you know, companies want to kind of push their own flavor of it and they want people to do that. And that becomes an operational nightmare because, you know, we can barely get OSPF to talk to itself to, uh, across different vendors.
GEAR and OSPF has been a standard for 20 something years now. Uh, and, and we're kind of on the, the forefront of things. And if you thought proto protocol development and prototyping was, you know, fast in the old days, you know, we're looking at things like model context protocol, which is on, they're not less than a year at this point.
And everybody is just on top of it saying, oh, this is the way that we're gonna do this and this is how everything's gonna operate. And I remember when APIs were the way that people wanted to do that, and that's a developer side of the fence. We're not even talking about the operational side of the protocols where people are like, well, why are these two things not talking to each other and, and how am I gonna get them to work?
And why do get, do I get rich statistics from this side of my network and you know, the bare bones from over here? When you go to the company and say, I need this to work together, what's their response? Are they gonna say, yeah, we'll bend over backwards to work with you because you're one of our most valued customers?
Or are they gonna say, well, for what it would cost us to develop that solution, you can just go buy some more hardware? Well, when it comes to Nvidia, um, Nvidia and Cisco are partnered now. Um, so we're integrating the GPUs inside the Cisco servers for the AI pods.
Um, so having that integration happening at the OEM level I think is really, really valuable, really important. Um, I agree with you. If it were my data center, I would not be mixing different, different OEMs in the data centers.
But you do see in some, um, in some environments you'll see one data center might be Arista and one data center might be Cisco. Um, Cisco, I personally wouldn't do it. It's too much of a management challenge, like you said, uh, working with different vendors.
Um, you also see companies that are migrating. So they might be migrating from Arista to Cisco or to Juniper. And so they've, in that migration process, they do have multiple OEMs could be even in the same data center.
So yes, it's not ideal, but we see it every day. And I thought, I saw some stats that the ultra ethernet, uh, se future seemed to be, uh, attracting more buyers recently than, uh, the NVIDIA proprietary technology. But all standards have that, right?
When the new standard comes out, people are gonna wanna adopt it and build around it, and then when the next leap comes out, then they're gonna wanna flock over there because, you know, 15% in performance increase could mean the difference between paying a huge cloud bill this, this month and paying a slightly smaller huge cloud bill this month. And Pete, if you're talking about InfiniBand versus ethernet, even traditional ethernet, you know, 400, 800 gig ethernet, you're gonna see performance unless it's an absolutely the biggest AI environments, your GPU to GPU traffic, ethernet done plenty of testing on that where ethernet has just as good performance. So, uh, I don't think we need to stick with, uh, the one-offs anymore.
We can go ethernet and then ultimately ultra ethernet. I think that's why we're seeing the big uptake on the ethernet is it's multivendor. As you can see, there are a lot of questions that are brought up when you start talking about the world future of net AI ops.
Uh, you have to know the costs that are associated with it, what is it gonna cost you in order to be able to deploy it. You have to understand whether or not someone in the room is smart enough to understand when what it's telling you isn't gonna jive with everything. And then of course, no network out there is completely greenfield or completely single vendor.
And if you don't have an understanding of how to interact with all the various parts of your network with a solution that knows how to talk to them all, you could find yourself with less information than you started out with. That will just about do it for this episode of the Tech Field Day podcast. We wanna thank you all for listening in.
If you enjoyed this discussion, please subscribe on our YouTube channel or your favorite podcast application. We don't want you to miss any of our episodes. And if you do subscribe, leave us a rating and a review so everybody out here knows what we're all about.
This podcast is brought to you by Tech Field Day, the Home for IT experts from across the enterprise. It is a part of TUM Group. com/podcast or check us out on Techstrong tv, including the Techstrong tv, over the top application.
Thank you very much for listening, and we'll see you next week.