69. Datacenter Networking Needs AIOps with HPE Juniper Networks – Tech Field Day Podcast
Transcript
Enterprise networks are huge, particularly since customers and our staff are accessing those networks across the internet for that huge complexity. You need ai, you need ai, managing, operating, giving some oversight on your entire network. Join me on the Tech Field Day podcast as we dig into the need for AI operations and Enterprise Network.
Welcome to the Tech Fields, a podcast. We'll bring together a group of IT experts to discuss a single idea around key concepts in the industry. This podcast features a variety of perspectives from members of the tech fields I Delic community, and also from the companies that support us by presenting a tech field day events.
The podcast is often a recorded in association with one of those events, and Tech Field Day is part of the Future Group. This podcast is also published on our sister company's website, tech Strong tv on the spotlight episode presented by June of the networks. We discussing the premise that enterprise networking needs AI operations.
But before this discussion, let's meet who's on the panel today. Hi, Alistair. I'm, uh, Jack Poller from Principal Analyst with Paradigm Technica.
I focus on AI and cybersecurity. That, and thank you for having me. Alistair Bob Friday and I'm chief, a officer at Juniper Networks, or HPE now.
Hey guys, I'm Ben Baker, data Center Marketing, HPE Juniper. Uh, I've been with Juniper for, for over 13 years now, in, in the data center world. A little, uh, a little over four, and I think data center is the hottest area in all of tech.
So it's a great place to be and glad to be here. And of course, I'm Alist Cook. I'm an event lead at, uh, tech Field Day.
And, uh, well, congratulations. First off that finally the, uh, DOJ has gotten outta the way and you've been able to close that acquisition by HPI think the, uh, combination together is gonna be absolutely awesome. And I think one of the overriding themes we've seen throughout that data center management and network management is that the massive amount of complexity that's out there means that it's not no longer tenable to manage everything through spreadsheets and iff looking at the configs on switches, um, hopefully we are a very long way away from my horrifying experience of network management back in the 1990s.
Um, but we really do need things that are more reactive and responsive. Policy-based management is awesome, but we really want to be seeing some more, uh, insight into changes and variations of what's going on in our environment. And that does tend to lead us to AI for all of the things.
Bob, you are deep in the ai from the Juniper side, what are the big challenges people are, are seeing and why is AI helping them, or how is AI helping? Yeah, I know Alistair, for me personally, you know, I used to be the CTO over at Cisco and, um, I was actually trying to sell a big, large retailer, this consumer mobile experience thing for their network. And they basically told me, Hey, Bob, we're not gonna put anything critical on this network, you know, until you can promise me that your controller stopped crashing, I had to make sure that we're gonna deliver code more than once or twice a year.
Uh, but more importantly, they told me, they said they had to make sure that that client to cloud experience was gonna be good. That was the big thing, you know, and that's when I kinda realized there was a fundamental paradigm shift from, hey, in addition to just, you know, making sure switches, aps and all that stuff was up and running and green, you really had to make sure that user experience was gonna be good. And that was kind of the beginnings of realizing that cloud AIOps was gonna be an architectural change in the enterprise.
And so that was one reason why I left Cisco to start Miss, was I knew there was gonna be this real time problem of how do you actually maintain that user experience on a very complicated network? You know, you brought up the nineties, uh, Bob brought up, uh, you know, going back a little bit, and I think that's one of the keys to think about is talking about delivering the user experience. We as, as IT professionals, as network professionals, we get very bogged down in the details, the technical details.
And for some people that can be a lot of fun for other people. It's just the, the, the nature of the beast, what's ultimately most important is making the network invisible so that we just have the communication we need between the, the, the entities in our environment, whether that's people or AI systems or data processing systems. And if we can use AI to do that and get people out of the way, I think that's the big win.
Yeah, I completely agree. I mean, so many technology discussions start in the wrong place. It starts with a technology and, and that's not the right place.
You know, we've all heard these edicts from above u use ai, right? Um, you know, how are you using ai? Maybe you're hearing it from your, your networking vendors, you know, you need to use AI in network operations, but it's a pretty aimless starting point and, and agree with the, uh, with Jack here and, and Bob, you know, the starting question should be, what are your goals and what are the problems to be solved?
And then you have to ask yourself, can we use AI and machine learning to help you solve those problems? And I think we're all aligned with the goal. It it's to optimize the user experience, you know, whether that's the, the network operator or that end user who's relying on that infrastructure for the, the applications that they use every day.
Yeah. Know what I usually tell people is though, AI is really just the next step in the evolution of automation. The only difference here is we're starting to automate things that typically would take a cognitive human to do, right?
I, I, I dunno, have, have, have any of you guys tried this new Waymo Uber self-driving experience yet? You think about, you know, ai, you know, when you get into that self-driving Uber, you know, if something slightly different. Now what I typically, you know, ask people is like, you consider that self-driving car automation or is that a, an example of great ai?
Yeah, my answer would be that there, there's no either it's, it's a both that the, the good solution has a foundation of some automation underneath it. And as you need more than simple automation, you start doing some of the ML kind of functionality of looking for trends and looking for over time, what should I be expecting? What should I be ex, um, looking ahead to?
And then you move a little deeper beyond that to what are the unobservable pieces. And that's where the, the large language model stuff starts to come in because it's seeing patterns that are beyond the capability of a human to comprehend and, and still trying to take action in a way that is more human-Like, there's a bit of a dichotomy there. Yeah, it, and I would say also the interesting things, right?
You know, when people say ai, now they think gen ai, right? Large language models, you know, but there's actually two components to it, right? There's kind of the classic training models to do something.
And that's what we have in our cars, right? Con you know, a lot of vision stuff in networking space. It's about training models to predict user experience.
And then I think ingen AI is kind of a new non-deterministic, non-linear programming language, you know, that is gonna help us really automate things in networking. Yeah. Well, I, I look at it and I think the last time I was hands-on with network equipment, I had an, uh, I was at a startup with 50 to 60 people, and we had a, you know, a single Cisco chassis switch with 6 48 port line cards.
Now, that's a complicated but not complicated environment. It's complicated when you're one person who doesn't know what they're doing working on it, it's uncomplicated in comparison to today's networks which have many, many, you know, hundreds of devices and thousands to hundreds of thousands of con connections. And we get to a point where just as with a self-driving car, the scale is beyond what a human can hold in their head and think about it one time, right?
And I think that's why, for me, that's why AI is so important, is because we can't, I, I just can't comprehend all the different, where all the different pieces go. We just can't keep it in our head. So how do you, how, how do you manage that if you can't do that?
Yeah, and that's why I tell, you know, the average IT person, if the signal's big enough, they don't need ai, right? If it's obvious, the, the DHCP server's down, I'm not sure they need AI to tell 'em that, you know, the whole building's down. But, you know, there's a great example.
We had a customer in India, they're having all types of Zoom problems, right? You know, people were coming in, they're getting all these Zoom complaints, everything. We were finally able to build them a model that could actually predict Zoom experience.
And they were able to use some fancy math like Shapley to actually figure out that what was happening. They had a, a misconfigured VPN router that was routing traffic from India to Australia, right? That, that's kind of the needle in the haystack problem.
Where, where I say the fancy math helps, right? Every now and then, that fancy math does help you find those needle in the haystack where you have multiple things going in a complicated network now trying to figure out what piece of the network is affecting user experience. And there's some real challenges with that end user client connecting across the internet, uh, where there's so much of it that you don't control, that you can only observe.
So this is why I think, Bob, you've, you've spent so much time on the, the client to, to cloud, uh, because you, you get massive numbers of clients connecting through various layers that are shared, not shared many of the layers like, uh, that the, the router in India is not being shared by people who are in Singapore. Uh, but other elements might be shared in, in their access across. And that complexity of having so many different endpoints and so many different path back into the server is one of the elements that gets to be really challenging to manage.
And those sort of, the subset is experiencing a problem, but I don't know what the pattern is that's causing this subset to experience the problem. That's the kind of use case where the, the analysis is so vital they automated analysis. Yeah.
That I would say the other interesting thing is, you know, when I talk to people about AI and where we are in the journey of, you know, do it, people trust AI now to drive their network. And I would say the interesting thing, right? They, the same person who doesn't trust AI to drive their network, they will definitely get in that self-driving Waymo Uber.
And so somehow they figured out how to trust that car to get 'em from A to B. So I think we're in that same journey of, you know, how to get it people to trust letting, you know, some AI fancy software start to, uh, play with the data plane configuration. I, I, I agree with, um, that the, the underlying problem with this I'll, I'll call it an epidemic, is, is complexity.
Um, and people are afraid to touch their networks. I mean, the network engineers, the operators, um, they're stressed out, they're nervous before, during, after a change, whether it's provisioning a new service, a firmware upgrade or whatever, it's stressful to them and they're usually we're drowning in data, but, but starved for, for insights. And it's, it's that complexity that's really causing that epidemic of, you know, just fear when people touch their network.
So there's definitely an appetite to, right, as Bob's saying, start to hand over the reins to, to ai, you know, something, something better to, to help rescue them. Well, when you talk about handing it over the reins, Bob's example was great about this misconfigured router, right? Is that if we are still operating in the, and I hate to use this 'cause it's such a trite analogy, but it's still applicable, the pets versus cattle analogy.
If we are still in the world where we're hands on to every device in our environment, when we get two thousands of devices, you will have somebody who misconfigure something. It's human nature. Mistakes happen with ai.
And we can go to intent-based management where we say, you know, I want this thing to do this, and we let the ai, which is the expert, configure it appropriately, and then we can get away from the don't touch my network because it works now. And if I make a single change, the whole thing comes crashing down. Well, I, I think it's a trust thing as I'm saying.
Like when you get into that self-driving Uber, for some reason they trust it. 'cause they've heard enough stories about, you know, hey, that self-driving Uber is actually safer than a human. You know, I think that same process is gonna happen in it.
You know, we, we will eventually start getting these tools to the point where it actually figures out that it's easier to trust that tool than it is a person who probably makes more mistakes than the tool does. Now, the interesting thing with these tools, they're deterministic. 'cause the other big paradigm shift is we're going from where we used to write scripts for API where our enterprise customers are starting to write agent codes, right?
MCP, they're starting to write this agent code, and they're starting to learn this kind of what a data center is on the other side. 'cause the data centers were eight, they're almost like a X 86 on the front end. And, you know, GPUs on the back end now, you know, and as you start to use these gen AI agents, you start to figure out the cost of, you know, how much is gonna cost per token to have my script automate something.
Yeah. The the cost question is an interesting one. And we look at that and I think I, I feel you're sort of tiptoeing along to the, the problem we see in the cloud of, well, the cloud is cheaper, so we move everything to the cloud, but then we, the cloud becomes more expensive because we're doing so much.
And we, you know, there was a non unpredictability to on operate on OPEX model versus a CapEx model. So people all of a sudden said, oh my God, I'm spending so much on the cloud, I have to pull that back in and repatriate it. And a lot of that is because they don't realize that in a, that the expense is going up because they're doing so much more than they could before, right.
To it, right? Once you get, once you, once you start writing these Gen AI agents, you get addicted. You don't realize how many tokens you're actually using to solve a problem.
But yes, so, so the cost is gonna, is gonna surprise people. You don't realize how much you're using it. But also, are you getting, is the ROI there, are you getting the, the, the advantage back from doing it, automating it with AI versus paying an engineer to be the expert?
Yeah, and I think, I mean, I think the other thing for those who watch it, right? The, the cost of a token is coming down. 'cause the, the data center guys are getting really good at optimizing these clusters now, right?
In terms of how they train and even the inferencing of these big models. And so that's the other piece of the puzzle right now, you know what I call, you know, AI for networking, networking for ai, networking for AI teams are starting to get better at what they do. And in part Bob, because of AI for networking, right, Right, right.
It's driving it, right? AI for networking is starting to drive the need for more networking for ai. Yeah.
I think the other thing that people don't, don't, don't sort of realize is that, you know, I talk a lot in, in, in security about data security. And that data resides on a physical medium. It's the, it's a mechanical part, right?
A spinning disc and networking is all digital except for there's mechanical component to it, which are these cables that get plugged in, pulled out. And that's the root of how many problems would you say is a cabling problem where AI can somehow find that We have that? So right now, I, you know, so we got models right now that basically find bad cables, you know, and so when people actually get this stuff up and running, they start to see how many bad cables are in their network.
They don't believe it until they actually start go checking them. And they find out that model is probably 90% correct. You know, you know, they find out they've got these half duplex, you know, half the pair is gone.
It's not negotiating a gig anymore, it's negotiating a hundred megs instead of a gig. So that is where they start to trust AI that I've seen that happening now, right? Once it gets that feel for it, then they say, please don't ask me anymore.
Just issue the ticket for that bad cable. And that, that's like adopting any new technology. I, we, we start with just tell me the, the conclusion so I can judge your fitness.
And we do this with junior engineers. We don't send a junior engineer to go and build a whole new infrastructure the first time we watch them do it. Uh, in the same way with, with AI and any other technology, we look first for reports and tell me what you think is wrong, tell you, tell me what you think is the remedial action.
Uh, we often start by manually taking the same remedial action if we have any faith in it, but eventually we get sick of taking the same action. I certainly get sick very fast, taking the same action, and we'll hand it over to automation. So yeah.
That, that trust has to be earned. I thought you were gonna say we send that junior engineer to go replace the cables. Yeah.
Well, at the moment we can't send a, a, uh, an AI bot to go and, uh, replace the, the cables, but there's a whole mechatronic side to, uh, to agents for the future. Well, y you know, it's a funny thing. I I'm sure you've tried chatt PTI would say chatt PT is on par with probably a junior IT person.
If you ask to configure, you know, gimme the configuration for a Cisco router, juniper router or something, it tends to get, it's tend to getting pretty good. You know, I suspect if you took a AI assistant and had to take the CCIE test, you know, it may actually get pretty close to passing that thing. Yeah.
If it's well known, lots of prior art. If it's something that's, um, that's well trained in. But that's, that's your general knowledge.
Uh, junior engineer, actually an incredibly well trained junior engineer. It's probably somebody who has had a, a previous career and transitioned in with the amount of general knowledge that these, uh, live language models have. But it's this specialist knowledge over time that you're building into your products that's understanding how networks operate and where, where the dragons are in your network.
And by the way, it's always DNS, uh, it's never the network. Um, you know, that that's, it is very different using chat GPT to give general purpose, general knowledge answers versus giving a, a fine tuned, well-trained specialist knowledge, uh, large language models to give you answers. Yeah, I would, I would say the interesting thing, the one thing I found, you know, in addition, when I started Miss, I knew I had to go build this real time cloud architecture to kind of solve this day two problem.
Uh, the other thing I found is I had to really take my data science team and, uh, support team and time to the hip. 'cause it turns out that the support team is the domain expert. Data sciences are great, but they really don't know the support problem.
And that is where the, uh, I found you really have to get that domain expertise tied in with the data science. And I think you'll see more organizations do that. Your support team is your domain experts if they, they know more about what's going on in the networking than anyone else in your company usually.
Yeah. Particularly the, the networking When it goes wrong, When it goes wrong, right? I mean, they deal with the real problems, you know, they know it better than the engineers know it.
Actually, One of the interesting things is, you know, you think about chat GPT and there is, you know, from a cybersecurity side, there is a security issue of course, with exposing your data, you know, a general enterprise data question to chat GPT, I would think that we also wouldn't want to expose our entire network configuration to chat GPT, which is why a, a specialist, uh, AI solution for AI net AI for networking right, is, is so important because it is gonna give you the, uh, the, the data protection, the security protections that you won't have if you just go out to a, a general purpose LLMI think most vendors, you know, most vendors like ourself or, you know, if it's public documents, they're okay with sending that to Azure open ai. Uh, but if we're talking about customer data or customer configs, that tends to have to stay in the BPC that cannot leave, that typically cannot leave the, uh, the data center. You know, you know, you cannot let customer data outside of your core Y Yeah.
And I, I think you guys are, are hitting on a, a really important point. I mean, there, there's, you know, so many real AI ops examples in, in data center networking and, and other domains to today that, you know, around predictive maintenance. We got our AI assistance, service level expectations, application assurance.
We have a ton of capabilities that, that are here and now, um, that have essentially been developed by, by the vendors like, like HPE Juniper. But I, I think the most important area of AIOps, um, in, in the coming years is gonna be just simple experimentation. I mean, these LLMs are these magical beasts that even the, the people who have designed them will even admit, they don't understand always the intuition behind them.
So it, it's important to just play experiment. And a lot of that experimentation needs to come beyond the, the regular corporate product development roadmap process because A, it's too slow and, and b, kinda what you guys are hitting on, there's some vendor liability issues. Um, you know, do I wanna put my name behind a customer, bringing their own LLM and, uh, tying it into my enterprise application, my, my network automation and, uh, management tool?
Probably not, but I do want that customer doing it because it's customer led innovation. They're gonna do some amazing, perhaps unexpected things. And I, we, we've seen time and time again in history, many industries, um, you know, you put the tools in the hands of your customers and, you know, they're gonna lead a lot of the innovation.
So I'm really looking forward to just what people in their garages are, are gonna be doing outside of that formal corporate development, product development process. I would, I would say, I, I would agree, Ben, you know, as I said this gen ai gent ai, this is a non-linear, non-deterministic programming language. It's very iterative.
It is totally different than the days of, you know, software engineering where I had a very deterministic Give me what, tell me what you want and I'll write you a, you know, it's taking, you don't know how long it's gonna take to iterate yourself to a working solution, and it is non-deterministic, so it's not a hundred percent. So, so let's think about the non-determinism a little bit. So LLMs are non-deterministic, whereas the, uh, the AI analysis engines and the other parts of AI are deterministic.
But when you go to design an AI ops, how do you deal with the non-determinism in the, the, the tools that you're building and present that to customers? What do customers think about that? I I, I would, I would say, you know, it's like hiring your new employees.
How often do your kids or some new employee does something go, what the heck were they thinking? You know, that's the point we're gonna get to. The point is, you know, these, uh, AI assistance aren't gonna be perfect, but they're, you're gonna get to a point where they are more perfect than your human counterpart.
You know, the deterministic piece will be actually less happening than you would with a, you know, an actual person. So I don't think you eliminate it all. I think you put guardrails around it and everything.
You try to get to 99%, you know, you want that, uh, AI assistant to be better than the actual human, the alternative. And you, you do things like grounding the results and, you know, make making sure that you are, uh, doing things like feeding the LLM UpToDate information about what is possible with the products that you've got deployed and what your policies are. And, and in that way you can end up with an l LM that is better informed than your average human, because the average human is gonna get bored reading that document and be thinking about their doing what they're doing at the weekend when they should be studying closely.
And so I think, yeah, Bob, be right, humans are not deterministic. I think that's one of the things that you gotta remember. We aren't It, but you do want that deterministic, like that self-driving Waymo, right?
Yeah. That is a very, it may make a mistake, but it's very deterministic from getting you A to B 99% of the time. Well, you know, I think that you, you go back to that, uh, that example, and I think that's a very apt example.
Um, the determinism comes in, in, we expect Waymo and Uber and the, like, on the self-driving cars to never have an accident. And, you know, every time they have an accident, there's a big, you know, big blast in the press, right? The car crashed or whatever drove down.
The last one I saw was it drove the wrong way down a construction road, right? So is that's, we can, we can put that type of constraint and, and have that expectation because there's a human life at stake when there's not a human life at stake when it's just the network. Does it have to be, and I don't think it does have to be a hundred percent perfect.
I think it, it has to just be better than what we as good as, or better than what we have today. I agree. And I think that's the same thing with the car, right?
I mean, you know, those cars are not, but you hope their accident rate is better than the human drivers. And so I, I did have a thought about, uh, that, that Waymo and the human driver is non, non-deterministic. If you've ever had a, a back and forth with your, your driver, uh, about where they're picking you up somewhere downtown or at a hotel or at an airport, uh, you can understand that, that there's non-determinism is acceptable in, in some places, right?
Whereas, as you say, that deterministic guards a bit of a pedestrian steps in front of my autonomous vehicle, the autonomous vehicle better, better stop. Uh, that, that, you know, there are places where deterministic is, is important. There are places where non-deterministic, uh, as you say, slightly better than a human on average, uh, is, is actually a, a, a really good outcome.
And if we can continue to improve that, I think the concept of iteration and there is also vital that getting a suboptimal result on the first iteration through is not so unusual. But if so long as we can start improving the result on each iteration, uh, and that those iterations maybe are even fully automated. So every five minutes we'll get closer and closer to the solution that we want.
So long as we're tuning things well enough that every five minutes we're not oscillating off into something that is significantly worse, uh, you know, that non-deterministic behavior doesn't mean that we can't approach the correct solution. No, I'm with you. But again, on my first Waymo Uber, I did have that feeling.
It's like I was okay getting in, and then I started thinking about it. It's like, where's he gonna let me off? How's he gonna figure this out?
I think if you, if you come back to sort of the networking side of the world, part of that iteration is having explainability, right? Is if, if the, the AI operations tool for networking, if it's a complete black box, it becomes a challenge for the early adopters, right? And they don't, they don't have an ability to say, to iterate and or provide correction to say you're thinking, you know, you went down this path and you took a left turn here and you ran into a brick wall for the card analogy, right?
When you should have turned turned right instead of left. And so, you know, ensuring that, I think we have to ensure that we have that explainability component. We can understand the chain of reasoning that an AI goes through, uh, as it's making its decision.
Otherwise, it becomes very hard for users to become, adopt, you know, to really buy into the technology. No, that's a given. Observability is key to trust.
You know, people need to see the evidence to earn, you know, to earn their trust. They need to, you need to show your homework. There's also the, uh, supply chain trust as well that has to be built out.
As we're recording this probably some time ago, as you are watching this, uh, relatively recently, there was a news story about the copilot ai, uh, that's, uh, has been attacked. Basically it's an open source, uh, plugin that goes into this. This was an open source plugin to go into, uh, visual Studio Code and a bad actor, um, put some code into it into the public repo that would delete your entire cloud estate.
And so governance of the supply chain for your AI is gonna be absolutely vital to acceptance and trust and enterprise organizations. Yeah. Yeah.
I think it's gonna become like open source code. I mean, people are gonna start looking at where'd you get the data to train it? Did you have the right to use it?
You know, you're right, there's gonna be a lot more, uh, like software bombs, you know, you wanna know exactly, you know, what's in that, what, what's in the components, what made that AI work? It's like the label on the back of your food, I want, yeah. I wanna know what's in that in that assistant.
Yeah. And it's, it, it, it's mostly a good thing that, um, I mean, you know, I don't know about anyone else, but we're all buried in, in AI and the information and, and the speed at, at which it moves and, uh, you know, we're, we wanna be informed, but not overwhelmed. But at the same time, everything that the information, the tools seems so democratized, you know, you can go to hugging face and download just about any AI model for, for free the tools to prepare and organize your data for free.
It's all there for people to use. But as you guys are pointing out, the flip side is wow, I mean, talk about some attack surfaces being, being huge and, and major, uh, security issues with that. So it'll be interesting to, to balance those two things.
And, and again, that's where, you know, we, we talked earlier on about, you know, having some guardrails in place and that's where guardrails are, you know, and policy is very, very important there. There's gotta be some policies that say, you know, you can't just do a delete or a delete. All right?
An RI of RF all for all of us old timers, right? Without having something saying, are you really, really sure that you want to do this? That's right.
We'll hand it from one agent to another to do that. But before we get down that rabbit hole and spend a long time talking about governance, I think it's about time we brought a close to this episode. So thank you all for joining us today on the Tech Field Day podcast.
But before we go, where can people connect with you, my guests, to continue this conversation? Uh, I'm Jack Poller. You can find me on LinkedIn most often.
That seems to be where all of the enterprise it technic Sessions are today. Yep. And, uh, Bob Friday, cloud AIOps topic, dear to my heart, feel free to reach out on LinkedIn.
Ben Baker, HPE Juniper also. And, and again, I'll uh, throw my hat in the LinkedIn ring too. And I'm Alice lead here at Tech Field Day, so you can find me on LinkedIn.
You can find me on all of the Tech Field Day estate, as well as you'll find lots of, uh, presence for us for Jack as well across on the Techstrong site. And, uh, Jack's been providing some awesome content across on Text Drop. So thank you for listening to this episode of The Tech Fields, a podcast.
If you enjoyed the discussion, please subscribe on YouTube or your favorite podcast application so you don't miss an episode. And do consider giving us a rating and a review 'cause that helps other people find this awesome content. This podcast was brought to you by Juniper Networks and Tech Fields, a home of IT experts from across the enterprise and part of the Futurum Group.
com/podcast or view us on techron tv. Thanks for listening, and we'll see you next week.