Root Cause Analysis with Nokia AI Operations Automation
Clayton Wagar introduced Nokia’s AI-driven approach to root cause analysis, focusing on solving difficult day-two operational challenges. The presentation highlighted the chronic pain of hidden impairments or gray failures, where traditional monitoring systems fail because physical links appear active while protocols or services are down. The goal of Nokia’s deep RCA tool is to move beyond simple port-up/port-down alarming by correlating end-to-end application connectivity (from VM to VM) with all layers of the network, including the underlay, overlay, and control plane, to dramatically compress troubleshooting time.
A live demonstration was shown on the EDA SaaS platform using a real hardware Spine-Leaf network. The team introduced a gray failure by impairing a fiber link in a way that kept the interface status “up” but caused the BFD and BGP protocols to fail. Wagar explained that the AI’s multi-agent workflow correctly diagnosed this. Instead of using one large, monolithic model, a planning model first determines which tool-calling agents to deploy. These agents gather specific, relevant data from logs, topology, and configuration, which is then filtered and passed to a reasoning model. This agentic-based curation of data is Nokia’s key to reducing costs and avoiding the AI hallucinations that would otherwise be a risk in mission-critical networks.
The tool’s capabilities were further demonstrated by successfully identifying a classic, hard-to-find MTU mismatch. Another key feature highlighted was the “Time Machine,” which allows an operator to select a past timeframe, such as thirty minutes prior to an event, and run the same AI-driven root cause analysis on the historical data from that moment. The entire process concludes with the AI generating a comprehensive report that provides a human-readable summary, a confidence score, and the specific evidence gathered by the agents, effectively solving a complex logic puzzle that would have taken an engineer hours or days to manually diagnose.
Presented by Clayton Wagar, Principal Consulting Engineer. Recorded live at Networking Field Day 39 in Silicon Valley on November 5, 2025. Watch the entire presentation at https://techfieldday.com/appearance/nokia-presents-at-networking-field-day-39/ or visit https://techfieldday.com/event/nfd39/ or https://Nokia.com for more information.
Transcript
I'm Clayton Weger. I lead our AI activity inside of our cloud and enterprise strategy group for Nokia within our IP division. Hi, I am, uh, I am just taking care of the senior, uh, PLM for the Nokia here.
So Rupa, and I'll be showing you the, uh, the SaaS version of eda. So we talked about edda has a number of different deployment models from on-prem, kind upon how you want to integrate it within, into your existing, if you have an existing Kubernetes infrastructure. Uh, what we'll be looking at now is the SaaS version.
So this is deployed into public cloud. It's used by enterprise customers of ours to manage their networks on their own premises. And so what you'll see is a lot of DNA shared between the IDA on-prem version and the EA SaaS version.
The user interface, user interfaces rhyme, but they, they don't match exactly right? So you're gonna see some differentiation there. That just gets to the, to the deployment model, right?
So we're providing a lot of ancillary services over and above that you might stitch together if you have an on-prem version based on your own requirements. Here with the SaaS version, we're obviously providing much more of a bow around those things. So there's things like a data lake and other portions that, um, that, that form the whole of the SaaS product.
So, um, when we look at sort of the key things for day two operations and a data center, what we're gonna be focusing on today showing you is app connectivity from end to end. Um, one of the key pieces that you wanna take a look at, obviously alarming, some of the earliest network automation was really around log collection, alarming and discernment of what's, what's important inside of a series of events. That's part of it performance as well.
So sometimes you have sort of hidden impairments that may not be obvious, uh, from the, the logs that you have, but you know, you have application degradation and so you have the performance element as well, of course the change request portion of that. We talked about that quite a bit, that we've had some engaging questions here talking about things like configuration drift and comparing between two different states. Um, I know my very first question whenever dealing with something that's changed is, what did you do?
What changed? Right? We always wanna get to that origin event.
Um, I'm not gonna drain the slide. We have it here and we wanna make the most of our time. But just the idea here is that there are any number of things that we can be collecting from the network if we're doing this.
Certainly in a an all manual operation, we have to go to potentially every different type of role of device in the network. We have to understand where the, uh, where the network is, is transiting over a large network, what, uh, what sub-component it's transiting. Uh, we have all of these different types of pieces of instrumentation that we can collect.
And so therefore in a purely manual model that becomes, uh, you know, uh, pretty much gated by the ability of the human to operate that. What we're gonna show you today is some AI and some non-AI tools, which will hopefully compress that, uh, down to your root cause analysis, getting to something much quicker than you otherwise would have. Um, last slide.
'cause the demo is the best part of this, but we wanted to sort of highlight some key pieces of what you're going to see completely model driven, right? So there's no sort of fancy scripts on the backside that are doing. Some have some sort of embedded logic.
All of this is based just like edita itself is based on the idea of being model driven. Um, when it gets to reliability, we have two pieces here, the reliability and the accurate reasoning. So as you look to implement AI tools on top of this, we have to be cognizant of number one, yes, hallucinations and context windows and accuracy of the model.
But in addition, we can't just be consuming AI tokens like crazy, right? If we put AI in everything, we're ultimately gonna be paying a big price for that if it runs in the cloud. And so we've put a lot of thought into the efficiency of gathering the information, again, being, uh, discerning what's important about that, doing the filtering that's necessary, and then reducing the token count that helps us with accuracy and also ultimately helps us with cost, right?
You'll see a, um, uh, what we have here is a multi-agent workflow. So we have multiple models that are working here. We'll walk you through exactly how those are, uh, interacting.
Some are doing tool calling, right? And some are being used for the planning purposes and reasoning purposes. So, uh, with that, I think the best part is to get the demo started.
What we're gonna show you is, uh, an actual hardware, real spine leaf network. It's a very large network. We're gonna show you, uh, sub components of it.
There we go. So we've got a classic, uh, three tier here. We've got a CLO and we're gonna introduce a network impairment here.
In between, uh, two of our application servers, we're gonna introduce a network impairment. We're not gonna do that. Uh, again, sort of drawing on what we've seen already.
Not gonna do it at the configuration level. We're not gonna do it by downing a port. We're actually gonna go and, and, uh, impair the fiber so that we don't have light down on either side, right?
So this is sort of a, a bit of a hidden impairment that doesn't necessarily start to cascade other pieces like BFD and the protocols on top of the fort. So we'll have to go in and do a root cause analysis on exactly what happened with a backhoe, right? Or if, uh, or someone tripped over the fiber in the network.
Okay? Uh, so we have just, uh, uh, gone to the slide. We have just seen the topology and, uh, the, the accent behind the topology.
We want to first understand where our VMs are located and then we can just try to find out the path between those VMs. So as a network engineer, it is very, very difficult if you have big network to find the whole path between the two different applications, which are just hosted on the server. So the first point, which we have to find out, is what is the path which looks like between the two different VMs here?
So for today's demo, we are taking the VM two and VM three as our source and destinations, and we will try to find out the path between them. What are the possible parts, which can be possible here. I'm not able to minimize this.
Yeah. So, uh, we go to the connectivity diagnostic page where we are just selecting the source, which is a VM two here, and then we will select the destination. So here we will just, uh, list down all the VMs, which are present in your network, uh, from the application perspective.
And you can just search between what VM you want to do, uh, see the path. In this case, the VM two and VM three, both are the layer three domain, uh, VMs. So they're layer twos are different, but they are connected to only the layer three.
So that's why we are just seeing here the IP addresses and the VLAN here. When we run the tool, it'll just try to find out all the possible path from that VM perspective in the whole network. And then it'll give us a very collapsed view of the network.
So this network is a correlated network topology. It is not the full view of your whole network, which may might consist of 50 or 60 switches. So here, if you see you have only 10 or 12 switches, and you have the path from the source VM perspective, that these are the possible paths which it can take if the traffic will be flown from the source VM to the destination vm.
So half of your traffic, um, has been reduced. So it is just a superset, uh, a subset of your topology. If you have a very big topology, you have reduced it to the 10 or 12 switches, and now your job has become very easy to find out the issues from here.
Okay? Now the question might be, okay, what I will do with this particular topology, there is no information on that. So we can just go on each of the module here of the source leaf spine, whatever modules you want, you can just expand it and you can just get the information about all the VMs, whatever information you need.
You don't have to do, go and do the SSH to the VMs. The operator don't have to go there. They can get all the information from here, which is required for your debugging purpose.
This is from the source instance perspective. Now, when you go to the leaf, you will get all the logical objects, which this particular VM will pass through this particular leaf. So these are all the logical objects mapping in that particular leaf.
And uh, from here you can find out the state of all the logical objects, what is the state and what is the value for that particular objects in that leaf perspective. So this is from one leaf perspective, but now the question comes, what if I want for the whole topology? Because then it'll give you a much more visibility that way.
The problem might be, so you can go to the deep dive and you can just open it here, and you will see for the whole topology, the correlated topology, which is for this VM perspective, even for the leaf spine, super spine, everything, you will just get it here. So from here you can find out all the objects modules, either it is on the leaf one or a spine one, or super spine one or the destination leaf. Everything is green.
That gives you a indication that all the path from the VM to the source VM to the destination vm, everything is fine. There's no issues right now in the system. So now let's try to go back and then introduce a problem here.
So, uh, what I'm going to do is I will just try to remove the link between the spine one and super fine one. Um, I just clarify, so you're cutting down a link. I thought you said you were just going to Yeah, and I, I was going to explain that.
Okay. Yeah. So we are shutting down the link in a way that it is not going to affect the interfaces status on the spine one and super spine one.
Okay? So the port status on the spine one and super spine one will still be active only. The protocol which is running on top of that will get affected, okay?
Like the B-F-D-B-G-P-O-F pf whatever the protocols you are earning between those links, that will get affected. But the link status is still up on the spine one and super spine one. Awesome.
It's obviously link down, But our budget ran out, so we now find a way to simulate it. Yeah, I was link up, down is easy. So, but the application stuff is Hard.
It's, it's a gray failure. So, so uhhuh, effectively we don't have protocol activity on there, but we're still seeing light, which of course light would be a trigger to a number of other things. Yeah, that's doesn't make for a great demo, so, exactly.
Um, Okay, so, uh, since we have done a, uh, interfaith shut down, which is not bringing the foot down, now, we will just see the status between the VM two and VM three. And, uh, tool will try to find out where the problem is and what is the problem. So it'll just give us a, uh, view of the whole network where the problem is right now.
So if you see here, the service status is still success, that means the tool is saying I still have a alternate path to go from the source VM to the destination vm. So my service will not get impacted from the traffic perspective as of now. But they, they, or, uh, amber bubble here that says there might be a problem on your redundancy or there might be a problem on your performance, or there might be a problem on your link, which is connecting between a spine one and super spine one, which can affect in the future.
So the network operator has to start looking into that from the future perspective. So now the question is, okay, I, I got to know that, see the problem on the spine one, but a spine one has so many modules which module the problem might be. So we can go there just a moment.
Yeah, so you can go there on the spine one and you can see the next top module has a problem. So next top module is saying that there is some problem due to that the spine one is not able to get the next top. Now why it is not able to get, to get the next top.
That is our valid question, right? So for that, we will just introduce our, uh, DR ca tool. That is the AI based, uh, root cause analysis tool.
So with this tool, we are trying to find out the root cause. So with the fabric inside feature, we are able to pinpoint where the problem lies. Now with the help of DR ca tool, we are trying to find out what is the problem, and then we will try to give you the, uh, probable solution for that, how to solve that.
So when we are in the deep rca, the the first phase is your planning phase. Um, the, so we are running a finetune, a small model, which will take the input of the topology and the config or the, and the connect issues from your tools. And then it'll just try to find out which all agents it has to earn.
So it'll decide if I have to earn all the agents. So if I have to earn only some of the agents, so based on the problematic statement, and then it'll just send that information to deploy those agents. Those agents will try to get all the information, the correlated information, so the agent will not get all the informations, which are for the whole network.
They will only get the information which is correlated and only related to this problem. So by that way, we get a very small amount of data, and then we get that data, we send it to the distill model to further fine tune it. And after that, the distill model will remove all the noise, send it towards the, my large model, which will go for the evening stage, and the hallo hallo nation will not happen in that way.
And then the large model will go for the evening phase, try to find out the patterns based on the patterns. It'll try to do the chain of dots, connect the dots, and find out the evening for us. And here, if you, we have just found here, the problem is, is A BFD failure, which is just creating a BGP ping issue here where we have not told anything to the model yet.
Along with that, we can just generate the report, and in the report you will get the evidence, the root cause summary and what should be the problem solution for you That report. Yeah, Just So while you open that, uh, I'm, I'm curious, uh, sorry, uh, Jonathan Davis here. Um, I'm curious on, you know, ultimately we brought down it, even though it was a maybe a gray failure, uh, we, we still ultimately brought it down.
What, how does this look? If we're the, the link isn't down, we're just losing half of frames, right? Uh, you know, so there's, there's something where we are, we either frames are being discarded or dropped along that link.
So, so it's still up itch. What does that look like? So, um, let's take a look at this report and then we'll look at the performance piece in there and we'll give you a sense, there's another module there for performance as well.
There's 2, 2, 2 more things I wanna show you, but let's get this, okay, this knocked out, but it's a insightful question. Question. So We're headed, okay, so here we are just providing what are the instances which are getting affected by this particular problem.
And, uh, this is the root cause, detailed root cause where it is saying it is down due to the control detection timer expired for the BFD, which brought the BGP down, and which is creating the problem in the underlay. At the same time, if you see here, all the configurations are correct, there are no issues on the conation perspective. The agent is just looking for the configurations also, and it says the configurations are intact.
There are no issues on that part, Okay? But the monitoring side of the system did detect that the pier went down, even though the configuration is compliant and in, in, in place, Yeah, the totality of the logs and the instrumentation that were coming in allowed it to go to reach this diagnosis. Is that the Y Yeah, so it definitely came in on the logs, but there's a monitoring system somewhere here too, right?
That's that's monitoring the status of the, The, yeah, the monitoring system cannot, uh, detect it because, uh, right now our monitoring system is detecting the status of the port and the port status still active. Oh, so the, the monitoring system is tagged to the port status, not the protocols that are running on. Yes.
So, so that's why we have the different agents, which we'll just try to find out. So our root call, our, um, demo purposes that even if you have the different systems which are monitoring different things, even if one of the agent can pick the problem, you will get the, your root cause. So in our case, the log agent has picked the problem of the BFT down, and we got the issues here.
So if you see here, Maybe I could ask a different way. Does someone have to notice that there's a problem? Or do, does that agent get fired off that goes out and looks and see, oh, there's a protocol issue.
So, you know, is it, is it take a human to do something to know that there's an issue? Um, okay, as of now, we are just, uh, waiting for the input from the human to give the source and destination vm because without that, in the, in the network, you can assume that you have thousands of applications which are running. And if I will just, uh, put this feature, which will just go and auto-create auto correct the problem, it'll create a more problem than the solutions.
That's the detection. That's all I'm asking about. Yeah.
Not, not applying any rem remedy. Okay. So, uh, in this current release, we are looking for the input from the human in the next release.
We, which we are bringing, it'll be the auto detection based on the different agents or the different monitoring system, which we will integrate with our system. And then we will take the action based on that. That That's where you, in this, in this case, the quality of your failure actually, uh, kind of hamstrung you a little bit.
Uh, because you know, I mean all literally every, every product out there i i is if the interface is still in that original po you know, status, it, it, the, all of the other, uh, overlay, you know, the, the rest of the network stack right? Never changes. And so you don't get that alert, right?
Right. And so in, in this case, your failure actually for that question hamstrung you a little bit, but we, but yes, The entry point to this is there's probably some dashboarding in place. There's probably some event triggers that are happening and yes, a human would have to come in.
Now we're kind of teasing out that autonomy level zero through five, right? At some point we're in the three range here where we have the human in the loop, there's human based interaction with it at some point in the fullness of time, will it be completely self-driving? I don't know.
Perhaps I'll be retired by then. But the reality is that there, you know, Google has been pretty clear about the fact that they do have completely au complete autonomy there. The difference is of course, they have a tremendous amount of parallelism in their network, and so they can make failovers very safely understand what They're Inconsistency within their own network, right?
You're, What's that Inconsistency that they manage control as opposed to Absolutely. When, when you're building it for an audience of one, you can have an absolute bespoke special snowflake and it's gonna be just fine. And here we're trying to get broad coverage across the number of, yeah.
So A quick question on the failure, 'cause I was trying to follow along what the, the value this was bringing. As I understand it, what you were trying to show is that you might have a scenario where forwarding on the port is down, the port's up, maybe you have received like mm-hmm. But you can't transmit traffic.
So you're trying, okay, why? And then it notice that PFDs down and subsequently your IGP or BGP is down, right? And it's making that analysis that, hey, even though the port's up, these things are happening and it's trying to point you in that direction and that, that's, I guess that's what I was trying to understand was a little bit murky at first that you, you brought back around, is that normally you would just say, oh, the port's down, that's the problem.
But this is in that, in this case, the port's not down, but things are still dying. That's right. It's A over hidden failure.
And then you compound that with Yes, exactly what you said, which is I have my, yeah, my physical graph, then I have my underlay graph, then my overlay graph, then I have some ECMP on top of that. And so all of a sudden you're sort of stacking these things together and it becomes a, a logic puzzle might be very fun for some of you. You're, when the network is down and everybody's sweating, you know, you're trying to, you're Collating failure domains Exactly.
Into an analysis. That's right. Yeah.
That, That, that would normally take a network engineer to look at and say, okay, there's this, there's this, there's this, that, that's, and you're wrapping it up into the agent of AI so that you can better understand the root cause. That's right. By letting it analyze it.
That's right. And when we talk about customized models, the question we had a little bit earlier, which is why don't we create a model, sort of a frontiers kind of model that knows networking, that is work that our Bell Labs group is doing. It's called the Noia language model.
It's in multiple languages, Arabic, English, others. Um, but it's not just knowing about networking terminology and terms and best practices. It's also then how do we build a model which understands specifically the graph, the different graphs that we have and be able to constrain those in order to do troubleshooting, right?
So this, this gets us to certainly is very impressive. First time I saw it, you know, you take a look at it and you go, well, this is pretty fantastic. I think we only can improve from there the understanding of it as we have better model technology.
If that's, I dunno if that's a term I'll just, So I'm curious if I take this to, you know, in the model that you use, and let's say I've got, yeah, I've got dark fiber and I've got this running on the other side, is it also able to understand that I don't have received light on the other side and make that determination As long As we have reachability, you know, if, if, I mean the platform can see all parts of that, Right? Assuming those pieces are in it. Yeah, I gotcha.
I know you can't just see into something that you don't control, Right? It's not a we took out our management routes. No, but in that case, if you're, you know, if it's dark fiber and you own both sides of it and you have this on both sides, then the analysis might be, instead of saying it might be this, the analysis might be you have Steve Light on this side, but not on this side.
That's Right. Right. And, and they should know, because this is the way it's set up.
We should know that, that it Should. So you would get a slightly different analysis if you saw both sides. That's right.
Okay. And um, and you know, speaking of which, we get to this confidence interval here. So actually the models go through, if they get score under 95% confidence, they rerun again, although in the four or 500 examples you've given.
Oh, we Have never seen That. Yeah, They're pretty smart. But I I I, I love the evidence-based piece here.
You know, when I, um, when I would commission circuits, you always have this idea of a birth certificate. You're running sort of performance testing. You're laying that down in the long-term documentation.
So you know your baseline here, uh, as you're doing RCA with an incident, you can tack this onto an incident. This is a PDF, it could feed into another level system. If you just want to take the data that's here and feed it into one of your bespoke, uh, operations platforms that you have, um, it's all right in place.
Yeah. So Just go add. Yeah, without a doubt.
Uh, Yeah, just to add one more thing. So, um, the reports which are being generated from the model, so we have to always put the evidence without evidence, it is of no significance. So that's why we are just keep on adding the evidence here and then we will just put the confidence number for that purpose.
Yeah. So Performance before we run outta time, two, two cool things I wanted to tell you. Um, so we have performance as well.
So if you're old like me and you set up a show command to iterate every one second on two sides and two terminal interfaces, so we can actually pick, um, elements here and get the same type of information. We can even live stream it, pick two and compare them. So have a live stream of the instrumentation that's coming off Interface.
So I, your issue, your question was related to the error packets and all that. So we are just here tracking the error packets from that particular vm. So whenever there is any error packets, it'll send the notification or it'll just start doing it here on the interfaith level that they package, which are received on particular interfaith or that vm, This is IA SaaS.
Is IDA on-prem, is this feature available? Not today, but stay tuned. They do have, again, they have, they have different capabilities, some core DNA together, different capabilities, but obviously you would think that we'd want to have great coverage of these features across all the different areas.
Yeah, this is a great feature. There's compar you know, obviously we're talking about ai, it's, we're all laughing because any, any little button that has sparkles on it, we're all clicking now, right? Because when we wanna see what happens with ai, the reality is obviously there's some infrastructure behind this to make AI work.
That would have to be planned on an, on an on-premise I installation. But that's, that's not the only reason. But yes, that's why they rhyme that sisters, but they're not quite stream yet.
Um, the last thing I wanted to show you and then we will, um, make sure to make good use of our time is there was a question about going backwards. We have time machine here. So because this is all powered through a data lake, right?
It's actually accessing and doing this AI root cause analysis on the data that comes off the lake. We can actually go back in time a certain amount of, uh, of time. So you've got time machine.
Now if we go back and say, what was it like 30 minutes ago? We run that connectivity test. This is now happening across the data set, not in the live network, right?
So where this might be useful is if you see those type of impairments that you, you know, maybe something happened which triggered something else, and then over time you have sort of a cascading set again, they become, they become someone hidden from you. Um, I know certainly plenty of times in my world you get to something and say, well this is what's wrong. Oh, well it actually happens from this thing that happened over here, time of day, whatever happens to be.
So you go back 30 minutes, you actually run these tests including the ai root cause analysis on the data as it was in situ 30 minutes ago, right? Which is, which is interesting if you have some sort of, again, impairment that that advances over time. I thought this was pretty cool, but it could, because it does allow you then to go back to time and take a look at the connectivity as well, so you can see what it was like in steady state versus what's happening now.
And do at the very worst, the steering compare. But of course you're gonna be using your AI and even So in, in this, can I see where the primary path is 30 minutes ago versus now? 'cause I, I found that was a lot of our failures is that some failover had happened and you get the report that 30 minutes ago I had this problem, but it's not there now.
And it's like, oh, it's because we failed from path A to path B or C. Right? Right.
OO Okay, so o on the time machine, we can just see the estates here, uh, 30 minutes ago. But, uh, what you are saying that is our, you our future plan from that data path perspective, I can give you the path, what path it has followed 30 minutes before. Yeah, that's the, because Yeah, so this is already in the, uh, future, at least for us.
So real quick question. I know you guys are are short on time, but I'm curious, you've got, you get pretty heavily into the overlay and the underlay there and the logical aspect of the network as well. You know, can you do comparative testing?
I say, okay, if I ping test the underlay it's this, but once I'm inside this overlay, whatever that is, that you can actually look at that and say there's a delta there between those. Uh, okay. So if you see here, the source to the leaf one that is your overlay and from leaf one to the leaf three, that is your underlay.
Okay? So we take care of, uh, both the parts and we just give you the information that if you just do the ping, at least from the latency perspective, what is the difference between the overlay and the underlip? Okay.
Or, or like if I change an M to u, let's say I change an MTU inside of a service overlay, maybe not in the underlay, you know, is it able to look at that and detect something like that? Yes. So, so actually we were preparing we're not gonna do it 'cause we don't have time.
Okay. Sorry, I'm, I'm Pushing you into all the good demo Because on, on Monday we, I came into town, I live in Atlanta, I came into town, we were prepping for this and he actually gave me an MTU based MTU mismatch demo because Mt U is never the problem. I laughed because I said it's 32 years later.
It's such a chronic pain. Like two things in my life are consistent things, not the ultimate diagnostic tool. And watch your mtus, like those two things will solve 50% of mm-hmm.
Know your MTU math. So we, So I run up the mtus and started up, we gave an example where there was an mt U mismatch and the AI was able to figure that out. Yeah.
'cause you'll spend all day chasing that. If I Had that when I was deploying broadband in, uh, you know, 1998, I would really great shape with DSL and mt U sizes and GT PE and all that. That might not be great.
Exactly. I'm only 26. Yeah.
No, but, but it's a, it's a great example. And we did actually play with that because again, chronic things that happen. Yes.
Well the reason I ask is a lot of things get into the underlay, the control plane, the physical links, not as many get into analyzing the state of the overlay and things that of that nature. That's a lot harder to analyze. Yep.
We have a little bit of time. Okay. Oh good.
Do you have one more thing you wanna show? Uh, Mt You want to say If you can do it quickly? I don't want if, if you know how to do it mt you quickly, that's good.
Yeah, just gimme to prove out what I just said. Just give two minutes So we don't have to trust you. It's just funny you mention that.
'cause that's the first thing I thought of too, is It's like software, it works really great in test, right? Oh, and, and uh, IP don't fragment bits. That's the other one that always, you can tell I've been in the Broadway world a lot 'cause I have those, those, uh, those pain points, those scars.
Just gimme a moment. So, uh, this is the, um, overlay MT u which we are going to change here. So right now it is on the, uh, normal default value of 1500.
We are going to change some value, which is, uh, pretty large. Then the interface MT. U and then we will see how it'll just, uh, impact and what is the root cause you are going to get Not able to commit it, Not able see the mo Oh, zoom is on top of your button.
That's what see it. Okay, So the real question is when are you coming out with the stop? Don't do that number feature.
Don't Use that number. Not sure. Zoom.
Can we move that? Okay, There you go. Okay, so we have changed the MTU and uh, what it'll do, it'll just try to find out what are the, uh, overlay elements which are dependent on your MTU or the vlan.
And then if suppose that VLAN is down, what other elements will be down? Maybe it is the bridge interface, B-F-T-B-G-P on the overlay perspective. So we will go back to our apology again and we will take the VM two to PMM three.
Excuse me. So we are trying to find out if it is able to detect the mt. So if you see there the, uh, defense here, now you bubble is red here and your service status is failure.
That means it is saying there's no path between your VM two to VM three because we are doing the MTU at the source leaf itself. So the traffic will hit the leaf and it'll just get dropped there. So now we are just trying to find out what is the module which has affected here.
So we are seeing it is a sub interface vlan, which got affected due to the Mt U change. So now let's end the D. So this time it might be possible that based on the problematic state, it'll launch the different agents.
It's not the same agent, which is it has launched for the previous time and it'll launch, it'll get some more data go through the same process again, fine tuning of that. Go for the evening and then get the e for us again, it'll start, uh, going for the chain of thoughts start, what is okay if VLAN is down, what BFD is done? If BFD is done, BGP is down.
So it'll try to correlate those data and then give us the evening for that. And Just, I just wanna clarify that Denise. Um, so it, it's calling different agents, it's loading different agents depending on like as it sees different things going on.
Yes, an agent, it is an agent, There's multiple models at play. So one is what we call a planning model. So it decides what it wants to go find out from these other tools that are, that it has to call them.
We have a tool calling model that specializes in gathering that data. Remember earlier I mentioned that, that compressing the number of tokens, doing that filtering upfront on what's important for this particular root cause analysis is important for two reasons. Number one, it saves costs ultimately, but even more so it creates higher quality outputs when we pass it off to, um, for analysis, right?
So being able to compress that data, or not compress it, but filter that data and discern what's important coming outta that, then we pass it off to the reasoning step, which allows it to formulate its opinions. Then a large language model used to produce the report. That's really just a human language sort of, uh, motion there.
But the, the, the models themselves are responsible for one, a planning model two, one that's responsible for calling the tool like you. Yeah. You're curating the context For your Exactly.
Curation step. Right? And another quick question, is this part of IDA also multi-vendor like the, so the other, in the other demo that we saw The IDA SaaS follows, um, in with a, with a short delay after an EA release.
So the capabilities that IDA are picked up. And so I asked that same question of the gentleman who's in charge of edda and he said shouldn't be any problem, right? So if, if you pull that multi-vendor piece in, because we're using Edda as kind of the core and the intelligence of the system.
Okay, you're gonna pick up that. So in theory, if I have a mixed network of Cisco and Arista and Nokia and the data center and I wanted to do that mt u analysis, that's possible. Yes.
Because that's a neat trick. It's not a lot of people that are able to do that. Only thing is we have to add some additional agents.
Yeah. That's the only thing we have to, so if you see here, your conclusion is there the L two M two U too large and due to that your VN was down between this, Like you said, where was that? 25 years ago?
Yep. At at three in the morning. Yeah, exactly.
Well by the way, it's never gonna be 9,400 or 9,500 something that your eyes will recognize at two in the morning, 3, 8, 4, 5, some number that looks like a port number, right? It's really gonna be tricky. So 1402 instead of 1400, what are the, what are the scale limitations of this?
If I've got a, you know, a really large data center fabric was, are there upper limits to scale? Where, how far can this go? Oh, Okay.
Go ahead. Very good question here. So if you see here, our network does not depend on how big your final network is.
I am only concerned about the connectivity from the VM perspective. Okay. So only the node that are in half the service or the, or that I'm analyzing.
This is the Sub graph of okay, of this is the sub graph that's eligible for traffic passing for these between these two M. Got it. Okay.
Which is why we kind of highlight here where this is, this is ultimately a test between an endpoint and an endpoint. We're all network engineers looking at this as a network. But the reality is we have a plugin here that can go to OpenShift or VMs, you know, vSphere or whatever it is, pull the metadata here to correlate the, the root cause outta there.
So we're doing app, this is application layer from end to end, not just the network itself. So we Are putting ourself at the source vm and then we are looking for the network from this vm, how the network looks like. Gotcha.
So it doesn't depend if you have a hundred switches or 200 doesn't matter for us. Curious about the data security, data protection, the models that you're using, are those Nokia's own that you're running yourselves? Are you using any frontier or third party?
In other words, I don't necessarily want my data about my network being passed off to a You said model. I went to YANG model. Oh, sorry.
AI model. AI model. The a, the small AI models we have are our own.
Mm-hmm. Right. They, we've done some reinforcement learning on that.
We've optimized those to do the tool calling the agent tool calling that you see, and then the, um, the large language model, we're just using a public frontier model to generate that report and, and communicate in the human language element. Okay. But the, the heavy lifting is being done by our own models inside of here.
Yeah. And So all that data stays internal. It's not shared with any third party during any of that process.
Yeah, completely our domain. That's right. Not that that would be important, would it?
Oh, not at all. No. It's a, it's a strong concern, which is why I mentioned before about Yeah.
In the fullness of time what happens here, you know, we run models which are specific potentially to large customers. Certainly you can see for a lot of our telco CSP type of customers, they have very specific types of model that they want to run based on equipment that was put in a while ago. You can see where the, this, this takes us as an industry is now there's gonna be sort of a model tech opening of, of, I'm Gonna say how many 5G mobility data centers are.
Uh, yeah. This is this running in Cloud and enterprise. So that's what I, that's what I'm focused on.
Okay. But yeah, certainly, certainly there's a lot of opportunity there.