Demonstration of Day 2 AI network operations, monitoring and anomaly detection with Aviz
Aviz Networks’ AI Infrastructure Field Day demonstration focused on Day 2 operations, monitoring, and anomaly detection for AI workloads. The core challenge addressed is the specialized networking requirements of AI, including multiple networks, differentiated QoS, and the need to manage compute as part of the end-to-end network topology. Aviz presented solutions for orchestrating AI fabrics based on Sonic and NVIDIA’s Spectrum-X reference architecture, showcasing a customer workflow that includes network design, Day 0 infrastructure deployment, Day 1 tenant onboarding and traffic isolation, and Day 2 operations like adding Pods, handling alerts, and troubleshooting.
The presentation demonstrated Aviz’s orchestration capabilities for Sonic-based and NVIDIA RA-based AI fabrics. For Sonic, the presenter showed how to orchestrate the fabric using YAML-based intent, validating configurations, and performing operational checks. The demonstration emphasized the ease of use of industry-standard CLI, built-in validation, and the ability to compare configurations to identify any drift. With the NVIDIA Spectrum-X platform, the presentation highlighted agentless orchestration, the use of NVIDIA AIR for simulating deployments, and config comparison.
Finally, the presentation detailed Aviz’s monitoring and anomaly detection features. The tool provides comprehensive monitoring with a bottom-up approach for networks, servers, and GPUs. The demo showed how to view various telemetry data, including traffic, queue drops, and GPU health metrics. The presentation also covered Aviz’s built-in anomaly detection system, which allows users to create custom rules and receive notifications through tools like Slack and Zendesk. The system includes curated rules, role-based access control, and configuration comparison capabilities to streamline operations and reduce potential errors.
Presented by Ravi Kumar, Solutions Architect, Aviz Networks. Recorded live in Santa Clara, California, on April 25, 2025, as part of AI Infrastructure Field Day. Watch the entire presentation at https://techfieldday.com/appearance/aviz-networks-presents-at-ai-infrastructure-field-day-2/ or https://techfieldday.com/event/aiifd2/ for more information.
Transcript
So I have an agenda for the demo today. There are a few things that I wanna show. So I'll go over how we orchestrate sonic based AI fabric using Sonic.
Then I will also show how we use Nvidia RA reference architecture to orchestrate the fabric. I'll also go over some data operations. So if you wanna change any configuration, you wanna do backups and restores, how you do that.
And I'll also show what inventory collect, how we do anomaly detection with the inbuilt, um, uh, rule engine that we have, and how you can integrate all the alerting and, uh, notification with customer tools, Zendesk, um, ServiceNow or Slack. And obviously between I'll, uh, I'll, um, have some questions also. So I'll start with sonic based AI fabric first.
So I have a small video. I'll be jumping in between videos and a live demo. So second.
Okay. So if you look at the topology here, I have two spines, three leaves. Those are Broadcom based switches running, uh, Sonic right now for the orchestration itself on Sonic devices, we have a small agent called F-M-C-L-I that enables us on the controller to actually talk to the switches directly in an industry standard CLI.
I'll show how the CRI looks like. So you don't have to learn any new CLI to use Sonic in your network right now. So you go to configurations.
And how you actually orchestrate is we have a simple intent-based yaml. Uh, one second looks something like this wherein you define your network. So your inventory, how many spines, how many super spines, leaves and tos you have, how they're actually connected.
So we have a plan to integrate LLDP here, but LLDP can sometimes be, uh, you want a source of truth right now. So here you just provide how the switches are actually physically connected. So you provide your inventory, the physical connections that they have right now, and what parameters you want to configure.
Do you wanna configure bgp? Do you wanna configure BGPN number? How are the physical interfaces look, looks like?
What as numbers you want to use for, and various other parameters. So if you look at the YAML here, we are talking about the whole fabric and not just per switch. So you define the intent for the whole fabric, whether it's a small deployment, large deployment, and you basically say what as numbers you want to use for the super spine, for the leaves stores, what IP pools you want to use, and various other parameters.
That's a quick question. Yes. All right.
So is this, um, do you fill this information in on, in the YAML script? Yes. Or do you have a GUI that you fill it in on?
We have a gui. Uh, I, I'll come to that also. So we, we, uh, we expect our customers for a large deployment to edit the YAML file.
And for that we have a open source re report where we have 70 plus test tested templates. And if you go here, if you have a normal IP cross deployments, we have a template for you. If you have multi with vxlan, we have template for you.
So what we actually expect is you go here, you, you take this yaml, you edit it for your specific inventory, for your specific connections. You fill out the IP addresses, your username and passwords. We also have all the QS configurations, all the various things that you need for the actual Sonic EEA fabric, all your schedulers, your PC if you want to enable EC, and the, the everything here.
So we expect you to download this, edit it, and basically orchestrate it. Yeah. So you get this, you put it into the system, and that does the orchestration for you.
So when you, uh, when you start the orchestration with all the QS and all the parameters that you have, we do two kinds of validation. One is configure validation, wherein when we are actually configuring the switch, we look for any errors that might be thrown out that gets the, uh, the status into a field state, and we stop for that particular node. The orchestration is stopped if it's a failure.
We also do operational checks. So we are also checking for control plane. We are also checking for data parts.
So if we have, let's say, three leaves right now, and so what, what it'll do at once, it'll check all the data parts from leaf one to leaf to leaf one to leaf three, all the data parts that I have, it'll do, it will, it'll run show commands to see if everything is as it should be. It'll also do ping test from all the leaves to all the other leaves so that we know the thing that we want to configure. It's actually reachable from here.
And one thing I wanna show is, while you are doing all this orchestration, you can also look at the, uh, running logs. So this is one of my live instances that I have. I'll come to the topology later.
What I wanna show now is you can look at the running logs. So you can come here while this is being orchestrated. So everything that is being done on the switch, you can come, come and view live.
So if you look at the CLA here, right, we go into our fm, fm, CLA, basically, and the agent. And then the C is industry standard. So you're going into configuration terminal, you configure interface, configure interface if you want to configure BGP.
So the UI is industry standard. You don't have to learn anything new. So we are doing show checks.
We are doing checks for the BGP itself. If everything is up, we are checking the neighbors, what is the state of the neighbors? And we are also doing at the end ping test for each of the neighbors or all my, all my hosts, all, all my, all the outsets I have, are they actually reachable?
Quick question. Yes. Uh, Brian Martin, signal 65.
Uh, while you're doing that, uh, configuration check, what happens if you find mis wired gables? Yeah. So, uh, the, okay, so configuration will pass because configuration was not an issue.
You were able to configure the interface. You were able to configure the BGP, but the BGP will not come up, right? Because the wires are messed up, right?
Mm-hmm. So, on, on the status here, one second. On the status here, it'll actually show, when it fails, it'll actually show the BGP between leaf one and leaf two, or leaf one.
And spine one did not come up because the interfaces maybe are not up, or the interfaces are up, but the BGP did not come up. Mm-hmm. Or at earlier stage, BGP is up, interface is up, but we are not able to bring it for some reason.
So it'll clearly show that message here as a popup, and you can then go down and look at the whole logs. And for that matter, we have a connect button here. So let's say you see this issue, something failed, you want to connect directly from here.
So what you can do is you can click on this connect button, it'll take to you, to your ui, you can connect to the switch directly via SSH, or if you have a console server, you can put the credits for the console server, and that will take you to the actual, uh, SS session for that switch. So it, let me check if I have a failed session right now. This is one of our, our second, uh, that I have.
Okay, so something failure. So it, so we are, we are orchestrating topology, and it's saying, uh, when I tried to ping the host from leaf on or LEAF two, something failed. 005 from source ten, ten, ten two with packets.
And this way I can go and look at the logs here and it'll basically show me everything passed. But the last step of the configurations is, is showing BGP. Okay?
So BGP here is stuck inactive. That's why that that is not coming up. Maybe now the wire is messed up, or, but it, it clearly shows you that something that the intent was for the configuration that did not happen.
So we are marking it as failed. And for that particular note, the orchestration stops, but it continues for your other host. So you can go back, look at that specific note to figure out why that failed.
Okay? So I'm confused. Was that an expected failure or an unexpected failure on your behalf?
This is unexpected. Oh, because we are expecting source of truth and we are expecting that all the information that to provide in the yaml, the interfaces that are directly ed, we are expecting that it, it has to be source of truth. That is why we are not going with LLDP right now.
We have a plan to basically have LLDP, but we are still testing internally. We are relying on you to provide the actual source of truth for us right now. Okay.
So You didn't construct that. I'm just trying to understand. You didn't construct that in the demo or you expected to see that failure?
I'm confused. Oh, no, no. Oh, so the failure, this is, I'm just trying to show, this is an, another demo that I have.
Another instance. Okay. This is just, I'm trying to show that if it fails, please proceed.
We will show you. Um, excuse me, another question. Um, we are seen in the yaml, uh, the, the credentials.
Yes. Is there a way to put the credentials outside the yaml? So are you implementing something like, Yes, we have request from, uh, customers who use LLAP or, or sorry, uh, LA or something else, right?
We do right now per customer, because every, uh, we see every deployment is a bit different. So we do support it, but right now on the public GI that we have, it's, it's, it's there in the aml, but we do support it. But it's right now, uh, per customer basis for right now for this current.
Okay, gotcha. Yeah. Okay.
So, uh, any more, any more questions for, uh, sonic based AI fabric? So I'll move on to NVIDIA based orchestration, how we do that. 0, the reference architecture, uh, we design your AML intent.
So we take the input as a number of GPUs you wanna use in the A fabric, the first IP octa you want to use. And based on this input, we actually construct the YAML for you. What we can also do is we construct the YAML for your physical real topology.
What we can also do is we can create Nvidia air dot files. So rather than actually going and deploying your physical topology, you can create kind of a digital twin with Nvidia Air beforehand to test out if everything is working okay, working, working fine, and then go ahead with the actual real deployment. What we can also do is the server itself can behave as A-Z-T-P-D at CCP or NTP server, but obviously you can use your own.
We also had an idea here for the network design itself, rather than taking input as the number of GPUs, maybe we can take the scalable units as input how many scalable units you want here, but we are still deciding on the UI part what input would be best for the users to, uh, to make it, uh, easier. So the UI is not ready here yet. The, this part will be, when we have a ga this part will be skipped.
But what, what eventually is happening is we take the input as a number of GPUs, the IP subnet that you wanna use, and we are creating this YAML internally. This will all be internally in the, in the g release. And basically, I'm just copy pasting it here because we don't have the UI yet.
So I have my YAML intent here. So this is all running on Spectrum X platform, the Spectrum X platform. We are doing agentless, so you don't have to install any agents there.
And while it does that, it auto discovers it based on the host team that you have. So it goes there, passes the YAML for you checks, uh, do an initial validation. If all the notes that you're mentioning are all the parameters that you're vening here are, are they actually valid?
Does that and starts all the orchestration in pallet. So this not only, uh, orchestrates your switches, but also does, uh, the, the, uh, compute notes because we wanna have a route back to reach out to the, to the network. So we also orchestrate the actual compute notes itself.
And this topology right now, this, I'm running in Nvidia Air, so orchestration has started. While this goes on, this will do all its validation checks, configuration checks. And while this does that, so this is the topology that I have in Nvidia Air right now.
And if you look at it, I have an out of band management switch that I have once. So the controller itself is running in Nvidia Air for this specific demo as an OV sitting here and reaching it, reaching all the switches over out of band. So we have four spines here and eight leaves, and with four actual compute notes down below.
Ravi For, uh, my fellow network folks who, um, aren't familiar with NVIDIA stuff, what is Nvidia air? So it's a, uh, virtual environment. So you might be aware of GNS three, right?
That runs locally on your, so it's A GNS three. Yes. Sorry.
Yeah. So it's an environment online cloud environment where you can go and simulate your topology. So if you, if you, if you have an intent to simulate, let's say one scale unit, those pines four leaves, and you want to test it out before actually going and buying the hardware, going actually configuring your physical switches, you can do that here.
And the configuration good part is all the configuration that is being done here, it's identical one to one with the physical actual deployment. So that is why we have it as part of our deployment also. So you make the YA intent, make the dot file, configure this topology here, and that basically gives you your real, real world kind of a scenario where the configuration is also same, the behavior is same, everything is same.
Cool. Yeah. Thank you.
Yeah, Yeah. Quick, and I know we, we obviously so much into this every day that Nvidia has basically simulation cloud that Nvidia runs a cloud simulation environment that they have where you can in instantiate Nvidia devices as as a digital twin, right? And so you really can simulate network fabrics, including the servers and the next, right?
And so this is very, very helpful. When you design a network, you basically put all the configuration together, you push it and you instantiate it a real as like a real topology, but it's virtual, right? Not physical devices in Nvidia Air.
And we use once our, our tool to basically validate that all the links up because the cloud environment is like a real digital twin of what a real network would be. And you can do the same thing. I can check all my interfaces, I can check what he just described for Sonic, right?
In the video world, this is gonna be running cumulus, but os right? Mm-hmm. Same thing.
But you can basically validate all the switch configuration that everything is reachable, all the things you can do in Nvidia ai. And when you're happy with that, right? And you're gonna see it on the screen, that's why the screen looks the same.
That's an our tool. We are using Nvidia s the backend to basically probe insensate, probe happy, and now I can go and deploy. So that's maybe a little bit more going slower.
Thank You. Yeah, no worries. Go Ahead.
Thank you, Ben. Yes. So it does all the configurations.
Basically it's a large fabric, so we have everything done. Basically it waits for everything to be completed. And one more thing I'd like to show now is if I go back to my live setup, I have this, right?
So for day two operations, we also have, we can also do, uh, uh, we can look at the intent, what, what the YAML that you programmed, what the intent of the configurations was. So for this specific that I have here, my intent was I wanted to create some ance on some lease switches with this IP submit supporting 64 hosts. And, and this is IP that I wanted on it.
And the controller designed this configuration that is actually pushed for this intent to happen, right? And if you look at the CLI, right? It's, it's industry standard.
It's easy, easy to understand, easy to configure. What we can also do is we can also do config compass. So if you, if you are applying configuration and you want to look at what I applied and what is actually running on the switch right now, we can do that.
We can, we, because we also do backups and, and, and resource. We can also look at the last happy baseline backup that I have. So we can do that.
So for the intent, I wanted to apply this, this is what the controller seems, things that should be applied, but maybe somebody went ahead and changed something in the background. So we see the route ID is now changed. We see VLANs are now changed.
So you can clearly see the intent was this, this is what we generated, but something happened in the background making, making the difference, right? So what we can do is, rather than looking at the running configuration, when the orchestration is successful, we also do a happy basin. So we take a backup automatically at that time.
So what we can do is we can look at that backup at that time, what was the intent? So we see the intent and the backup at that particular time. When the con, when the actual, um, orchestration was completed, this looks good, but something happened in the background.
So this gives you a clear idea of if some, somebody is changing any configuration in your, in your fabric. What we can also do is if you are looking at running configuration, we can, we can, we can basically compare configurations on what happened. We can compare configuration between different switches, what is running on leaf A, what is running on leaf B, we can do comparison between them also.
So all the running configuration is actually pulled in real time. So agent contacts, the, uh, the con the agent, sorry, the controller contract contacts, the agent gets the real time whatever is running on the switch. And we can basically compare here Maybe to step in, you probably, I don't know if you ever have done this, you probably wonder know why this is actually matter.
Let me step out of the picture here for the gentleman PRA as if you do like modeling type template driven, uh, deployments, the configuration tend to be the same except the IP number in, right? Right. And so what does allows you to do is, is there any drift happening over time, right?
If somebody deployed is everything is happening and doing normal operations, maybe somebody configure port up, port down, and you can easily compare leave to leave when you know they should be the same and they're not. Right? So that's one right?
Between same type of devices. That's a comparing two devices. The other one is, uh, you had a golden configuration when you push it and something drifted on that same device over time, not they had a backup.
That's, I know that was a good configuration, the current one of this. Tell me what's the difference, right? And then obviously you can use as a normal networking to roll back if you need to, right?
There are a lot, a lot of different scenarios, but the point really what you look here, it's very, very easy to do configuration comparison on a per device basis. Uh, and then just adjust, right? And this is going back where when we started this is it was a pure monitoring tool.
It becomes really, how do I operationalize anything that has to do with configuration and, and templating, uh, and just, just to give some contact, uh, context. Um, so yeah, keep going. I have, I have more, but Okay.
I'm pretty sure you have more questions. Yeah. What we can also do is, uh, do data operation.
So let's say you have, you have, you have orchestrated your fabric. Now you wanna do a VLAN change, you wanna maybe add a network, add a vlan, so we can do that. So for any switch, you go to NetOps, it gets your current running configuration on the switch, and you can basically pull it here and make a change.
So let's say I want to change the router id, right? I want to add another VLAN now. So you can make the change directly here and click on apply.
And basically what the controller does is it realizes that only two, we are adding a and changing the ip. So it only does these two changes, uh, on the actual device. You don't have to reorchestrate your whole, uh, uh, fabric here.
So we can do that also. And last, we can do obviously backups and resource. So you can give it a tag, some tag, start the backup, and when you have the backup, we can store multiple backups and you can do restores for that particular switches or, or your fabric directly from here.
Any any questions At the risk of asking a heretical question? Um, this tool looks phenomenal for network engineers. Yep.
Uh, in simplifying, uh, their day-to-day task. Um, what about the accidental network engineer? And you could tell me in a deployment this size that doesn't exist.
Um, but what about the accidental engineer that knows enough to get himself in trouble mm-hmm. Herself to get in trouble, but not enough to understand the edits you just made and putting those VLANs in the right place? Yeah, you would Go ahead.
So, uh, we control it, uh, and in a few ways. One is we have different roles you can assign for the UI itself, wherein you have admin rights for a user. You have where you can create telemetry, clear, clear, create notification, look at all the alerts, and point them back to your zenex or server desk so that some, uh, agent can look at it, or you have view only access here.
Mm-hmm. Uh, so basically how I, I think that that is how, but when you actually do a VLAN configuration mm-hmm. And let's say you enter the wrong syntax or something, it's, it'll tell you that the syntax that you're entering, while it is orchestrating, it'll tell you that syntax that you're entering, it's wrong.
It's not going to actually go and configure it. So the UI will show that there is a difference between the running config and the config that you're trying to do. But if the configuration is wrong, it's not going to push it, it's going to say the syntax is wrong, or the the thing you did that you're trying to do, it's not correct.
Okay. And this is where copilot might come in in the future. And yes.
Let me say what I want to happen and you'll translate it to the right words. Yeah. Yes.
Yeah. Let me, let me expand on this. So there are two things.
You're absolutely right. This tool is very powerful from an, like if you had to operation support for the network mm-hmm. You literally have control over everything, right?
You can literally change the device configuration line by line if you want. To your point, that's not what you want everyone to do. So we do have the multi role capability where you say, Hey, if you user have this role, you cannot change.
You can only view, or there's certain things you can only do. So that's one, one way to, you know, uh, adjust this. The second one is, what we showed was templating, right?
Where you saw, when you did the NVIDIA scenario, all you have to enter is literally three things, right? Tell me how many GPUs you want, what's the starting subnet for the next, and that's it, right? You don't have to know anything, right?
You don't have to know how to, because the number of GPUs we generate, how many switches do you need? It will tell you, well automatically what the topology is. Mm-hmm.
It will automatically tell you the IP numbering on every port on these switches. Uh, and the auto, the, the conversations automatically out of all of this input automatically con uh, con, uh, generated. And that just applied simulated.
And you don't have to, you can see it, but you don't have to know, right? You can if you want to, but you don't have to. So that's the, the other extreme, right?
This is what we see where a lot of vendors go with, uh, reference architecture valid designs, right? What Ravi described. And so we have this for the, uh, NVIDIA Spectrum X reference architecture, but we also have, and this was the comment to the YAMA file, which is a public good, uh, uh, repo that we have where we are actually having YAMA files for different device configuration that basically pre-populated with the right settings.
And then all you need to do is adjust, right? Because we created them. Let's say you want a four leaf, two spine apology.
If you say, Hey, I have more service. I need six leaves and two spines, you just need to copy the two leaf confirmation, always adjust the numberings and you're done. Right?
So that's more the template driven where you reduce the amount of the technology you need to have, right? But as an, as an operator, once it's deployed, you want to have full visibility, right? If you have to troubleshoot, you need to be able to go to every individual device.
And so we have both of these modes, right? Yeah. And then the third one you touched on is the copilot, which we don't demo, but clearly all the data is available to an API, right?
What we just show is config one, config two. You could literally type the question saying, Hey, I want to know the conflict differences between the backup that was taken on day two and day three, and please summarize this for me. And you basically got a nice little written summary, what, what the differences were.
So that's why I personally think a lot of this is gonna go just to make it simpler and faster instead of trying to figure out staring at the screen, like what you guys did. I saw the template for the NVIDIA configuration. Yeah.
You also have a template for Sonic environments As well. Yeah. Yeah.
All this, everything is, it's Sonic, It started off as Sonic, and we are adding, so if you think about Cumulus as a very close, and I need to be careful the way I word it, but it's same roots as as Sonic, right? So what we do all this templing, all we do for Sonic and you can do easily apply to, to cumulus environments. Mm-hmm.
Right? And it's, it's important, right? The reason why we have both is Nvidia supports on their spectrum switches, Cumulus, as well as Sonic.
Mm-hmm. So we will give customer the flexibility. If you wanna deploy Nvidia or Cumulus, we can do this.
If you wanna deploy on a GPU side, cumulus, and on the front end for whatever reason you want Sonic, we can do this, we can give you both with the same tool. And then obviously you can use any other switch that runs Sonic on the front end. You probably have you, if you catch, if you call what we had, this was like a, there was a real lap network work that is running there.
So we actually, I'm actually amazed, Ravi is like when, when he went from the video to the real network. This is a real running lab network. As you see, we have Cisco devices running Sonic, and we have Acton device running Sonic in there.
Mm-hmm. Uh, pretty much every major hardware right now. Good.
You can run Sonic on these days. And so we will have that. Can I use these?
So it's a define, I defined a configuration. Can I use it like a defined state where I compare what's actually happening out there to what I specified originally? Yeah.
Would that be a regular feature? Something that copilot could help me with Both. So you want to go back to com, compare ConX.
Mm-hmm. You basically use your original, let's say you created your original intent and we use a yam file right now, but you can pull it from wherever you're gonna get it, right? Yep.
And saying, Hey, this is my config zero day zero, right? That's intent config. And then you can pull the actual one and say, compare this space.
You see it here. And now either you're really a geek and you look and says, oh yeah, I'm very comfortable with this. Or saying, Hey, this is like, because these confis can be very long right now, this is where copilot can help.
Right? And you basically just look into copilot and saying, please compare these two configs and summarize for the differences. Both of these options are there.
Thank You. Yeah. Okay.
So last thing I wanna cover is, um, how you can monitor this, this project. So when it comes to monitoring, we have a bottoms up approach. So Sonic, we started once from with Sonic.
Right now we support 300 plus, uh, unique matrices. And if you look at it once was originally built as a support for, for the network engineers, right? And when you have an issue or an error, we want to be able to find all the matters that we have very quickly.
So this is how we started as a support pool. So right now we support end-to-end monitoring with networks such as servers next and GPUs. We support Rocket Elementary, whether it it's utilization, counter skews, drop, PFC, easy and everything.
We also do GPO metrics per GPO metrics for health performance utilization. And obviously we have alerting and integration with tools like Zendesk, uh, ServiceNow Slacks, right? So for, for the actual compute itself on the server we have, we can get GP inventory, GPU health, and basically health of the server itself, plus all the GPUs that are running on the server.
We can get all that. And for the switch methods we have for the AI specifically, specifically for the ai, we can get the Rocky configuration, all the maps, schedulers, debit configurations that are actually running on the switch right now. And we can get all the counters associated with it.
So I will show that on the live demo itself. So this is kind of the landing page that we have right now. And if you look at, we have different platforms running Sonic, and we have all these different basics in a single multi highway skew controller, you have the visibility here, we have different views that you can look at this topology at.
We also have a ML view, but what I'll focus is we can go to traffic to it, filter out all the qs, basically the ai, AI driven devices. We can say, only show me PFC enabled devices. And on that only show me all the, all the, the interfaces that actually have the QS configuration for the AI fabric to actually work.
So this shows me, these are all your interfaces. So I can go to that particular interfaces, look at basic things, in and out packets. They are utilization, the discards and everything.
I can go to advanced to look at frame level telemetry. I can look at unique queue packets for all my queues, the queue drops. I also have visibility for my bump traffic.
Quick question. Yes, sorry. Uh, is there a way to explore this data or are you integrating that with the third party tools or?
Uh, Not right now. We, uh, no, not right now. No.
But I think, so you can, the agent itself, the one telemetry agent that sits on the switch, yeah, it is GRPC based GM and GRPC based. So you, if you have a collector that can talk G GM, GRPC, you can probably pull directly, directly, rather than using our goi, you can use your GUI and talk to our agent. So that Works.
Yeah. So for the, uh, for the, obviously you have the rocky telemetry here for this specific interface, what are all my DCP mappings W profile schedules? If I have PFC enabled, if I have watchdog enabled, what are my packets?
So is it, is it normal traffic or do I have any rocky specific traffic here? So you can visualize that. Also, all your PFC counters, all your queue drop, if you have any condition notification packet for which queues, you can look at that.
And all this data is visible from, for a window of one to a maximum of two weeks for data older than two weeks, we have, uh, connectors for data lake. So we can push it to a cloud or uh, store it somewhere else if you want to do analytic analytics on top of it. But the UI itself, we do one hour to two weeks.
We also have health of your fabric. So you can look at CPU memory and all the basic telemetry, right? We can also do capacity.
So we can look at the asic, we can find out how many, so for the a c itself, how many acis it can run, what is the current utilization? Same thing for IPV four routes, I PV six routes. So you can look at how my ASIC is actually visualized.
How Much, when you say the asic, you don't mean the GPU? No, we do the GPU as well as the ASIC network. Asic, yes, We can do both.
So I'll come to the GP part, but this is specifically on the switches right now. So the asic, right? And obviously we have link level telemetry and we have protocols.
So we can look at bgp, the I-B-B-G-P kind of and go format. We have L-A-C-P-M-C lag for the fabric itself. We have qs.
So you can come here, look at what is actually active on what interfaces come here for details. And if you come and click on any of this, it'll show you the actual configuration that, that it pulled. So these are all my queues that I have right now.
All the mappings, my w already CN profiles, the schedulers that I have, and if I have watchdog enabled and what, what I'm doing with that. Mm-hmm. And the GPU stats are coming from a, a host agent or?
Yes. For, for the, uh, for, for a server running GPU, we have an agent, a small lightweight agent that you can put there. But for, for the ulus, the Spectrum X platform itself, right?
It's agentless. Mm-hmm. Um, I think we are almost out of time.
So I will quickly show we have GP level telemetry here. So the number of hosts, we have, the number of GPUs running, we show, uh, the all basic telemetry about your GP memory per, you can also look at per GPU and with all this telemetry that we have, the last thing I wanna show is we can create some rules. So we have an inbuilt anomaly detection system wherein you can come in and say, I want to create a rule for GP utilization.
We have per device, per interface, many things and per server, obviously, right? So I'll say, I want to create a notification when the GP utilization average for the last 10 minutes is greater than equal to, let's say 70%, show a warning threshold critical for 90, and send me a Slack message. Whenever this happens, send me a Slack message.
And what can, what I can also say is you can also raise a Zen Desk ticket directly from here. But what I'll say is if raise a Zen test ticket, only if the critical threshold is actually breached, and send me a weekly test. So what happened for the week?
How many notifications, what kind of notifications do I have? We did my, so summary of all the alerts that happened for the past week, you get it as a weekly on your, on your Slack. And because you don't want to get annoyed.
So we have a control here. So send me a Slack notification only once for every one hour, no matter if the alert keeps coming. So you don't wanna be annoyed.
So we have that and we also have control over, oh, stop after five times. So you send the notification five times, stop after that, I'll take a look and you have some payload. So here you can write custom debug steps that you want to see.
And we can basically come here and, uh, look at the alerts. Do you have a set of curated rules, like a smarter Kit? Yes, yes, yes, yes.
Okay. That comes preen enabled. And this is what the payload looks like.
So this is what you'll get in the Zendesk, or if you raise a, uh, uh, a Zendesk ticket, ticket or Slack notification. So you have what happened when it happened, where it happened, and how to actually dynamically reach there.