Day 2: Operating your AI data center with Juniper Networks
Juniper Networks presented its latest Apstra functionality for AI data center network operations at AI Infrastructure Field Day. It focused on providing operators with the context and tools to manage complex AI networks efficiently. Jeremy Wallace, a Data Center/IP Fabric Architect, emphasized the importance of context in understanding the network’s expected behavior to identify and resolve issues quickly. Juniper is leveraging existing Apstra capabilities, augmented with new features such as compute agents deployable on NVIDIA servers, and enhanced probes and dashboards, to monitor AI networks. This presentation aims to equip operators to maintain optimal performance and minimize downtime in critical infrastructure environments.
The presentation highlighted the evolution of network management for AI data centers, transitioning from traditional methods to a more proactive and data-driven approach. The core of Juniper’s solution involves leveraging telemetry, including data collected from GPU NICs and switches, to provide real-time insights into network performance. This enables operators to monitor key metrics, such as GPU network utilization and traffic patterns, and respond to potential issues swiftly. The Honeycomb view, traffic dashboards, and integration with congestion control mechanisms (ECN and PFC) demonstrate how to provide visibility into the network’s behavior. The goal is to provide context and the tools to diagnose and resolve problems faster.
Finally, Wallace demonstrated a live demo of the platform, showcasing features like real-time traffic analysis, heatmaps of GPU utilization, and auto-tuning load balancing. The auto-tuning functionality dynamically adjusts parameters like inactivity intervals to optimize performance and eliminate out-of-sequence packets, increasing the likelihood of successful job completion. These power packs are essentially Python scripts and are evolving, with Juniper actively working on creating more of these power packs. Juniper is also working on deeper integration with other vendors for their customers’ environments and solutions.
Presented by Jeremy Wallace, Data Center/IP Fabric Architect, Apstra Product Specialist Team, Juniper Networks. Recorded live in Santa Clara, California, on April 23, 2025, as part of AI Infrastructure Field Day. Watch the entire presentation at https://techfieldday.com/appearance/juniper-networks-presents-at-ai-infrastructure-field-day-2/ or https://techfieldday.com/event/aiifd2/ for more information.
Transcript
I'm Jeremy Wallace and I work with Kyle. I run what we call the product specialist engineering team in Abstra. And we just kind of, we're like the abstra janitors.
We do just about everything. So, um, I get to talk to you guys about taking all this stuff and now running one of these networks with it, right? So we've seen a lot of the theory, but the load balancing, we've seen a lot of the theory with how these AI data centers are built, how they run.
Cal talked about standing them up, and I get to talk about it from the perspective of the operator, right? What is, if, if I'm a, if I'm sitting in a knock, what am I going to be looking at when I'm looking at these things? So the thing, the first thing that I want to, there, there's a word that I like to get across your minds, and it's the word context, okay?
So when we build these networks, we have the full context of everything that should be happening, right? So we know exactly the way that it's built. And because we know the way that it's built, we know exactly what routes to expect and what locations.
We know exactly what interfaces should be where. So this word context is critical, right? So that's gonna come up as a, as a repeated theme.
The other thing that I wanna impress on your minds is this isn't necessarily new. It's a new model of a network for the AI networks, but it's not new what we're doing. We've been pulling this data outta the networks as abstra for quite a while in different kinds of formats, whether it's an IP fabric or an EVP and VX line network.
The AI is just a new model, right? There's some new data in there. We've got, we've got some new agents that we can deploy out on these Nick or the, the GPU Nicks to pull data directly out of the servers.
There's a little bit of new things, but what we're doing isn't actually new. So, with that in mind, these are some of the things that we're going to talk about, right? So if we're operating these networks, we know that any kind of an issue as in any other network, is going to be catastrophic.
We have a small glitch in a fabric or in a, in a cluster. It's going to cause, uh, immense issues with the training jobs to be run, or, um, customer, customer utilization, whatever it is, these issues are going to be catastrophic. There's a ton of data coming at us.
It's amazing how much data actually lives in the network, in these devices, in the Nicks, there's a ton of data coming at us. And, and as operators, I remember my days as an operator trying to, you know, sift through the CLI and trying to find where the problems are, right? It's, it's just an immense amount of data coming at us all the time.
And then taking that data and digging down and finding what is the actual problem? Where, where am I congested? Where's my load balancing broken?
Um, those, those kinds of things. So again, it's all about the context. We used to take these networks and we'd stand up all these devices all over the place, right?
So we have a whole bunch of dots, we have a whole bunch of devices and interfaces, whatever it is sitting out there. Well, if we can actually look at it and say, this is a cat. I now know how to, in, I, now, I now know how to, um, now my brain just, uh, interact with this cat, right?
I know that if I go and I touch it in this way, it's not going to like it. But if I stroke it nicely, we're gonna have a happy cat. Networks are no different.
If we go and we push the wrong configuration, now we're gonna have bad reactions. We're gonna have things go sideways. We're gonna have these, these critical issues.
It's all about the context. If we have the context of what's actually supposed to be happening, we can pull out the right data and get to the root causes of the issues much quicker, right? So now with that in mind, now that we're, we know we're looking at a cat, right?
We're looking at this AI network. We're going to look at some of the slides here. And then I'm gonna move from, I'm gonna move through these pretty quickly.
I'm gonna move into a live demo. I like doing demos better than slides. This is just a representation of the body of the cat, the body of the network, right?
So this is a call the honeycomb view, and this shows all the GPUs that are available in the network. We can break it down by different, different kinds of views. So we have the view of the overall network, alright?
We can dig down a little bit deeper and we can get down into the pipes, into the traffic dashboard. So we can look at what's actually flowing through the network. We can see where, uh, where there's buffer utilization.
We can see where there's out of, out of sequence packets, um, congestion notification. We can talk about that a little bit more in detail as well. We can get down deeper and pull this information out of the, out of the networking devices and outta the GPU Nicks themselves, we can, let's see, skip forward.
We can pull data out of the GPU Nicks as well. We have, like Kyle mentioned, we have an agent that we deploy out on the deep on the GPU, uh, the GPU servers. So now we can pay attention directly to the Nicks themselves, and we can correlate that data with what's happening upstream in the leafs, and then from the leafs into the spines, or, you know, and across the rails.
However, the, the network is built. Uh, we can look at the congestion control stuff. So we have, and I, I have a, I'll talk about this in just one second, but there's different mechanisms for congestion control within an AI network, right?
When you're talking about Rocky and someone's might ask a question, what does Rocky stand for? RDMA, which, you know, remote direct memory access RDMA over converged ethernet. So it's just doing RDMA calls over our ethernet network within Rocky.
There's a bunch of congestion mechanisms to help us notify when, when things are getting, uh, bottle bottlenecked and helping us to avoid that congestion and, and provide back pressure to, to slow things down. So we can go and we can actually configure these thresholds to, to meet whatever needs we have. Every workload's going to be a little bit different.
Every network's going to be a little bit different. Every customer's going to wanna see slightly different thresholds and, um, drop at different levels, that kind of stuff. So we can, we can, we can configure all those different, uh, thresholds and we can, we can look at, look at it to say, I want to create an anomaly.
If I have data that, or if I have, um, if I have a certain number of drops over 10 minutes, create an anomaly. If I have one drop in one second, I might not care about that, right? Or if I, if my buffer utilization gets to 90% for half a second, it drops back down, that might not be a problem.
But if it's at 90% for 10 minutes, now we've got a problem. So all of this is configurable to each individual customer's needs and, and training jobs and networks with the, the agent. We can actually look at the health of the servers themselves, right?
So not only are we now looking at the CPUs on the switches and the memory on the switches, we're paying attention, paying attention to the servers as well. What's the CP utilization on the servers? What's the ram utilization on the servers, right?
So that we can have all the way out to the far reaches of, of the environment and have it holistically, um, monitored. Okay? The goal for all of this stuff is nothing more than keeping everything moving as fast as possible for as long as possible.
And the two mechanisms that we have for that to help with that are congestion control. So we have to make sure we can get cars on the freeway, we can get the packets on the wire at the appropriate time, the appropriate speed, make sure we're being delivered to the other end. That's the congestion control mechanism.
The other mechanism is the load balancing, and they work hand in hand, right? So the load balancing makes sure all of our paths are utilized, and then the congestion controls, making sure the packets are going on at the right time, and all the stuff is working together to make sure no cars end up flying off the freeway inadvertently. You know, that's when we start losing packets in a job.
Things go sideways. We start losing the effectiveness of our, of our, uh, of our expensive servers. So the two me, two of the mechanisms that we use for, and this is just a primer.
We're not going, we're not gonna go deep into the congestion control, the load balancing that's been discussed in other field days and in other sessions. But I just want to give a little bit of a refresher on what these are because we are going to see these in the, in, in an upcoming demo here. So the, we have two different mechanisms here.
The ECNs, the, um, explicit congestion notification and the priority flow control. So the explicit, the ECNs would be like, you're on the freeway and it's getting congested, and now the brake lights start coming on going forward, right? And you're saying, Hey, we might need to start slowing down a little bit and that's going to get to the end point, and he's going to send back a congestion notification packet to the beginning, to the source, and he's gonna say, pump the brakes a little bit.
You might need to, you might wanna slow down a little bit. Okay? So that's good.
ECNs are good. That's, that's indicating we're gonna start behaving nicely with each other. PFCs is like, I'm coming up on a traffic jam and I smash my brakes and everybody else behind me gets backed up, right?
So now I'm telling everybody else behind me, you gotta slow down, but it's just, it's a hard jam on the brakes. Okay? So there are two different mechanisms.
So ideally we have more ECNs in the network than we do PFCs Question, is there also a, a management view specifically about PFC negotiation where I might be able to see the PFC negotiation failures between switches or between my endpoints? Um, we can, we can look at, we can see which device is initiating the PFCs backwards. Is that what you're mm-hmm.
Is that what you're asking? Right? Is there monitoring specifically for Cs?
Yeah. Yes. I'll show you a dashboard specifically for the PFCs being initiated.
We can see exactly which device and which interface is initiating the PFC backward Okay. As well as who's sending the ECNs forward. Okay.
Is the P-C-P-F-C that you're talking about here, is that part of the data center bridging? Is that the, the same type of, uh, priority flow control that was available through data center bridging? Yeah.
Same kind of idea? Yes. Okay.
So this is, this is, yes, yes. Short version. Yes.
Okay. Yeah. So I, the way you were describing it, it's typically you, you do that by assigning it to a particular vlan.
Correct? And, and you know, that's the way that you know, which has priority. So I wasn't sure how you could configure it other than just plain these vlan.
Yeah. So this is, we're actually gonna see how this actually applies to a queue within the, within the AI workload. So we're, we're, we're looking at this from the rocky perspective, and we can see a specific queue sending the PFCs back.
Okay? Right? So we don't have VLANs all through the network.
It's all layer three, it's all BGP signal, it's all ip, but it's all doing it back backwards, um, for the PFCs based on the, the queue and didn't Mean to do braille. You just wanted make sure that the definition hadn't changed on me. Yeah.
Okay. So a year ago, and, and I'm not gonna redo this demo, but a year ago, um, we, we presented a demonstration on what we called an auto tuning D-C-Q-C-M power pack. Okay?
So D-C-Q-C-N, that's a crazy acronym. Data center quantized congestion Notification. Notification, I think I got it right.
Um, it's just a fancy term for class of service in advanced class of service in, in the AI data center. That's basically all it is. So this auto tuning power pack, what it the, for D-C-Q-C-N, what it was doing is looking at the PFCs and the ECNs, and it would slide, make, move the sliding window for drop profiles left and right, right?
So the idea is I'm gonna get as fast as I can and as close as I can to the edge of this cliff without jumping off the edge of the cliff and dropping packets, right? So I'm gonna back up a little bit and I'm gonna move forward and get it exactly to where I need it to be. Okay?
So this is, this is a, a tool looking at the APIs, pulling data outta the APIs, looking at these PFCs, looking at the ECNs, and moving that sliding window to get it exactly right so that the AI network is moving exactly where it needs to be. The reason I bring this up is because we're going to do, I'm gonna show a video of a, of another demo similar to this, but it's another power pack that we've leveraged for load balancing specifically. So when we talk about load balancing, and you'll notice the one that's missing here is the, um, the new, the new one that RDMA based load balancing, um, excuse me, within these mechanisms of load balancing, they kind of get more advanced as you go from, you know, kind of weave through it.
And there's different places where you're gonna use different versions of load balancing. So the demo that I have specifically shows the dynamic load balancing, primarily because the network that it's built on is based on tomahawk three, asics, which supports the DLB, the tomahawk five asics supports the GLB, right? So this, when when we look at the difference between DLB and GLB, the DLB, like Vic was talking about, I'm paying attention to the quality of my local links, right?
So it's a hop by hop mechanism, local links look on the spine, my local links, GLB takes a whole bigger picture, right? So I can look at the whole end to end. What we're gonna do in this demo, when I get to that, that demo for the, the DL Bs, we're gonna see that we're paying attention to the entire network.
Go back to that word context. We're paying attention to the entire network and doing almost the same kind of thing as GLB by looking at the whole big picture and being able to move the different, um, settings, the inactivity interval specifically to where we can have optimized load balancing and, and eliminate out sequence packets, right? So that's, I, that's what I just described here.
This is what we're going to see in the demo, is this auto tuning, another auto tuning mechanism. So if you have, you have the auto tuning for D-C-Q-C-N running, that's going to manage your congestion, move that sliding window, then you have the auto tuning for load balancing running, and that's going to, to change that inactivity interval. And as those two things are moving together, we're gonna get that, that network moving as close to the edge of the cliff as possible and running as hot as it possibly can.
Okay. Any questions about any of that stuff so far? Great.
So with that, I am going to move into the RA ui here. Hopefully this works if I'm kind of angled a little bit. We got it mirrored.
So I'm, I'm going to do a live demo, and this is running directly on our AI ml pock lab that we have living in that building over there. Okay? So this is, this is real hardware, real GPUs, real, real switches.
Um, nothing's, nothing's virtualized here. This is all, this is all real stuff. And I figured this is, this is more interesting than just kind of going through some theoretical stuff.
This is, this is where I like to live. We're gonna try and jam some of this stuff into like 15, 20 minutes of demo. And I recently spent four hours with a, with a customer going through this stuff because there's so much data coming at it.
So, on that note, one thing I want to point out, jump in here real quick. One thing I want to show is exactly how much data there is. So if I come back over here, and I look at this, this tab here, this is, this is a representation of our cluster built on a 100, um, Nvidia, a 100 servers, and it's a relatively small cluster.
It's going to be eight servers. This is the, this is the danger with live demos. It'll get there.
This is gonna, it's gonna, it's built on eight servers with 16 leaves and eight GPUs per, per server. So you can see it's, it's a relatively small cluster, right? So there's 64 GPUs.
Um, it's all, it's all rail stripe optimized. So we each, each one of these leafs is a rail. If I zoom in here, we can see that each one of them is color, is that you see that well enough.
So each one of these, each, each RI is identified by a color and each, um, each one of these groups is a stripe. So you can almost, you can almost say a stripe is like a pod almost, right? It's just, just kind of a loose analogy there.
So this is what our network physically looks like. We're talking about topologies. This is what our network physically looks like.
Now this is something that I like to show, just to identify how much there is to keep track of. When I go into the graph explorer, this is looking at the actual graph database, and we show the full blue, the full blueprint. Every one of these dots is something on the network.
And every li every line in here is a relationship that all this stuff is a non-zero entity. Some you have to track it somewhere, right? Either in your brain, in spreadsheets and multiple tools, whatever it is, this is where the context comes together.
These are all these dots that make up that cat, right? We're not gonna deep dive into the graph database, just this is just kind of a wow moment to say, there's a lot of stuff to track in here, right? And those have properties and stuff on them.
Yes. That's what makes it, Yes. So if I, I can come in here and I can hover over this and I can get contextual data out of every one of these dots.
The UI for abstra is nothing more than a first class citizen, A API caller that talks the graph database. That's it. So everything in this graph database, if a, if there's something that's not in the UI that a customer wants to do, they can use, uh, whatever restful API tool they want to use to create their own API calls directly into the graph.
And we have, we have a number of customers that do that themselves. So, And the relationships that can have properties as Well. Yes.
Yes. Yeah. So, so everything in here has its properties and the, the relationships will identify, like, this one is a link that goes between a leaf four and ixia.
I'm glad you can translate that. Yeah, it's, it's kind of, it's hidden. If I click on that, I think it'll stay up.
No, it's hidden right down in here. So there you go. It's pretty, Yeah, I like it.
It's, it's cool. And, you know, we can drag things around. If I wanna look at, ah, here we go.
I wanna look at this one and then I can see where it links into, right? So we can play around with the graph database. We can look at it the real power, and this is where I am not an expert.
We have other guys who are, is you can go and you can create queries to pull the exact data you want out of it. Then you can take that query, throw it into an API call. Yeah, pretty cool.
Alright, that's enough. Took four Hours to demo this thing. How long does it take for somebody to learn And use this?
Um, it's actually easier than you would, than you would think. Okay. And you actually, I'm not a graph query expert myself, and I almost never interact with the graph.
The graph is just there, the graph DB is just there, it's just A pretty thing. So the CEO can show pretty Pictures. It's, this is what you spent your money on.
This is, this is the secret room. This is the secret sauces, like the nursery. Yeah.
And in, in reality, it took abstra a number of tries to get to the right kind of graph database to make this work, right? So yeah, it's, it, it looks, and, and I like to bring this up because this is what you have to keep track of, right? And, and it's difficult.
And if you take this and you turn this into an EVPN VX line network Gets even bigger. Okay? So with that, I wanna step back and show this.
So now here we have all these different networks that we're managing with a single instance of abstra. We have one cluster that is based on EVPN, right? So this is where our H one hundreds and our, um, A-A-M-D-M-I three hundreds live right now.
They're in this EVPN. We're building multi-tenancy. JVD, uh, Juniper Validated Designs based on EVPN vxlan.
We've got the different, um, o other backend fabric clusters. So the, the A one hundreds that I'm gonna be showing you are based in this cluster one. We've got a front end fabric.
We've got a backend storage, storage fabric all managed from the same instance. We understand the context of all these different networks and what's required in all these different networks. Okay?
All right. So when I go into my cluster one dashboard here, we're presented with what an operator would be looking at. The operator would be sitting there in his cubicle looking at pretty screens all the time, right?
And ideally, and I'm gonna come back to these probes here. I'm not, I'm not glossing over that. Ideally, when he's in here looking at this, this is what it looks like because we have the reference design, the best practices, we push out the correct, um, configurations.
And one story that I actually like is, one customer called us up one time. They were like, what do I do with this? I never see anything red anymore.
I don't have any problems. I was like, well, that's because we pushed out the correct, and this was an EVPN customer. This is some years back.
This is an EVPN Von customer. That's because we push out the best practices that you need for your environment. They're like, ah, got it.
So that, that's a real story, by the way, that was, I, I talked to that customer personally. But this is ideally what you see. And we have representations of our BGP connections, our cabling, this is what Kyle talked about, the cabling map.
If we have, if someone goes in and, and replaces a server and puts a cable in the wrong switches, the ports, we'll, we'll tell you that, hey, these cables are switched, they're reversed. Um, there's ways we can automatically remediate that, or we can tell 'em, nope, go physically remove it. Or we can just use an LLDP discovery and pull that back in and fix it dynamically.
So there's different, this, this, this is what we want to see. All healthy, all green, right? User.
Uh, before I jump into the probes, this is the interesting part. And I can tell you that something is now broken on our ixia, because I've been playing with this enough. But this is the honeycomb view, and I showed you a screenshot of that a little bit ago.
This represents every GPU in the fabric. And I can take this and I can group it in, in different groupings. So if I wanna look and see what each one of my physical servers itself is doing, this is now per server.
I can do that. I wanna see what my rails are doing. I can look at each rail and I can see that I've got some rails misbehaving.
I'm going to explain these colors here too. This is a heat map of our GPUs. It's not a status indicator.
So if you think of a heat map, the hotter it is, the redder it is, the better it is. This indicates how hot our GPUs are actually running. Ideally, we want them all to be running at this, um, this brown, like 81% and higher, right?
The hot, the closer we can get to a hundred percent, the better we're going to be. The, the, the better our training model is going to run. The quicker the job completion time is going to be.
I happen to know that the ixia is broken right now because this is the exact traffic pattern that we saw when the Ixia broke couple days ago. So the Ixia is doing something weird. We're sending a bunch of traffic with the Ixia creating a bunch of congestion.
I'm gonna show you where the out of sequence packets are and stuff like that. So, but what this indicates, and this is actually cool, I actually like this because now that if I'm an operator, I can look at this and go, ah, crap, I've got some of my GPUs are not operating at their peak efficiency. What's going on?
Um, right now I'm sorted by rails. I can take this and sort it by servers and I can select a specific server. I can see these, I zoom in here.
Um, and each one of these GPU zero GPU one, that is a rail. So zero rail, zero rail one, rail two, I can see that rail one and rail three, there's something broken there. Now I can tell you that rail one happens to be a switch that all the ixia portrait coming in on.
And rail three is the other switch where all the ixia portrait coming in on. So it makes sense. The ixia dumping traffic in there, creating a lot of back pressure, a lot of congestion, and that sending it back to the, the servers.
Okay? But if I'm an operator, I can look at that and go, dang, you know, what's going on here? So when I come up and I look at my probes, this is going to give me an example of what's going on in the network.
And this is gonna tell me, there's gonna be a whole bunch of out of sequence packets and a whole bunch of, uh, ECN frames and, um, things like that going on. Quick, quick question about how you're getting that data. I'm assuming it's some sort of telemetry subscription, and you're just getting the telemetry off the ASIC for that, correct?
Correct. There's a combination of RPC calls out to an agent. So that's actually a great question, and I'll take a step back and explain that.
In abstra, we have the abstra, uh, here we go. This is the danger of real demos. There we go.
Okay. In, in appra we have the abstract core server, and then there's agents that live out in, in either on the devices themselves or that's an on box agent. Mm-hmm.
Or we have an off box agent that sits as in a small docker container with the abstra server. The ideal is to get them on the on box out to the devices. So we're, we have, and then we have a, uh, either an RPC call that we will go out and pull data.
Mm-hmm. It's not SNMP, just an RPC call, or we'll do GRPC streaming data back in on some, some of the, some of the counters, right? There we go.
Okay. So here we, we have a whole bunch of outta sequence stuff, and we can look at this and see, here's a system id, and we'll correlate this a little bit easier here in just one second, because when I look at a serial number, that doesn't mean much to me, but I can see that on GPU three eth, I've got a whole bunch of outta sequence packets coming in outta sequence packets is indicative again, of core load balancing somewhere, right? I've also got CMPs going on, and the CMPs is indicative of congestion.
So I've got something going on within my network where I've got both outta sequence packets and I've got CMPs. So I've got bad load balancing and bad congestion happening. You know, it's, the ixia is doing all kinds of crazy stuff to, to make this a mess.
So if I go back over here to my dashboard, just gonna scroll down and show some of the parameters that we then, that an operator would go and look at as he is trying to figure out what's going on. So we can create, we can pull this data outta the network, and some of these graphs are based on seven day intervals. Some of them are based on, um, one hour intervals.
And the user can go and create these graphs based on what they wanna see the intervals they wanna report back up to their VPs, right? Hey, over the past seven days, we've had this, this much bandwidth in our AI network kind of a thing. So as we scroll down and look through this, I'm just gonna show you some of these graphs that we've created really quickly.
Um, we've got spine, spine traffic, and if you hover over these, you can actually see the, the amount of, I'm not supposed to leave the white line. You can see the amount of bandwidth on, on each interface on the spine, right? So one of is operating at like 400 gigabits per second.
That's, that's cool. That's good. That's what we want.
Um, and same thing, you can go down and here's, here's our leafs, our leaf for stripe one and our stripe two leaf. So we can, we have this broken out now per stripe. We're getting just a little bit more granular as we keep scrolling down through here.
We can look at total traffic over the last seven days. So you can see you've got bandwidth. Um, this, you know, and here's, here's where gaps where the job stopped running.
Ideally, it's all just running really hot, right? 95 Tera, right now, if I come down, I get a little bit more granular, now I'm gonna look at my intra stripe. 1 terabytes.
76 terabyte. 76 and the three, you know, and that adds up to the total traffic, right? So we can keep getting a little bit more granular as we keep moving.
And these, these dashboards can be sizzled and moved around. However, however an operator wants to see it. Yes.
Hi, Denise Donahue. Can you set it up to notify you proactively when things start creeping? Yes.
Yes. There are, there are different ways to, to set that up. Um, yeah.
In fact, in the, in the demo that we did last year with the congestion notification, they actually had that set up with ServiceNow to where the congestion notification would pop up and ServiceNow would open a ticket so the users would get that scrolling right through ServiceNow. And as things changed in service in the, in the network, the ticket in ServiceNow would change. And when the congestion got relieved or got completely fixed, the ticket would get closed in ServiceNow.
Do you provide any integration with the other third party, Uh, observability stacks or monitoring? Um, the, the, the, we don't necessarily provide that DI or from, from Juniper's ourselves, but that's all available with the rest APIs. So there are other customers doing things with other kinds of observability platforms.
Okay. Yes. We do have a telegraph plugin Yeah.
That we can use. And we work with Grafana a ton. Grafana.
Yeah. So that, that's all there. And then, and, and of course we have the, the flow where we can get lots of flow stuff and all that can go to different places.
So yes. Okay. Gotcha.
Other questions? Okay. Yeah, maybe a question, um, um, I'm looking further as well to miss that one.
Is there also, because notifications is always a good thing if there's a problem, but would it be better that there's also a recommendation engine saying that, hey, there is a problem and we advise you to do this or that Yes. Switch the button happens, so Yes. Yeah.
Push the button that it happens. Yeah. So with the, with the cloud services, with the missed piece Yes.
Um, that, that is going to be more part of that mm-hmm. Where you have a, a, a, the, um, AI ops engine running up in the, in mist and we'll be able to have different kinds of things you can visualize in there. We can already do that with the EVP and VX line stuff where you can, you can look at a service and you can see traffic flows through the network with a service.
You can see what would happen if I, it, it's called a failure analysis. What happens if I create a break here? You can do things like that and it will create, or it can create, it can create recommendations, but so That's another UI that I need to use then.
So it's not in here yet, Correct? That's not part of abstra. Okay.
That, that would be a different window that would be running and you can launch ABSTRA from a CS. Oh, Okay. So you can, you can click on it, it'll say, this instance of abstra is, is reporting this issue, click here to go do this thing.
Okay, thanks. Yes, Jim, rinky, CDC, the whole concept of a rail, I'm still trying to get my head around that. And is there a but to elucidate on that, is there a specific thing you would monitor a rail for instead of some of the other things here?
Yeah, so let me just jump back over here. So I'll show you, I'll show you a comparison of what it looks like with a rail versus a non rail. Okay.
So think of, let me just zoom this in a little bit better so we can get a little closer. Okay. So if you, if you look at these, you can see that you, I have, I've got four purple links and four blue links, right?
Gotcha. Okay. So a rail is nothing more than every server has eight nicks.
Gotcha. I've got eight leafs. Ah, Nick zero goes to leaf zero, Nick one goes to leaf one.
Got it. So rail zero is everything going to that one leaf. Got it.
Okay. That's, that makes sense. That's Somewhat simplified.
You can, you can do it where if I have a leaf, maybe I have two links in a rail, you know, so there's different things you can do there, but that's kind of the idea behind it, is the next map to the same place. Okay. And so would there be a, there's obviously specific metrics that you're watching there, and if you saw a failure on a rail, that would indicate what typically that something's wrong at the, in, at the server at the nick, or It depends on, it depends on, I mean, the failure's going to be wherever the failure is, right?
I see. Kind of, kind of a dumb statement, but if, and I don't see that necessarily too terribly different from any other network if I've got, so let me show you what I mean by that real fast. Please.
If I can zoom this out. I'm just gonna close that window there. Okay, let's go back over here to, um, our blueprints.
So if I look at this particular instance, the, the EVPN instance, this is where, like I said, where the H one hundreds and the MI 300, the A MD as six are, if we look at this one, this is our multi-tenant. So this would be like your GPU as a service cluster. Okay?
Right. Inside here you'll see that this one is not built necessarily in a rail and stripe fashion. Okay.
But if I have a GPU failure in a rail, it's still a GPU failure. If I have a, an interface failure, ah, I'm, I'm just gonna get it reported in a, in different language. So it'll say, um, stripe one, rail one interface ET 0 0 0 Gotcha.
Had a failure. Okay. Versus here it's gonna say it's just The classification, essentially.
Exactly. Okay. We're, we're still working with Rocky between the GPUs, right?
So we're still looking at the, the f the PFCs, backwards, ECNs, forward CMPs, we're still looking at those same kinds of things. Okay. It's just the terminology is going to change a little bit.
So it's Just the topology terms is really what it comes down to. Yeah. Got it.
And one thing, and one thing that we are actually working towards, it's not in Apria Yeah. But we're working towards, is having rail optimized and EVPN together. Okay.
Okay. So now imagine that one, okay, Uhhuh, That's not the system blowing up. That's like my mind just blowing.
I, I work with this stuff every day and I'm still blown away by the things that we're moving forward, the, the speeds we operate at. Um, I actually loved our hashtag and this, this kind of goes along with what we do with abstract. I love the ha, the, the tag.
We used to use engineering simplicity, Uhhuh, Get it, get it. Mm-hmm. This stuff is not simple.
So we, we try to take, and, and there's a guy that I used to work for, and I, we were in a meeting with a bunch of people talking about APIs and all this stuff, and he just stopped the meeting. Prol might remember this. He stopped the meeting and he goes, hold on, nobody cares about this piece under the iceberg.
What's the very shiny tip that our users are going to interact with, right? Mm-hmm. This is abstract up here.
There's a whole bunch going on underneath here. This is the important part. This is what we interact with.
That's how we interact with all these other, right. You know, think about the cat, right? We interact with the eyes we have with the fur.
We don't, we're not touching the intestines in the heart, but we know it's there and we know it's important. We gotta keep it healthy. We've All had the experience of a user that we really respect walk up and go, it's running slow.
We get it. Yeah, Exactly. And so, so, so back to this, as we, as we get a little deeper down, it's running slow, right?
That, that is a, it is such a common thing that we don't know. It's slow until three days later, right? And someone comes to us and says, Hey, my job ran slow three days ago, right over the weekend.
The cool thing, one of the cool things about abstra is we can do all this stuff in a time series database. We can keep all this data, and I can take, and I can run a dragging slider back and forth. It's actually something that I'll, that's I think is kind of cool.
I'll show you real fast here. Um, if I go to the, the active tab here, and we look at this heat map, um, so with this traffic key, I can turn this into a time series. And I don't, this is on the, this is on the EVPN cluster, so I'm not actually sure what's there, but I can take, and I can drag this back and forth and I can show, and they can say, Hey, Saturday morning at 2:00 AM there's a noticeable slowdown and we can go and look at the network and it will show us exactly where congestion happened, where bottlenecks happened, right.
That, that kind of a thing. So it was Probably a DNS failure. Yeah.
It's always dns, it's always DNS, right? Yeah. It's always DNS Now, uh, because it holds a lot of data, um, and it collects it from all those devices, um, and so on and interfaces.
Where does it store the data? It's gotta be a really big database. Is it in that apps press server?
Is it somewhere else in the cloud or a combination of how long do you keep the data? That's a great question. So on, on the APPRA server, we keep 30 days worth of data uhhuh.
If you want to keep a year or two years or more worth of data, then you export it out to something like Grafana Telegraph or up into the cloud services, and we can keep the data longer in, in the cloud on the server itself, in order to keep it from becoming like terabytes in size. We just keep 30 days worth of data. Because it's a vm, right?
Correct. It's a vm. Okay.
Correct. Yep. So That's 30 days of completely non aggregated or compressed data, Correct.
Yep. Granular. Yep.
That's all the raw data that we pull from all the agents, all the RPC calls, everything that's coming into us. We keep that for 30 days. Yeah.
Okay. In the interest of time, I'm gonna jump over to this, um, auto tuning load balancing demo. And I'm just gonna start this.
And this is playing at two x speed. So this is, it sped up a little bit and you're gonna see what's, what, what's happening here is initially the inactivity interval timer was set for 64, right? And so it's set really, really small.
And when it's set really, really small, we're gonna start seeing a lot of out of sequence packets. And as, as we're monitoring the network for these outta sequence packets, we're gonna start backing that number up to get to where we eliminate completely the outta sequence packets. Um, the way to think of the inner, the inactivity interval in, um, in DLBI, I, I like to think of it as like the professor timeout interval, right?
We've all been there with college or high school when you walk into the class and you're like, how long do I have to wait for the professor? Right? How long?
And, and if you wait for 30 seconds and you bail out, that's too quick. You might, you're probably gonna miss something. You're gonna miss a professor.
Something important is gonna happen if you wait for too long and you sit there for three hours. Well, now, um, you know, who knows what's changed in the world of the news in the past three hours and all you, you miss all kinds of errors and all kinds of things happen. The inactivity interval isn't really any different than that.
If it's set too small, I'm going to create my own problems. There's, there's going to be, uh, out of sequence frames, I'm going to receive things on the wired incorrectly. If it's set too big, I'm going to miss the opportunity to detect the actual problems on the network, right?
So this is the whole purpose of this is to monitor the network continuously and then make incremental changes. This is like, I'm touching the whole cat now, right? So I'm looking at the whole cat and saying, okay, there's, there's a little piece here that's unhealthy.
Let's holistically heal the cat, um, kind of a thing. So I'm just gonna speed this up a little bit. You'll see right now where we've got the, um, the, it, it's the old, the out sequence packets is at 1675.
There's a value up there. If I click forward a little bit here in time, the value is now dropped down to six outta sequence packets. It might jump, but back up as, as the timer's going back and forth.
And the idea with this auto tuning mechanism is to get this down to where it is actually down to zero out of sequence package. And now we have an ideal inactivity interval for this specific job. The goal here is not necessarily to be faster than every algorithm in the world or every, every computer in the world, the goal is to be faster than humans, right?
If, if we're sitting there looking at this and I'm waiting for an outta sequence packet, I now have to go and touch every device I have to figure out where it's coming from. The goal is to be faster than humans. So if, if we start a job and this job is going to run for three hours, or it's gonna run for three weeks or three months, whatever the job's going to run for, however big it is, if I can resolve some of these things, the congestion, or I can resolve the load balancing within the first couple minutes of that job running, I dramatically increase the odds of my success of the jaw brain to completion in a time that I want it to run.
Right? So that's the, that's the idea behind these auto tuning things. And one of the cool things that we can do with abstra is, and this is something that we're work actively working on within my team, is building more of these power packs as we're calling them, that we can move features to ab extra just a little bit quicker.
Mm-hmm. Right? And then we look at 'em and we say, okay, this is a good one.
Let's pull this and actually pull this into the product. The power packs are just Python script. This one's a Python script.
Yep. And the, the congestion mat notification is a Python script as well. Yep.
How often are you seeing like, rate of change in this particular perspective? And, and what I mean by that is how often am I, am I reconfiguring the cluster enough to have additional problems created for myself? Or are you saying this is more of an enterprise perspective where I maybe have a cluster that I don't mess with very often?
Um, what would you say is like the give and take on rate of change and how often I'm redesigning a cluster? That's a good question. Um, probably something we'd have to take more and discuss a little bit more offline.
I, I don't have, I don't really have good data behind that necessarily. Um, ideally not very often. Right.
You know, ideally you get it to where the, you, you solve the problem with a few configuration changes at the beginning. Maybe it's just with an inactivity interval timer. Mm-hmm.
Ideally, you're not touching it a whole lot, you know, so I've Built it day zero, day one, day two. Right. And I'm running it.
Yep. I'm just kind of doing a, a reactive monitoring at that point, right? Yep.
'cause I'm not rebuilding my cluster very often, so I have to imagine it gets into a stable state pretty quick, right? Correct. That's the idea.
Now, the, the place where I can see it changing is maybe I run job type A and it has certain size of packets and certain inter um, inactivity interval, right. Or the, the congestion notification pieces, or have a certain place in that sliding window. Job B, it slides a little bit to the right here and the inactivity interval slides a little bit down here, you know, so having those things running constantly, I may have to make a configuration change here, depending on the job that's gonna change.
Right? Could potentially. Yeah.
Okay. Yeah. The idea is that every customer's gonna be a little bit unique.
Just like every cat's gonna be a little bit unique. You're gonna have a little bit different requirements for food. This one's gonna make 'em sick, this one's not.
So changing those things when needed as early as possible, and then having it just set stable is the idea. Yeah.