Day 1: Managing your AI data center at scale with Juniper Networks
This presentation by Kyle Baxter focuses on how Juniper Networks’ Apstra solution can manage AI data centers at scale. Apstra simplifies network configuration for AI/ML workloads by providing tools to assign virtual networks across numerous ports, an essential capability in environments with potentially millions of ports. The core of the presentation highlights the ability to provision virtual networks and configure load balancing with ease, using an intent-based approach that simplifies complex network tasks. This reduces the burden of manual configuration and allows users to quickly deploy and manage their AI data centers, regardless of the number of GPUs.
Through its continuous validation features, Baxter demonstrates how Apstra allows users to pre-emptively catch configuration issues, such as missing VLAN assignments. This prevents errors before they impact operations. Furthermore, the system enables bulk operations to streamline the assignment of virtual networks and subnets across the entire infrastructure. The solution includes options for selecting load balancing policies, with clear explanations through built-in help text. Juniper focuses on simplifying tasks through an intuitive interface to minimize the need for command-line configuration and facilitate a faster and more efficient deployment process for AI data centers.
Finally, Baxter mentions the delivery of Apstra as a virtual machine, with the solution available for download on the Juniper website. It can be deployed on-premise to manage a network and integrated with MIST AI for AI-driven network operations. The upcoming release of version 6.0 and support for advanced features like RDMA load balancing were also discussed. The key message is that Juniper’s Apstra enables efficient deployment and management of AI infrastructure at scale, regardless of the deployment size, simplifying complex tasks with its user-friendly interface and automated features.
Presented by Kyle Baxter, Head of Apstra Product Management, Juniper Networks. Recorded live in Santa Clara, California, on April 23, 2025, as part of AI Infrastructure Field Day. Watch the entire presentation at https://techfieldday.com/appearance/juniper-networks-presents-at-ai-infrastructure-field-day-2/ or https://techfieldday.com/event/aiifd2/ for more information.
Transcript
Welcome everybody. My name is Kyle Baxter. I'm gonna be talking about how in, in deploying and managing your AI data center, how you can do that at scale with ra.
So the, the challenges is there's potentially lots of supports. How do you assign those virtual networks that you need to be able to, to run the train jobs across all those ports? You know, you don't wanna do that manually one by one by one by one.
'cause there could be thousands or millions of different ports when you're talking about some of these larger, um, deployments. Um, and then there's ever evolving requirements. Um, there's, there's new techniques they're coming out for things like load balancing.
Um, we saw in a, in a previous session about some of the new, um, RDMA based load balancing that's coming out. Um, but what also about how do I configure load balancing without having to go read through hundreds and hundreds of pages of articles to become an expert? And that's what we can do within Astra is simplify that and give you just a few click operations to, to get those things assigned and up and running.
So starting with virtual network provision. So how do I assign virtual networks across my entire network? So the first thing that, that we talked about a little bit earlier was, you know, how do we catch things before we deploy?
So that is exactly one thing that we can do here that I'm gonna show you, um, is the first step is we've determined before you've deployed this production in that staging version of abstract, that there is warnings that there's missing VAN assignments to rails. And so we've caught that before you've actually deployed. And what we can then do is give you an easy button to say, go ahead and provision A-A-V-A-N across those rails and those ports.
And then afterwards we can see, yep, I'm all set. We can see the, the virtual network, the connectivity template all assigned to it. So what this does is it simplifies that burden of managing and gives you the ability with just a few clicks to assign virtual networks, make sure your connectivity templates are all connected to the ports and the rails that you want and you need.
So let's take a quick look at this. I'm gonna move to this demo real quick and we'll see this in action. So we're looking very similar to those screenshots, but we're gonna go here to that uncommitted tab, which is showing what's the difference.
So what we played around with in staging between what's in staging in our production, and we can see that the warning tab is highlighted. And so when we click on that warning tab, we'll immediately see that there's interfaces associated with a rail that are expected to have a VLAN for untagged traffic. And it's not there, but we have resolutions where it says, Hey, do you want to assign that vlan?
And yes, I do. I could do it one by one or what I'm gonna do is I'm gonna look at all of my rails and say, let's bulk do it. Let's, I don't wanna do it one by one by one by one by one.
I can do it, but it's tedious. Why do I wanna do multiple things at once when I can do it in one click? And so I'll select all of my rails and say, let's provision be lands across all of them and we'll see that we can review it, we can look at it, they were all missing before, but we'll add 'em in there and we'll we'll see that, uh, column go from missing to now they're all assigned.
Now the other thing after does is it still catches that there's still some red in there. So we're gonna go over there and the one thing that we didn't do is tell it what virtual network, what subnet to use. Hmm.
So I'm gonna pick an IP pool. I'm just gonna pick one of the default ones. I can see there and there how much of it is used.
So I know there's some, some space in there. And I'll pick that IP pool. And so then you can see the page all of a sudden start everything turning green, showing me that it's validated.
This is that continuous validation that we're doing an app store that's continuously looking to see, did you do everything that we, that you need to do? And so it caught that, you know, we, we didn't have the, the BLANs set and we didn't also have, um, a virtual network, um, set for a subnet of IP dresses. But once we've fixed all that look, a warning tab now went green.
And so there in just a few minutes and a few clicks, I was able to assign a virtual network and an IP subnet for that, for a virtual network across all of my ports and and rails. So really cool to be able to see how it catches that, how it does that continuous validation and helps make that smooth. So I'm gonna switch back over to the slides and we'll look at load balancing.
So a question came up earlier about how do I do load balancing? Um, and in Astra we have the same kind of simple intent based driven models where we're looking at what is the intended outcome that you want? What do you want it to do?
Well, how do you want it to work? Not let's go switch by switch and configure DLB or GLB. It's do you want DLB or GLB?
Do you want flow that per packet? You know, those kinds of things. And we'll see how that, how that works.
And we can go through that. So there's a simple walkthrough configuration where you can go and you can pick, do I want DLB? Do I want it to be flow lit per packet?
Do I want to set some of the inactivity in the, in the intervals, but in there, and I'll show in a quick demo, you can hover all those, those little question mark helps at the end. It's a little small, but um, you can hover all of those and you'll be able to see what do all those parameters mean. So you don't have to go research and be like, um, well what is the inactivity interval?
I don't really remember what it should be. It'll tell you right there when you hover over it that Hey, this is what it is. This is the default value we set.
If you wanna change it, you can, if you wanna keep the default value, great. And we can also do validation on there's specific things that maybe only work on certain hardwares. So like GLB only works on, um, the 52 40, um, because of the specific hardware it needs, specific hardware, we can do that validation.
So you don't just, you know, say, yeah, I want GLB and you don't have hardware that can do it. We don't let you deploy it if it's not gonna work. So we can do that validation.
Is RLB handled separately or is that coming soon? Coming soon. Okay.
So it's some of the, the latest innovations coming out. Mm-hmm. Um, so we have DLB and GLB in the product.
Some of the great things about abstra is we have flexibility to do things like that. Um, separately from our intent based models, we can do what we call confit. So you can do custom configurations when there is, you know, like brand new innovations that come out on a switch hardware.
Um, so if you wanna keep up with the latest and greatest, but we're actively working as we speak to get things like the, the RDMA load balancing, um, into the, the configuration that we're gonna see right here. Got it. Thank you.
Mm-hmm. So how this looks like is, is a few steps. So starting with, we assign a default, um, policy that's, that's DLB across because that is, that is on all the hardware.
So we assign that default policy, but if you want to build your own, it's only two steps. First you create a load bouncing policy and you go through the selections on what options do you want. Again, we have the help techs that tells you what all they are, where all the default values, and then you then assign it and you can do it just like we saw before, one by one.
Or you can do it in bulk. So if you wanna do 'em all at once, you can do that. If you want different ones, maybe different ones on the spine versus the leaf, you could have that option where you can be able to do them across different ones.
So let's see this in action real quick. Um, so actually go here and then pull up this one and we'll walk through it. So it's in our staged again, we'll go to fabric settings where we can look at the load balancing.
And as, as I mentioned, there's a default policy that comes already assigned that's running DLB. Um, but I wanna build my own. So let's do it.
So the first thing we go to is over here, the load bouncing policies. We will create our own load bouncing policy. We give it a name, call it my policy, or whatever you want to call it.
And then we can go through the options. And here's where I was talking about the help text that you just hover over it and it tells you exactly what those mean, what are the default values. So you can see do I like that default value?
Do I wanna change it? What do I wanna do? I can pick, do I want GLB and helper load balancing.
You know, you can pick all those settings and then you can simply go over to the assignment. And as I talked about, you can do it one by one. So we could pick, you know, just one by one here or if I wanted to, and this is what I'll do 'cause I don't like doing things one at a time.
I wanna do it in bulk. I want the quick, um, I want to get it done fast. Um, I can do that in bulk, assign it.
And now all of a sudden here in just a few clicks, we've created a new load bouncing policy. I didn't have to be the expert knowing what a deal, what are all, all the settings to know? Do I, you know, what is an activity timer?
It was already there for me. I just made a couple selections and I'm now off and running. I have a question.
Uh, it looks like on every leave and every spine switch, you can set different load balancing policies. Yes. Is there a reason why you wanna do that?
It's the same network. You have different policies. Probably not, you wouldn't be advised that, um, you know, some of them like GLB maybe you want, um, different at the spine versus the leaf.
Um, there's, there's a little difference there that you might, uh, want there, but in, in general, yes, you wouldn't probably want to. Um, so that's where the bolt comes in. Bang, is you, you just want everything to have the same load balancing policy.
Um, but if you really wanted to get crazy, you, you could, but I wouldn't advise it. Okay, Thanks. Uh, Jack Poller paradigm Technica speaking of getting crazy.
Yes. Is everything exposed through this interface or is there still stuff that you have to drop down into command line and tweak and lunge and stuff like that? For, for things like DLB and GLB?
No, it's, it's all configured there. We saw there was a big list of items you could manually configure. So, so for those, no you don't, um, if you wanted to use like the, we talked about just a second ago, the, like the RDMA load balancing, we don't have that yet modeled yet.
That's coming soon. Okay. But if you wanted to use that today, you would then have to drop into what we call confit to, to help set that up.
But the, the, the, the point is you're going to capture everything that you possibly can in this tool. Yes. And stray will stay away from Yes.
Old style Yes. School stuff. Yeah.
The goal is to stay away from the C-L-I-C-L-I. Okay. We can still show you, as I talked about in the earlier part, right, right.
The, the rendered config. If you still love to see it and you wanna see, you know, did it match exactly what I thought? Mm-hmm.
Um, but, but the idea is, is yes, you'll drive everything through the UI or like we talked about, um, earlier APIs via like rest, Terraform, Ansible, things like that. You can drive it all through there if you wanted to, to help automate that. Right.
Okay. Thank you. Yeah.
So to let me go back to the, to go back to the slides. So to sum up, we saw how we can manage this network at scale from, from deploying it. And it's the same process no matter if you have one GPU 10 GPUs, a hundred GPUs, a thousand GPUs, a million GPUs, no matter what it is, it's the same process.
Um, it wouldn't take me any longer if it was, you know, a million GPUs. 'cause I could do that bulk edit thing for everything and, and have it all set up. So we can take that complexity, remove that, that burden of complexity out there, give you the expert to help you deploy your AI data centers faster.
So any last minute questions before we go to the Yes. Oh wait, Before we go to, I wonder if the next section's better to ask it or you're done not to this. I am, I'm finishing now and we're gonna be moving on to, um, kind of the day two, the day-to-day operations.
And we'll see a lot about the, the visibility we get in, in monitoring the networks, heat maps, all that fun stuff. Okay. So my question is, how is this delivered to customers?
How do they get it? Mm-hmm. What is, how is it consumed?
So how much of it is through si? How much do you get it from, uh, MSPs? How do they get directly from you and how much of it is bundle the complete solution?
I'm just curious how it, how it gets to them and how they consume it once it gets to them. Yes, yes. Great question.
So thank you for asking. Um, so Abstra is delivered as a virtual machine. So on Juniper dinette and our downloads page, we have OVAs, KVMs, um, Microsoft HyperV.
Um, we're looking at adding Nutanix versions as well. Um, 'cause that's getting popular now. Um, so it is deployed as a virtual machine.
So almost all of our, our users and customers, they get that OVA, they deploy it inside their, their network, inside their data center in the management plane. So it has that management connectivity to all of the managed witches. Um, so that's typically, that's deployed.
There's, there's some cases where, um, you talked about like SIS or MSPs where as Part of the bigger solution, they come into your data center. Yeah. And they'll kind of, they'll help deploy it for you.
Um, you know, same thing like professional services could do to come in and help, you know, deploy it for you. Um, but it is traditionally, um, delivered that way as an on-prem application. Um, we do have, as I talked about in the very beginning, in the previous session, was the integration with, with missed ai.
Um, so we do have a component that we can take a lot of this data and a lot of data we'll see in, in the upcoming session and be able to build AI ops on top of it. So, um, AI for networking, um, where we can use AI and ML to help improve network operations. Um, so we'll be talking about that probably in, in other sessions.
'cause today was focused on the, the building of the infra. But we can absolutely do that to be able to bring more AI and insights and be able to help you troubleshoot faster. So more direct sales versus going through partners.
Well, we still have, you know, partner channels, um, but even through a partner channel, they'll still get the software and deploy it, you know, on premise. Okay, thanks. Mm-hmm.
When do we expect six O to be Released? Thank you for asking. Should be in the next couple weeks.
Excellent. So we are in actively in early trials with several customers that are actually using this and deploying it in their own AI and ML trading jobs as we speak. Got it.
But it'll go ga here in, in about the next couple weeks.