Maximize AI Cluster Performance using Juniper Self-Optimizing Ethernet with Juniper Networks
Vikram Singh, Sr. Product Manager, AI Data Center Solutions at Juniper Networks, discussed maximizing AI cluster performance using Juniper’s self-optimizing Ethernet fabric. As AI workloads scale, high GPU utilization and minimized congestion are critical to maximizing performance and ROI. Juniper’s advanced load balancing innovations deliver a self-optimizing Ethernet fabric that dynamically adapts to congestion and keeps AI clusters running at peak efficiency.
The presentation addressed the unique challenges posed by AI/ML traffic, which is primarily UDP-based with low entropy, bursty flows, and the synchronous compute nature of data parallelism, where GPUs must synchronize gradients after each iteration. This synchronization makes job completion time a key metric, as delays in a single flow can idle many GPUs. Traditional Ethernet, designed for TCP in-order delivery requirements, doesn’t efficiently handle this type of traffic, leading to congestion and performance degradation. Solutions like packet spraying using specialized NICs or distributed scheduled fabrics are expensive and proprietary.
Juniper offers an open, standards-based approach using Ethernet, called AI load balancing, which includes dynamic load balancing (DLB) that enhances static ECMP by tracking link utilization and buffer pressure at microsecond granularity to make informed forwarding decisions. DLB operates in flowlet mode (breaking flows into subflows based on configurable pauses) or packet mode (packet spraying). Global Load Balancing (GLB) enhances DLB by exchanging link quality data between leaves and spines, enabling leaves to make more informed decisions and avoid congested paths. Juniper’s RDMA-aware load balancing (RLB) uses deterministic routing by assigning IP addresses to subflows, eliminating randomness and ensuring consistent high performance, in-order delivery, and non-rail performance without expensive hardware.
Presented by Vikram Singh, Sr. Product Manager, AI Data Center Solutions, Juniper Networks. Recorded live in Santa Clara, California, on April 23, 2025, as part of AI Infrastructure Field Day. Watch the entire presentation at https://techfieldday.com/appearance/juniper-networks-presents-at-ai-infrastructure-field-day-2/or https://techfieldday.com/event/aiifd2/ for more information.
Transcript
In the AI Data Center Solutions report peripheral. Today I'm gonna talk about how to maximize, you know, ai, um, cluster performance using juniper's, uh, self-optimizing networks. So, um, I know earlier you had asked like how, uh, you know, can you show a topology?
So let's look at like, you know, what are the challenges? How AI application traffic is differs from traditional data center, right? So, uh, most of these, uh, Rockies over UDP, and also what happens is, uh, there is very low entropy.
So there are very few flows, there are busty flows, high bandwidth, and, uh, some of them live throughout the training, right? So that's the first challenge in how the property of the flow. And the second one is if you see training where when you parallelize it over multiple GPUs, hundreds of GPUs, um, one kind of parallelism is data parallelism, which is very popular in that case, GPUs are lockstep.
So they all load their batch of data that they're going to train on, and at the end, end of that iteration, they have different gradients that they have to synchronize amongst each other. And until they do that, um, you know, they cannot move to the next batch. So they're kind of idle.
So that's why there is that synchronous compute, um, during training that, uh, you know, even exacerbates this, this problem. So a single flow that is delayed in this synchronization could really, you know, keep a lot of GPUs idle. So that's why job completion time is a, is a key metric.
Now in this diagram about what you see is it's a physical topology strategy that you say, Hey, uh, I can ease the burden of my load balancing. So this is, Nvidia has been prescribing this, um, and it's called rail optimized. Um, so the first nick of every, so every PU has its own nick, and the first nick or first GPUs are connected to the same switch.
Um, so there are eight GPUs in the DGX form factor or similar, and you see eight switches or eight rails being created. And then, uh, you have the spine layer in the leaf to spine. There is one is to one or subscription to just handle this capacity.
Um, as in these are, uh, physical topology, uh, strategies that will give you the best chance on, on, you know, uh, to keep the performance up. Now, uh, profu alluded to this and, and, and he said he, he, he mentioned that, okay, traditional in the traditional, um, or ethernet or what we call static ECMP, that catered primarily to most non-AI applications that were built on TCP and TCP said, Hey, just give me, uh, in order packets, uh, delivery, right? And that was because, you know, TCP has this exponential back of whenever, you know, packets are out of order.
And also it has a slow start. So the bandwidth hit whenever that happens is pretty, uh, bad. Uh, whereas in Rocky's case, there is no slow start.
There is, uh, you know, all these flows are always bursty. They come in, uh, you know, and are long li So the, the traditional ethernet, which was like do a hash and pin a flow for its lifecycle on a path, which was tcps main ask, um, worked well in that space because there was so many flows. There was enough entropy, some long lived short lip.
So a natural distribution or a random distribution worked well in this case, if you apply that, you can get in, uh, even if you have one is two, one or subscription, you will get in, uh, trouble where, you know, um, you can make forwarding decision or load balancing decision, uh, at the leaf and over subscribe a a a a leaf to spine link or two leafs independently sending traffic could overload a spine to leaf link. So really what we need is more efficient load balancing that is designed for handling rocky traffic. And again, con uh, peripheral mentioned this, that congestion is, uh, can really bring down, um, you know, the performance of these, uh, AI workloads.
Again, because a single flow can delay, can cause, um, cause that, and so that's why avoiding, um, that congestion because you have enough capacity is the key strategy, uh, for maximizing performance. Now, if you see here, like I I've just said up currently, and there, there was a discussion in the previous session as well, like, uh, if you see packet spring appears to be the, uh, you know, uh, current, um, best method of load balancing, and there is truth to it. So if you see smartnick based solutions where, you know, these flows hit the first leaf, um, and then you, the packets are sprayed across, um, all the ECMP paths, and then they are delivered out of order to the next.
So you need specialized nicks or super Nicks to today, uh, which can basically handle those out of order packets. And the other thing is, you know, the other approach is distributed schedule fabric where same thing flows, hit the first leaf, then they're sprayed over, just like you do it in a chassis based, uh, fabric solution. There's a credit grants based mechanism, and then packets are basically sprayed to load balance across spine.
But in this case, some of them chop it into cells, some don't. And they would reorder at the, uh, reorder these packets because again, spraying is, is, is, uh, you know, uh, will cause out of order, but they will reorder the packets before they ship it to the GPU nick. So in this case, they, of course, you need deep buffers, um, to enable that.
And if you compare, um, you know, the cost is, is, is, again, both of these solutions are expensive. The super nick ones. Now you need specialized nicks, which are more expensive to begin with, but are also more expensive in terms of they, they consume more power.
Um, and you know, already these gpu, uh, servers or racks are power constrained. Now you're adding almost 400 wat more power to each server, uh, using this. And on the second one, yes, of course, um, these leaves are de buffer and their, hence they're more expensive.
Um, uh, you know, uh, and also both of them are proprietary. So you are logged in, in a, in, in a, uh, on the left hand side, you will be using vendor, a proprietary nick to achieve this. And, uh, you know, uh, so you'll be logged into any expansion of your AI data center because, you know, because of that, and also in the case of the distributed schedule fabric, you are definitely logged into that, uh, the whole, uh, fabric ecosystem there from a single vendor.
So now let's look at, um, uh, peripheral, I mentioned this in the previous session, but I'm gonna go dive it a little bit deeper at how juniper's open and standards based approach, uh, using ethernet, uh, solves some of these problems. Um, and the umbrella, we are calling it AI load balancing. So the first technique is dynamic load balancing, right?
So this is an evolution or enhancement or static ECMP where we take more informed decision instead of just doing a hash and randomly pinning a flow on all of, uh, on one of the available paths, we track ev uh, almost at a very every microsecond granularity that, Hey, what is the link utilization? Or what is my link quality? Um, uh, at, at every leaf where the first, uh, uh, these, this traffic hits, right?
So as you can see, the most opportune point where you can take a good load balancing decision and avoid congestion is the first leaf, because it has the most available paths. So what we do is we track link quality, which is a function of how much traffic is going. So link utilization on every link, we track that and also see that, hey, is there any buffer pressure on any of those links?
So those two metrics are combined and we derive, um, link quality. And here I'm just showing for simplicity, three, three bands. Um, uh, but there are actually eight bands of link quality.
So we classify every link from a zero to seven, and when actually these flows come in real time, at that instance of time, at a microsecond granular granularity, we say, Hey, which is the best link based on this link quality available? And we start placing the, the, the flows there. Hey, Denise, um, who's making that decision?
The switches, the leafs of themselves? Yes. So the link, this is all in the a, um, in the ASIC logic.
So this is done real time and line rate yes. To, to follow up on that is to what Denise is asking there, where's the state being managed, right? I'm monitoring the flows, I'm, I'm holding onto that information.
Is that some sort of packet broker or docker container running on the leave or the spine? What's ha what's handling that or managing all that flow? Yeah, so this is none of that.
This has, because these 800, so this is like, let's say our current generation is 52, uh, sorry, um, uh, 64 by 800 gig. And when all of them blow up, you have no time to, you know, so this is all burned in the logic in the ASIC itself. So these decisions, um, along with our Juno software is actually a function of, of that.
So that's why it's able to, it's a non-blocking switch. It can, if all 800 gig 64 ports blow up, this can do all this at line rate for all of those. Um, and that's the only way, because, you know, otherwise you'll introduce delays and, and you can get in trouble.
So, um, so there are certain, um, so two primary methods on how you do this, right? So as I said, there are, these flows are long lived and very few of them. So to extract more entropy, the first one is flow let mode.
Right? Now, what the flow flow let mode, mode does is extracts more, more flows or sub flows in time, right? So if a flow pauses for a configurable amount of time and the next burst comes in, we can treat it as a new flow, even though it's like the fi same five topple, right?
Uh, in the, in the traditional ethernet. So that way we can break these, uh, flows into more, uh, sub flows and then start, uh, you know, fairly, uh, distributing them, uh, based on that instance of time, whatever has happening. And, um, this works very well for AI workloads because, you know, in AI workloads there is this natural tendency of compute and synchronize.
So all the GPUs that are participating, let's say in all reviews, they will come take their load, their batch of data, and as I said, they're lockstep. Once they're done with their competition, they'll synchronize. So there is a natural pause, and that fits very well in this.
And this doesn't introduce any out of order packets because you know, the previous burst of our previous flow lead, in this case, we, if we put it, placed it on a path, let's say to spine one, and the next bust, we after, uh, waiting, let's say 20 microseconds or 16 microseconds, we place it on a, on a better path. At that instance of time, the previous packets would have reached the destination. That's the assumption.
That's why you don't see, uh, you shouldn't see out of order packets. Now the second mode is packet mode, as the name suggests. You can see we make this decision on a per packet basis.
So this is the classical packet spring. So we will spray the packet, uh, across these, uh, ECMP paths. We will still factor link quality in it.
So if some link still gets degraded, um, you know, we will avoid that link even spraying. Um, but this is, um, this is packet Request. Quick question, sorry.
Um, because now you say every switch will make that decision, we have those two modes, but if that leave one switch will make a decision, I will go through as one or S two or S3, they need to be aware of each other, uh, of each other. Is there some kind of a stacking mechanism between those switches or not? A great question.
And that's the next technique that is GLB that actually takes a global view in this one. So I'm just gonna, uh, say, so this one is effective. Yes, you are right there.
These are based on local decisions, local link quality, and still you can still send some extra traffic to a spine and that the GLB will will solve problem. Okay, but this, well, We're still on the local Brian Martin here, signal 65, correct? Uh, for the flow lit mode, uh, those periods of inactivity, you mentioned 1620 microseconds.
Is that tuneable? Is that Yes. Something you're, so we can set that, Yeah.
So the lowest you can go is 16 microseconds, and then anything up to, uh, yeah, you can do that. Okay. Um, so yeah, so this basically solves, uh, if you see this picture, this solves this problem, right?
The local, you're gonna fairly distribute, um, uh, traffic over all the, the leaves, and you will not, you'll avoid getting overloading a single, uh, because you're now tracking each link, right? But the other problem is still there, and I'll come to, uh, your question, What does ECMP stand for? Equal cost multipath?
So, Okay, So because these are symmetric low architectures in data centers, so you, you know, a routing protocol will, uh, build all the paths that are equal cost, and the leaf will pick one of them based on a load balancing that you'll choose. Okay? And since we're right here, what is IXIA?
Oh, IXIA is a traffic generator, so Keysight, okay. Oh, So, um, yeah, so here what we did is we, uh, in the lab, uh, that we have here with, uh, with the, uh, NVIDIA GPUs, um, we saw, we put these three, uh, uh, techniques to, to, uh, test, right? So the first one is static ECMP, that's the hash base, the A-C-M-P-D-L-B flow lead mode, and the DLB per packet mode.
That's the packet spring. And we used a ml, common standardized benchmark, um, DLRM for this, uh, test. And what we found is, you know, the static ECMP deterior rates as, and, and, and, sorry, one more thing is we use the IC traffic generator to introduce congestion so that we can see the performance under congestion, right?
So as we increase the load on the traffic generator, started introducing more and more congestion, you see that the static ECMP starts to deteriorate. Uh, DLB was very deterministic, uh, highly deterministic, and it, it actually did pretty well as the, as the more congestion was introduced, uh, best was still packet spring, as you can see from that. 08%, right?
So the flow lead really did well without you needing any special nicks to handle or, you know, expensive nicks to handle this. So, uh, really there's not that much of a, uh, ROI on that, on that, right? And if you see pictorially, this is a graph that shows you what was the link, how did, how much traffic was going on, all the paths, um, between leaf and spines, right?
So as you can see, static ECMP, it was a random distribution. Some links were, were highly utilized, some were not. Whereas in DLB flow lead, you see all the links are fairly utilized.
Um, and that's, that explains the performance predictability. And the packet spray also looked exactly like it's indistinguishable. That's why I don't show you, but that's how it looked like.
Um, and in terms of variability, so the number of times, so this is a graph on, uh, we did three consecutive tests, and this is a nickel test, which is a standard Nvidia, uh, open source tool. And you can see the variability in performance of static ECMP, which is basically lottery based. You're pinning, you know, you're picking a path.
One time it got really lucky and had really peak performance, but the next time it was 1 96 and third time it was around 2 63, right? So, and this is in a controlled environment, whereas in, in real, uh, clusters where there multi-tenant traffic is hitting this variability can really hurt, right? Uh, whereas the DLB really reduced that variability band to a very narrow one, and you have, you know, 3 73 to 3 53, which is, uh, which is really good.
Um, now coming to the global load balancing. So the second problem is still based on local decisions, DLB or these leaves can still send some traffic where, you know, even though you have built an equal, uh, one is to one capacity because of these local decisions may skew some traffic or, you know, uh, can create imbalances by sending, uh, traffic to a spine. Um, and in this case, what we do is, so for the DLB, uh, the AC quality calculates all of its local link quality in GLB, what it does is it actually sends it one half of, uh, upstream, or in this case, the spines are going to send that link quality data that they're measuring locally to these leaves, right?
And now these leaves, because leaf is the first, leaf is the most effective point of, uh, you know, um, of, uh, uh, where you can take a good load balancing decision in clo the, at the top of the tree, you pretty much have, um, only one path, uh, down, right? So that, uh, now these leaves are armed with that information. So what they do in, in the, in the example that you brought up, so let's say GP two is going to talk to GPU six.
What the leaf does is first it reference its local link quality and says, okay, one of the links to spine one is already a lot of traffic is going there. I will avoid that, but I have two, two paths now, right, uh, to spine two and spine three. And if it was only DLB, it would have picked one of them, right?
And in this case, now it is armed with this information that spine two told it like, Hey, I, one of my links is congested. So in this case, the leaf one is going to say, okay, I'm not gonna send traffic that way because I have a clean path all through via spine three. So let me load balance this flow new flow onto that path, right?
And two things need to happen here. One is an ASIC to ASIC real time update that's updated. Uh, but one more thing we have done is we have an IETF draft, uh, which we have published on, you know, enhancing BGP, where now we are tracking next, next hop, right?
So these leaves, and that's done like in the control plane. It doesn't have to be in the, in the runtime. So each leaf knows, like for every destination leaf, which flows may take which path because let's say when the real time update comes from S two S two says, Hey, my, uh, my link to S two and to L three, that one in the red is congested, or is, is, is highly utilized, then it needs to determine which flows may take that link, right?
Because it'll, it still needs to send a, all the traffic this time to other 62 ports on that spine two, um, still to spine two, right? So that's how we just, uh, we just determined that by the BG P'S next, next hop tracking on the leaf, and then when the update comes, we just pull those, uh, routes out. So for the temporarily time being, the traffic will not be sent to that spine and when we get the next update, and if that link is fine, we just restore them and then, you know, load balancing is restored.
Yes. So, so I think back, uh, to the, to the earlier question about what's handling state, I think you just answered that, right? We're moving that into the BGP control plane and saying a BGP is gonna advertise, keep up with what's available, what's not available, uh, any potential performance hits from that perspective.
I mean, it sounds like that might be a lot of BGP updates unless we're looking at very similar traffic types in this cluster. Yeah. So BGP, uh, uh, we are relying on BGP to learn the paths.
And also we have done this enhancement to know, hey, not just next hop, I need to now know, hey, what next, next hop, so which links, so that's done maybe when even training has not started. So you build this and that's fine. In the runtime when traffic is running, there is a ASIC to ASIC GLB update.
When that comes in, I just refer to this table and say, Hey, spine two has reported me this link is down. What did BGP tell me? Which destination may take that link on that spine?
We just take that out and put that. So BGP really is not in the runtime, so there is no reliance on BGP update to this. There's a GLB update that's mapped to already built topology information that BGP was used for.
Gotcha. So this is like, as Perfu said, uh, like a Google Maps, right? Like, so it gives you end-to-end visibility, not just like, Hey, what does my, uh, freeway look like right now, but what's the next connecting freeways?
So Google maps is a good analogy there, uh, because the color can go from blue to yellow to red to black, uh, I see low and high there. What's the granularity on those paths? Great question.
Yeah. So again, so here just to explain, I had three, three, um, three levels. It's actually eight levels.
Yeah. So eight bands of quality are there, uh, zero to seven, and that's what is used for DLB decisions as well as, um, G lbs, DLB. Thank you.
So now then, uh, you know, so these two, uh, are quite effective at actually solving a lot of congestion. And as you can see, they're re uh, dynamic because they're reacting to, um, wherever congestion is created in spite of, you know, you fairly distributing the flows. And then we are juniper, we try to challenge ourselves like saying, Hey, is there, is there a simple and better way to actually avoid congestion?
Right? So this RDMA workload balancing, um, I will, um, I'll just take a moment to describe it Now, whenever, um, the, uh, these, uh, the training is occurring at the fundamental level, whenever the GPU needs to write, um, something or synchronize its gradients or whatever it's doing, or RDMA to another GPU's memory at the, uh, fundamental level or lower level, what happens is, so these are orchestrated by a middleware called nickel, uh, which is communications library from nvidia, and it sets up like these, uh, I mean the analogy could be like these sockets in the TCP world, right? But these are, uh, created by a middleware and then reused over time.
So what happens is there is a memory region associated on, let's say the left side is the center, GPU, and then there is a receiver GPU. So if it has to write two gigabytes of, uh, data over this RDMA and ethernet, and over the network to this, at the RDMA LE level, there is a Q pair, which is map to a memory region on both ends. And then this Q pair is, um, then used, uh, over a UDP and IP and ethernet and send over the, the network, right?
Yeah. Okay. Just, uh, for listeners, our DMA, what does it stand for?
Uh, remote direct member access, access, remote direct, direct memory access. So DMA is like whenever you, you know, whenever your CPU fetches from your local laptops memory or ram, uh, you know, and then remote is like when it can fetch or do operations or, or remotely on, on another, uh, on another CP or G GPU's, uh, ram. Thank you.
So that's what it is. So remote direct memory access, thank you. Yep.
And so, um, so, so this is like, you know, so you have a giant flow of 400 gig if you use a single Q pair on both ends. So one of the techniques, um, is like, Hey, can I extract more subflows so that, you know, I can spray them, I can create more subflows so that the fabric can usually load, balance them across all the available paths. So what, uh, this does is, hey, can we do four Q pairs?
I mean, Q pair is like, you know, on both ends. In this case, what, what the nick that's anchoring the RDMA does is like, okay, I'm going, you have to write that same two gigabyte. I'm gonna break it into 500 megabytes, four regions on both ends, use four QPS and transmit like a hundred gig four flows, right?
And, uh, this is not new. This is when this is enabled. Usually what happens is each q pair uses a random UDP unique source port so that the switches, when this traffic hits the switch, it can hash, it can catch that, okay, there are four flows now and it can then distribute these flows, right?
Correct. And because it's based on random, so it is like the, you know, you have to hash and there is no determinism, right? What we, what we did is very simple idea.
We added determinism, but we said, Hey, can we, instead of using a random port, uh, which has been done for, for 10 years, can we assign an IP address to each floor? And then if we do that, the beauty of that is we create this to a, instead of random and hashing and all those things, a deterministic routing problem, and you can now do a nice traffic engineering and pretty much carve out a lane for it throughout the fabric because the capacity exists. It's just that the load balancing, which usually hashes on, on a source port, uh, takes this decision.
So, so in essence, what you're doing there is you're, you're taking away, uh, using just a random number to try to load balance correct on instead doing the work to actually map all those Flows. Exactly. Right.
And still using routing protocols, ethernet's distributed routing protocols of your choice, but BGP preferably, uh, and what, what that does and uh, is, so these same four flows, think of it as like, you know, so now you have a deterministic flow with its own source and destination ip, and we have a pre card power path. So let's say you have a, the first, uh, flow is from subnet A, the second is B, C, and D kind of thing. And then you can now use traffic engineering to pre card, uh, this path, like saying that, hey, all the first flows, now deterministically can take the first spine or the, a preferred path.
So it's, you're pretty much, um, taking the randomness away. You say, Hey, I, uh, because the capacity exists, you have build, you spend money, and having this one is to one hour subscription. It's just the load balancing was not, may get you in tricky situations.
At times it'll work most of the time. In this case it is predictable, deterministic, uh, thing. So if all the flows blow up, we disaggregate these flows into so that there is an important concept.
N So we break these flows and as many paths you have from the leaf, so in this case you have four paths through four spines, we'll break these sub flows into four, right? And, uh, if all the flows blow up, in this case, let's say there are eight nicks, and you know, so each one will be a hundred gig flow now, and they will, because there are eight GPUs, you will have eight by a hundred gig, which is that link is eight by a hundred gig. So there won't be congestion, there won't be any out of order packets because they're not gonna take any other path.
It's a, it's a reserved lane, uh, for each subflow, right? Very simple idea. Um, but very, very effective.
Right? So, Uh, sorry, question. Um, mm-hmm.
Now you are assuming that, uh, all those GPUs require the full bandwidth. Yes. Is that always the case?
Uh, it's not, but you have to build for that case, and they are built for this. So you have, for example, like, let me show you the next picture, right? Um, usually these are like, um, six, um, 64 by 800 gig, um, uh, switch.
So what we do is one is to one, um, what you do is you, because these nicks are 400, so you split the eight, uh, uh, 800 gig, um, uh, into two by 400. You attach 64 GPUs, and then you have to budget that in case all of the traffic comes to the leaf. You will have 32 by 800 gigs.
So that's what I mean. One is to one more subscription, meaning whatever, uh, ports you reserve for the GPU facing or the host facing, you have to have in down exact capacity towards the spine to handle if all GPUs are sending full line rate. So in this case, what we do is what, as I show you, there are, if there are, is a scaled, uh, 4,000 GPU uh, topology.
If you have 32 paths in this case, 'cause 32 by 800 gig is equal to 64 by 800 here, you divide these, uh, you break your end becomes 32, and you will do 32 Q pairs. So that if all of them, all the, the GPS here, uh, uh, start sending traffic, it'll fit nicely in a pre deterministic fashion. And here you eliminate, uh, load balancing.
This is pure routing, uh, because this fabric is built on, you know, BGP and when the flows come up, we just, uh, look up and there is a preferred path for each of the, uh, you know, colors. And that's how this works. Now, the way routing works is like for each of these colors shown here, we will have a higher preferred route, um, to a spine, and the second color will have the blue spine and so on.
But we will still advertise a lower preferred route through the backup for just in case that link fails, right? Because failures are a reality. So this is how we will still advertise ECMP paths, lower preferred or backup PA paths, but a, a preferred high, high path.
So when, when everything is in steady state, you have that determinism and let's say, let's say a link or a switch fails. In that case, what happens is for this, uh, for this flow, which was preferred path, was that spine, all the backup paths are activated and it'll do dynamically DLB and GLB across all the other paths, right? For the time being.
And as soon as that link gets restored or the switch gets restored, it'll immediately sta snap back to, um, the deterministic for forwarding. So I just noticed you switched from earlier in the slide deck you were using rail or rail optimized network topology, correct? And with these optimizations you've dropped back to a standard leaf and spine, Okay?
Hold your thought. Yes. Okay, great point though.
You, that means you, uh, I am, uh, able to explain this well, okay. You know, that's a, that's a great point. Yeah, I have it covered.
Okay, so I'll go over it. Um, so now if you see the variability, right? Like I flashed this like, so again, uh, quickly rehash, ECMP gets lucky sometimes, or, uh, mm-hmm.
And this is like in a controller environment, three consecutive runs DLB improved, it considerably reduced that, uh, unpredictability to a narrow band. But look at this. So, uh, this RDMA aware load balancing, what we call RLB, is consistently hitting the 3 73 highest performance mark all the time.
'cause as you see, there's the, in the design, there is nothing in steady state. Um, you know, it, it always performs no conflict. There is no, um, load balancing, uh, introduced, um, randomness anymore, right?
So that is a, so it'll keep the network running at peak performance, uh, by, by this design all the time, which is very desirable. And again, here I'm just showing, um, you know, uh, the DLB and R lbs some other metrics that we measured. So as you can see in the DL BS case, yes, uh, all the links.
So again, these are all the, the link, uh, the traffic on all the links between leaf and spine. And as you can see, and we, we bombarded with all our, uh, GPUs, the traffic using nickel test. And what we saw is with the, um, with the DLB, yes, the traffic is fairly distributed over a cross links, but still at some point it starts, um, causing micro, uh, congestions, right?
Uh, but in this case, there is no congestion because, you know, there is a reserve path and it's, it's just a straight peak and a flat line up there. Uh, during multiple runs, we never saw any congestion, uh, metrics. So ECN and PFCs, uh, uh, and D-C-Q-C-N is, uh, primarily, uh, congestion avoidance, uh, mechanism in this, uh, ethernet.
And we didn't see any of those triggered, and no, out of order packets because of course, uh, it's designed to follow a single path, uh, through the fabric anyways. Uh, whereas in the case of, uh, uh, the other one, and this is for rail optimized, and I'm gonna come to the non rail, um, now this is with the point you were trying to say. And, uh, so in 2018, 19, when NVIDIA started recommending these rail optimized designs, it was like, hey, uh, traditional ether, uh, load balancing techniques are not, not enough.
Mm-hmm. So, um, you know, we will connect them in such a way that, you know, you have rail optimized, uh, designs. And that is true because, you know, it takes off some of the, the load balancing, um, requirement from the leaves.
Then we said, Hey, if we are achieving this kind of a performance in rail optimize, is this good enough to now go back to top of rack switch where, you know, you have all these connections, uh, in a, in a, uh, top of rack switch that allows you to use now copper, right? Right. DAK cables, um, and d cables are, uh, you know, way cheaper, uh, less expensive.
They consume very low power, uh, they're more reliable, um, compared to optics. And you know, that traditional sense was, hey, performance, you will not be able to get, because you need to connect rail, so you have to run the longer cable. So we said, okay, let's put this to test.
Right? And what we found is almost this brings not even non rail, uh, which is basically the top of the rack. It brings, brought it to almost, uh, peak performance.
Um, because, you know, we saw, I mean, uh, this is, this is actual lab results from, from multiple runs. So really, um, and this is again, the, uh, similar graph. There was still no congestion, uh, no out of orders delivered, uh, for, for the non rail.
Um, whereas DLB was still had it in, uh, that explains it, right? So, so really, um, if you, I summarize the benefit of, uh, this is a new, very simple idea, uses all the components of what Ethernet's routing protocols and you know, just by, um, assigning determinism to these flows and converting it into a routing problem, you have a consistent high performance, like predictable in steady state, uh, you have in order delivery because in steady state, they, you know, they, they will all follow the, the consecutive packets will follow the same path. And this is done without the need for any expensive hardware like, um, you know, extra nicks or super nicks or, you know, deep buffer switches and things like that.
And this also delivers non-real performance, which is highly desirable because of, uh, you know, um, the property. So, But for those, for those of you that don't know the difference between rail and non rail, uh, in a rail environment, any communication from a GPU on a different rail has to travel through the PCI connection to the other GPU and then across the network. So one extra step real big.
Yeah. So actually that's, uh, that is it. Yeah.
Sure. So any questions? One More?
Yeah, sure. Go ahead, please. Uh, when you look at scaling up, when we get into very large clusters, um, do you see the le the, the regular costs, um, scaling up that high, what do the topology start to look like or the tradeoffs start to look like between rail and cloth?
So in the rail one, uh, rail topology again, uh, and that'll apply to non rail as well as you go like through the five stage. Usually we start seeing, um, over subscription. Um, seven is 2, 1, 5 is two, one that kind of, because the traffic that's traversing to across that is, is becomes lower.
And that's because also because, um, you know, a lot of these orchestration orchestrators have a locality, uh, property where they will try like, Hey, you need GPUs of a hundred GPUs, or, you know, a thousand GPUs. They will try to orchestrate the workload in a local area where, you know, you don't have to traverse that high. You know, they're all in a, in a cloud topology where, you know, you don't have to take more hops.
So that's why you can go over subscription and because, um, in the case of RLB, you, um, that over subscription, we are using, uh, you know, BGP communities and stuff. So you can still architect and factor any or subscription on that layer mm-hmm. Where you'll start assigning multiple spines, um, you know, uh, to handle the same color.
Got it.