Cisco AI Cluster Design, Automation, and Visibility
Cisco’s presentation on AI Cluster Design, Automation, and Visibility, led by Meghan Kachhi and Richard Licon, aims to simplify AI infrastructure and address the challenges of lengthy design and troubleshooting cycles for GPU clusters. The core focus is on enhancing cluster designs, automating deployments, and providing end-to-end visibility to protect a competitive edge. The session outlines Cisco’s reference architectures, key components for building AI clusters, and upcoming updates to its Nexus Dashboard platform, which is expected to streamline design, automation, and monitoring at scale. This comprehensive approach is crucial because the battle for AI success lies at the infrastructure layer, ensuring GPUs are not underutilized by network inefficiencies.
Cisco leverages three unique pillars in its AI networking strategy. Firstly, its systems feature custom Silicon One platforms, offering programmable pipelines that quickly adapt to evolving AI infrastructure demands, and a partnership with NVIDIA that provides NX-OS on NVIDIA Spectrum X silicon for full-stack reference architecture compliance. Rigorously tested transceivers and a mature NX-OS software, now optimized for AI workloads, complete the system offerings. Secondly, the operating model includes the Nexus Dashboard for on-premises management and Nexus Hyperfabric for a full-stack, cloud-managed solution, complemented by an API-first approach to seamless integration with existing customer automation frameworks. Thirdly, extensive AI reference architectures serve as validated blueprints, spanning enterprise-scale deployments (under 1024 GPUs) to hyperscale cloud environments (1K-16K+ GPUs), providing detailed component lists and ensuring a consistent networking experience across vendors such as NVIDIA, AMD, and storage solutions. An AI cluster is broadly defined to encompass front-end, storage, and backend GPU-to-GPU networks, with a growing trend toward convergence enabled by high-speed Ethernet to unify operating models.
Designing an efficient AI backend network requires a non-blocking architecture that maintains a 1:1 subscription ratio, keeping every GPU within one hop of others for optimal communication. Cisco employs a “scalable unit” concept, enabling incremental expansion by repeating validated blocks while adjusting spine-layer connectivity to maintain high performance. For smaller-scale deployments, such as a 32-GPU university cluster, Cisco demonstrates how front-end, storage, and backend networks can be converged onto fewer, high-density switches, simplifying infrastructure. A critical consideration for such converged environments is Cisco’s policy-based load balancing, an innovation leveraging Silicon One ASICs. This enables preferential treatment of critical traffic, such as GPU-to-GPU training, over storage or front-end traffic, ensuring AI jobs run with minimal latency and maximum GPU utilization, even when sharing network resources.
Presented by Meghan Kachhi, Technical Marketing Engineering Technical Leader, Cisco and Richard Licon, Principal Technical Marketing Engineer, Cisco. Recorded live at AI Infrastructure Field Day in Santa Clara on January 28th, 2026. Watch the entire presentation at https://techfieldday.com/appearance/cisco-data-center-networking-presents-at-ai-infrastructure-field-day/ or visit https://techfieldday.com/event/aiifd4/ or https://www.cisco.com/ for more information.
Transcript
My name is Megan Kuchi. I'm part of the AI Infrastructure Networking Group at Cisco as a technical marketing engineer, and today virtually joining with me is Richard Ko who leads this entire team. And let's face it, as much as we are obsessed with different models as engineers, we know that the real battle is at the infrastructure layer.
And that is what exactly we are going to talk about in terms of how Cisco is trying to simplify from an infrastructure perspective. The cluster designs, the automation, the end-to-end visibility so that you're not finding needle in a haystack sort of a problem when your GPU take say months to design weeks to troubleshoot, we are not losing that time only. We are losing the competitive edge.
So that's what we are trying to emphasize on for this session. And before I get started, I just want to make a quick call out to previous session on networking field day 39 where we went in detail and Paresh did a fantastic job of covering all the AI cluster design components, the challenges from bits and bytes level, all the way to congestion control mechanism with dynamic load balancing and cabling plans and everything. You can refer to this for more of a deep dial, but for today what we are trying to go ahead and cover is more of the reference architectures.
How do we build this AI cluster and what are the key components tied with this AI cluster? And at the same time we'll try to look at a demo of what new updates we have made with our Nexus dashboard platforms and how we can design automate right at scale and provide that end-to-end visibility from this dashboard. After that, I will try to hand it over to the hyper fabric team.
We'll try to focus on more of the cloud managed piece for it. So with that open to any of the questions because we want to keep it as interactive as possible. And let's get started with Cisco AI networking.
Now being part of this technical marketing engineering team, we get to interact with lots and lots of customers across different segments. And after all these AI infrastructure boom, we help key customers who are hyperscalers, who are scaling from tens to thousands to millions of GPUs which are out there. They have separate needs, separate, uh, you can say expertise and everything which is out there.
When we talk with new cloud vendors, what the GPU as a service vendors, they are building thousands and thousands of scale in typical multi-vendor environment fashion. And then we are seeing this uptick in terms of enterprises that we are targeting and talking more and more in terms of universities, automobile manufacturing plants. And I will give you the ex actual examples of the customer deployments where we are designing.
But those designs might be just tens or thousands of GPU that they are looking at. And how does Cisco AI networking fit into it? Number one thing that we have as part of our three unique pillars is the systems.
We build our own custom silicon, silicon one platforms if you're familiar with it. And sometimes we get asked the question why not rely on merchant silicon? But the key differentiation that we have with this is that with ai, the infrastructure requirements are all the time changing rapidly.
And we are working in lockstep for, with this large scale customers, to make sure that we go ahead, leverage our programmable pipelines and make sure the features that they're looking for are available without respin of asics and then go ahead and support, start supporting it on our platforms. At the same time, we have this partnership with Nvidia where if customers are looking for that full stack reference architecture compliance, you can have the same Cisco switching platforms running the Cisco NX OS as the customers are already having NX OS deployed in their uh, data center fabrics, but have the Nvidia spectrum X silicone running on top of it. And we will talk about that as well in detail.
We have our own transceivers, which are rigorously tested. The key benefit that you get is if you're connecting with NVIDIA Connect X seven eight N or a MD Polar, we have these transceivers validated as well as checked at the unit level testing, uh, stress testing, entire performance testing, everything is done for these transceivers. And we have tried our own NXO software, which is across the data center mature platform now being efficient for AI workloads.
The second important pillar that we have is the operating model. We have Nexus dashboard, which is for on premises management and for full stack cloud managed solution that you want, you would have Nexus hyper fabric. At the same time, when we talk to new cloud customers hyperscalers, they have that expertise where they do not want to rely on these management vendor specific management dashboards.
They want to integrate with their own APIs. And that's why even we take this API first approach to have those APIs included as part of their automation framework. And third important thing is that you would see lots and lots of AI reference architectures, which are, which are validated design architectures with performance benchmarking teams along with all the different vendors that we try to support, which is Nvidia, A MD, Intel, all the storage vendors like TD and CCA and so on.
So the key thing is that we want to make sure that we have consistent networking experience irrespective of what vendor you are interacting with. So this is exactly where we are with our AI networking story and we are constantly expanding and enhancing on top of it. So for the first piece, what I will try to do is focus on the systems and reference architecture to understand that what we are trying to solve over here.
Now, first of all, we need to understand the definition of AI cluster because sometimes whenever we talk with customers as well, AI cluster is nothing but GPU to GPU backend communication network, which is east west traffic. Mm-hmm. Well, that's not true when you are trying to design the network.
Actually it's more than that. So say for example, if you have your GPU nodes which are out there, the first and foremost thing that you would require is to make your GPUs accessible to users, applications, services, schedulers. And that's where frontend network would come into picture.
The traffic pattern over here would not be that much burst. It may be moderate in comparison to east, west GP communications, but at the same time, your inference traffic runs on this front end. So it'll be bursty in that capacity, but at the same time, latency sensitive on top of it, you would have storage network, which can be high speed storage leveraging say vast cca, DDN and so on.
And historically these two have been separate networks where due to different QS profiles, latency profiles, this have been running separately. But more and more we are deploying an actual customer, uh, scenarios. We are seeing that they're converging both frontend and storage because of more and more 400 gig, 800 gig proliferation that we are seeing.
And we are seeing that front end traffic is anyways limited or less compared to the storage traffic that you would have from a high bandwidth perspective. Megan? Mm-hmm.
Uh, Jack Poller with Paradigm Technica. So you earlier mentioned you're talking about the, the type of traffic for inferencing. Do you make a distinction in the architecture between a cluster targeted at inferencing and a cluster targeted at training?
Uh, yes. We do have different, uh, load balancing mechanisms and so on. Mm-hmm.
So again, that was covered during the networking field day 39 session where we have different modes in terms of load balancing the traffic, right. And based on the policies and so on. And I'll try to cover during these slides also that how we are trying to have that, Well, my, my question is more directed at in this discussion, are we discussing something targeted specifically at inferencing or is this more generic?
Well, This will be more generic based on the customer deployments. It can be either training or inferencing and so on. So it's not going to be just inferencing focus.
Thank you. Yep. Alright.
So Meg, mm-hmm. Does, does uh, Cisco get involved in the GPU to GPU east west traffic kind of thing? Absolutely.
That's the next thing. So that's where, uh, based on your designs, say if you're scaling out from one node to two nodes to all the way up to say whatever end number of nodes that you have, that's where we have this dedicated backend network. And this is where all the key buzzwords that you might hear about, like lossless, bursty, traffic, long live flows, elephant flows, lowest latency, lowest jitter, highest job completion time, all those things apply over here because you are doing those collective communication jobs like all to all, all gather, all reviews, which is happening in the network.
And that's where also Cisco has all the solutions for different challenges that you might have in terms of congestion and everything. So this is very sensitive in terms of even minor congestion network delays or anything that might happen in the network because you want to make sure that your GPOs are being utilized at a hundred percent or high capacity so that you're not wasting your infra investments that you have in your network. So I'm not exactly familiar with the technology, but uh, Nvidia has their own NV link and things of that nature are, are you playing in that space?
Yeah, so Right now NV link and all those technologies are specific within the servers where you can go ahead and scale up your designs, add more GPU nodes and so on for the internal communication that you have. So that is specific to NVIDIA and compute servers. We are not playing in that scenario.
What we are trying to right now, go ahead and how as of today is playing the scale out when you want to distribute the nodes across uh, different uh, say GPUs and so on. So like rack to rack versus server To server Server, yeah. Within rack.
Yeah, that's called like a scale up design and then we have the scale out and I we are also playing now and scale across meaning that YouTube power constraints, cooling constraints and so on, data centers like new cloud vendors have data center facility in one location and the other location is also there. We are connecting those via our new asics, which are called P 200 A six and then having it across like hundreds of kilometer distance a cluster which can be a megawatt megawatt cluster together, it can form like a gigawatt plus. Yeah.
So Frederick Van Herrin, uh, high fence consulting. So when I think training I think in finna band. So are you, is Cisco then Yes.
Injecting ethernet to replace in finna band? Correct. So this is mainly ethernet based solution that we are trying to showcase.
And the key benefit that you might get out of this is that if you have separate networks, you are breaking up the operating model, right? And I will try to show you in the demo as well that how when you go with an ethernet route, you can manage the entire cluster as one, add configurations, profile visibility as one as well. So that's coming up in the demo as well.
The benefit you get and then you have so many different components you need to manage your switches, you need to manage your compute nodes, you need to do firmware upgrades, you need to go ahead and have the schedulers and everything pushing down the configurations. This is where out-of-band management comes into picture. But that's what the entire AI cluster constitutes of.
And this is one of the misconception people just focus on the east west backend network and think that this is what is your AI cluster. That's not true. Backend network East west can be for your distributed training and inference perspective, but you need to make these GPUs accessible to your users application scheduling logic as well as the storage network which is out there.
And the key reason I discussed this AI cluster is because we have dedicated reference architectures which so as a validated blueprint so that you can eliminate all the, so you, you can think about guesswork and get a starting point in how your AI cluster would look like and following the NVIDIA's enterprise reference architecture principles, we have our own Cisco enterprise reference architecture for less than 1,024 GPUs. Where based on your operating model, if you ma want to manage it via on-prem Nexus dashboard or via hyper fabric AI in a full stack form, you can get examples of like 96 GPU cluster, 1 28 GPU cluster, 1,024 GPU cluster along with each and every components. How many optics do I need?
Uh, what switches would be, right? Not just for my backend but even for the front end storage, all those infrastructure which is out there. So that is the power of this enterprise reference architecture where it sort of gives not just a network diagram but validated certified components which are surely going to interrupt it with each other.
So question about that is your reference architecture, Ken now I'm sorry and then introduce myself. Is your reference architecture meant to be implemented as part of a broader reference architecture from Nvidia say, or compete with it as customers are looking how to, you know, deploy in their environment? So I I, it'll come up to the NVIDIA versus Cisco reference architecture in the next slide.
Okay. So just wait for that. And this is where again, this is for enterprise reference architecture.
At the same time we talk with large scale cloud service providers, the new cloud vendors who are building in capacities of one k, 2K, 4K, eight K, 16 KGPU clusters and that's where we help cloud reference architecture. But this is where we are not only showcasing the backend network, which might be easier to understand, but even the end storage converged network which would come into picture with the Nvidia, uh, say HGXH 200 platforms, specx and so on. The examples are there where say there might be thousands of optics, but what optics to use, do I need to have O-S-F-P-D-R eight?
Do I need to have cable lens of five meter a hundred meter? It gives you entire description. So it's not just a network diagram, it gives you different example.
If I want to build one KGP cluster, how do I do that? If I want to build 4K GPU cluster, how do I do that? What would be the components and everything required?
So coming to your question, right, are we competing with them? No, we are working together and making sure that for the right set of customers we have the right set of platforms. So see if you want to add to the Cisco reference architecture, have that unified operating model with silicon one based asics, which are our custom silicon, asics have your front end backend storage network, everything connected via these switches.
We have the N 9,364 cross 800 gigs switches as of today. And there can be requirements from customers where they want to make sure that they're compliant with NVIDIA's reference architecture in a full stack manner where you might have Nvidia, Nicks, Nvidia storage e everything which would be compliant with that. And this is where we have, because of this partnership apart from Nvidia Spectrum X switches, we are the only vendor to have N 9,100 series switches with Nvidia spectrum X switch silicon so that you can go ahead and have N-C-P-R-A compliance but at the same time NXOS becomes your common operating point because your expertise, like we talk with enterprises, we talk with new cloud customers, they are already having Cisco operating systems as part of their data center fabric.
They can leverage NVIDIA's silicon, have the Cisco designed with NXOS running on top of it and at the same time manage it via the same dashboard, Cisco dashboard which they're using to manage their current data center fabrics or even their uh, you can say AI clusters that they're planning with. So we are not competing, we are partnering more of more together and making sure that we have right fit of platforms based on customer's requirement. If they want to go full NV route, we have N 9,100.
If they want to go with Cisco route, have Nvidia like GPUs Nix and so on with Specx capability you have our N 9,300 series. So is that the distinct spray ese from Silver Tank Consulting? Is that the distinction between the hyper fabric versus the non-hyper fabric?
Is that because you're using the Cisco N three and And hyper fabric, uh, is applicable for both reference architecture? So even uh, the CERA that we showed the uh, I understand that, but I mean is a is a, is a switches themselves the Cisco N 9,100 versus N 9,300 different between the two fabrics. So right as of today, yes you would have a hyper fabric 6,000 series, but there are plans moving forward where hyper fabric will be supporting the nine nexus 9,000 series switches as Okay stand back hyper fabric.
Okay. Stand so the camera gets uh, with them. So, uh, we'll talk a little bit more about this in the second half of the session.
Um, the operating model is really what's different between hyper fabric but assume that the hardware will be able to uh, use b, we use any of those sets of hardware as we build out for both, but we'll dive more into the operating model and sort of why we built it later in the session. So I guess, I guess the other question is are you limited to using the the 9,100 series or 9,300 series? I mean can you intermix the two?
Absolutely you're gonna have your front end network within ninety three hundred and ninety one hundred, but at the same time, if you want like a full stack compliance in terms of architecture, you will go with N 9,100 because that's why the customers want that compliance and that's why they have. But it's running the NXOS operating system so you will be able to have this common experience irrespective of what switch you're going through. I understand the operational experience would be the same.
I guess the question I have is like for the backend network Yeah you would be required to use the N 9,100 series in order to support that. Yeah, N 9,100 or N 93. So there is one.
Yeah, either, either one you can, you can support that. So thank you. Yep.
Alright, so I will go and give you an example of how do we design a backend network in an AI cluster and once we get these principles right, it's pretty much straightforward where say example number one thing you want to make sure in an AI backend network is that you have a non-blocking architecture. What does that mean? Say for example we are taking 64 800 gig port switch that we have, we cut it exactly in half, not literally, but just in terms of port connections.
Uh, we make sure that based on the GPU Ns that you're connecting to say if you are having connect X eight NS with support one 800 gig OSFP port, you can connect up to 32 of those connect X eight nicks. And then because you have 32 going to your GPU Nicks, you make sure in terms of uplink you have 32 ports going towards your spine. This is what a non-blocking architecture means.
Or at the same time this switches, if you're connecting to connect X seven nicks with 400 gig ports, you can make sure you can use the optics, right optics and the right cables to break it down into 1 28 6, uh, 400 gig ports, which means 64 go down to your GU mix, 64 go up and we have the highest redx redx meaning that we can even break it down further to a hundred gig ports and have 512 ports that can be powered up from this single switch that we have. So that's how typical designs are based on the customer requirements. But when we extend it out to say scale out, say if you want to go from one node to other node and so on to have your distributed training or inferencing, one of the critical or foundational design choices that we see is how you connect your GPU nick to the LEAF switches which are there.
Mm-hmm And how many of you are familiar with rails optimized design? Alright, I see two of them. So what does that mean is that when you connect your GPU NIC one you need to make sure that it connects to LEAF one, your GPU NIC two connects to LEAF two and so and so forth.
If you are eight GPU NICs in this particular server, so I have eight leafs which are out there, the key benefit that I get is when I try to go ahead and extend this, I make sure each and every GPU is just one hop away from each other. So that say if GPU one of node one wants to talk with GPU one of node 64, they are just one hop away from each other. And this is how we can power up to say with 400 gig CX seven nicks that we talk about 64 400 gig ports down.
This is called rails only design. You can get away with it, but what happens when say Nick one to Leaf one connection fails, what happens when Nick one fails all together, that's when the traffic has to travels through an alternate path. And even in terms of scalability, you require that spine layer for that non-blocking architecture and how that one is to one subscription ratio.
So the animations are a little bit slower. So hopefully I'm not talking too fast, but this is where you would go ahead and have your spine layer coming up. I uh, spine layer coming up in one-to-one or subscription fashion.
Yeah. Finally it's there. So it takes 30 seconds to go ahead and plug cable them and now in 30 seconds we have the ports also up and running.
So, so that's how we have this non-blocking architecture tied together and you make sure that you have 64 ports from Leaf one going down, you have four spines, so 16 400 gig ports going towards your spine as well. So Frederick and Hern from ENS Consulting. So, so I understand the design, so if I compare this with Infiniti bands, there's not just the hardware component, it's also the logical component like you UF FM has, you know, for Infiniti band.
Do you have similar No, uh, no we do not require that. We have like right now this particular like fabric which can be configured with ethernet and we can, which we can scale out accordingly leveraging all the ports. So is it fair to say then that your existing ethernet tools are still working?
Yeah, Yeah. Those are working the technologies that we use configurations and everything, it's beyond the scope of presentation. We covered it in networking field day 39, right?
But we have gone through it and you can use exact same configurations and everything that you have in terms of networking com component, uh, concepts and this becomes your sort of a scalable unit. That's the term that we use, which means that this is a repeatable block and you can just go ahead and keep on adding more and more components, uh, or like, uh, like repeating this modular block and you will be able to scale out further. So for example, we are taking this uh, 64, 400 gig port scaleable unit.
The size can vary based on your designs and everything. This is not a hard stop or fixed solution. So that if I want to go from this 512 GPU 2024 GPU cluster, all I need to do is repeat that exact same block and incrementally upgrade it into scalable unit two.
But just to maintain that one to one or subscription ratio, I need to make sure that the links earlier, I had 16 links of 400 gig going through four spines. Now I have eight spines, so I have eight, 400 gig links. But you can look at how many cables, optics, and everything is going through this, right?
Because of this, the reference architecture can give you that cluster bomb analysis where number of optics cables, everything can be under By by one-to-one oversubscription. You mean it's not oversubscribed? No, it's not.
It's just the term that we, I also don't like it and whenever I use it I use one-to-one subscription ratio, but that's what the industry term is like one is to one or subscription because from data center's perspective, we are always thinking about five to one, seven is to one or subscription ratio. Okay, well, or or it's one-to-one under subscription. Yeah, exactly.
Exactly. And that's why when I try to go ahead and build like a smaller clusters in front of customer, and I'll show you the example that what we do is, I always call it one is to one subscription ratio that we need to maintain. So you are absolutely right, but just using the industry from that we are, This is uh, Arian Newsome, uh, I'm from Ethical Tech Matters.
I actually have a question about, uh, the GP node failures. So does that hyperscale fabric design prevent like casing failures? We used to see the old traditional three tier architectures.
Uh, sorry, the question was, Uh, so does hyperscale fabric, does that help prevent some of the cascading failures? So, so we can identify those failures and we will go into demo with Nexus dashboard and hyper fabric as well, how we can go ahead and look at those failures and try to remediate that which are out there. So for sure we are coming to that in a couple of minutes.
Okay, thank you. But we get all the time when we talk with say not new cloud vendors or hyperscalers that hey, we do not have this much of power. Uh, we do not have these many GPUs, we don't want to build at this scale.
And I just want to share one experience that I had with one of the university customers in East coast who got National Science Foundation funding for a couple of million dollars and they wanted to start small. All they had is the budget to have four nodes with eight GPUs each. They want wanted to start small with 32 GPUs and 32 nicks.
And based on all the research and everything for their PhD and every, uh, what what they had, they wanted to make sure once they get funding for all the universities in that area, they wanted to share this infrastructure. So some of you might be familiar with that, but we were working closely with that. And for them they were saying that do I need a separate backend?
Do I need a separate front end? Do I need to build at like eight, leave four spines sort of a scale? The answer is no, because you can just go ahead in this particular design and have two switches just to limit the failure domain to one switch.
You would have two switches take those principles for a high end like high AI cluster scale use, say four ports going from each of the servers to switch one and four ports from G Pix going to switch to, and by the way, the over subscriptions piece I always use when I try to design maintaining one is to one subscription ratio is what we try to go ahead and do over here so that if we have four, four ports going, each means that inter switch connections need to have four ports between them. So similarly for this like say four servers that we have, we will have 16, 16 ports for the entire block going to the switch and so on. And then I can just merge my frontend storage management network because anyways, these switches have 1 28 400 gig ports, right?
We have plenty of capacity in terms of port density that we have and we can just converge your frontend backend management and even your uh, the back uh, storage network altogether. So I'm gonna ask kind of a maybe obvious question. Mm-hmm.
So we're talking about cluster technology and you just mentioned an example of somebody that got funding and blah blah. Yeah. So what, what network are we looking at honestly from, from what stage of the AI lifecycle?
Is this helping them clean their data? Is this doing the inferencing? Is this doing the training or like what does, what traffic flows in this super?
So this is your converged network. So for this in university customer, they are mainly doing research work, more of computational analysis and all those things which can be tied to an AI job, which is distributed training or a distributed training for HPC, like high performance computing that they're using. So this can be mainly from a research perspective that they're using.
It's not necessarily inferencing that they're trying to do in most of the cases over here, but for sure they can be looking at once they train the data, how quickly they're able to access the train data with the storage network and so on. So this is more of a converged environment that we are talking about. So it's, it's When you say converge, you mean the environment from a technical standpoint or do you mean it from some usage standpoint?
From a technical as well as usage standpoint. So that's what we are trying to say. So that next time when they get like more funding or adding more GPU servers based on the same principles, say if we have term this as scalable unit one, they can keep on adding based on we can ports more and more servers which are out there and then scale out to more GPUs while Rev leveraging the same networking infrastructure.
So, and they can, can they leverage the same networking infrastructure? So as this one set is trained and you have whoever accessing the results to train data, then you can also have use the same network to do those other two correct funding rounds but also consistently being able to present that data for people to access along the same network, Along the same network. That's how we are trying to do that convergence.
Mm. Question on that Brian Martin, uh, signal 65. So I'm managing a handful of 32 GPU and 64 GPU clusters, just like you're describing here.
Uh, when you converge front end traffic, storage, traffic back end GPU traffic, are there any special considerations you have? Absolutely. Okay.
And that was the point I was going to make. I have, I don't have the detailed slide because I didn't want to uh, go into detail conversation and spend the entire, but that's why we have a Cisco innovation over here leveraging our silicon one asic, which is our differentiator based on the exact requirements with customer mentioned. Mm-hmm.
And even large scale customers are looking for this when they converge their north, north, uh, north-south networks. Mm-hmm. So this is what is called policy based load balancing in mixed mode environment.
So what we are trying to do is, because you would have your storage traffic, your backend GP to GPU traffic, you want to give that preferential treatment, right? So that your GPUs are not sitting idle. If there are, say for example an AI job is running, there will be, it would be distributed into a hundred tasks.
Those a hundred tasks would be distributed to a hundred threads within the GPU, right? And if you complete 99 threads and one thread is just waiting for the entire process to happen congested in the path, they're all waiting. They're all waiting.
So you need to make sure that the GPUs get the highest preference rate. Treatment then can be your storage because you're feeding all the data for training traffic to your storage network and the least can be for your front end network. So that's what we are able to do based on the policies.
Say, say that for my training traffic GP two GPU traffic, I want to make sure that it gets the best path possible with the highest load balancing scheme that I have. Then for storage can be the second priority and the third priority can be my front end traffic. Perfect.
So DLB and GLB can run on GPUs And Other products? Absolutely. Yeah.
Excellent, Excellent. Yeah, so that's what we are trying to do. So just to be clear, so that sort of supports a, a non-blocking network for the GP to GPU traffic, whereas blocking elsewhere?
Yeah, blocking in the sense that, say your storage can have eight, 400 gig ports based on the storage nodes that you have, and then front end can be anyways, the traffic would be like very lower. So it would be just one 400 gig port and so on that connects to all.