Introduction to Cisco AI Cluster Networking Design with Paresh Gupta
Paresh Gupta, a principal engineer at Cisco focusing on AI infrastructure, began by outlining the diverse landscape of AI adoption, which spans from hyperscalers with hundreds of thousands of GPUs to enterprises just starting with a few hundred. He categorized these environments by scale—scale-up within a server, scale-out across servers, and scale-across between data centers—and by use case, such as foundational model training versus fine-tuning or inferencing. Gupta emphasized that the solutions for these different segments must vary, as the massive R&D budgets and custom software of a hyperscaler are not available to an enterprise, which needs a simpler, more turnkey solution.
Gupta then deconstructed the modern AI cluster, starting with the immense computational power of GPU servers, which can now generate 6.4 terabits of line-rate traffic per server. He detailed the multiple, distinct networks required, highlighting a recent shift in best practices: the front-end network and the storage network are now often converged. This change is driven by cost savings and the realization that front-end traffic is typically low, making it practical to share the high-bandwidth 400-gig fabric. This converged network is distinct from the inter-GPU backend network, which is dedicated solely to GPU-to-GPU communication for distributed jobs, as well as a separate management network and potentially a backend storage network for specific high-performance storage platforms.
Finally, Gupta presented a simplified, end-to-end traffic flow to illustrate the complete operational picture. A user request does not just hit a GPU; it first traverses a standard data center fabric, interacts with applications and centralized services like identity and billing, and only then reaches the AI cluster’s front-end network. From there, the GPU node may access high-performance storage, standard storage for logs, or organization-wide data. If the job is distributed, it ignites the inter-GPU backend network. This complete flow, he explained, is crucial for understanding that solving AI networking challenges requires innovations at every point of entry and exit, not just in the inter-GPU backend.
Presented by Paresh Gupta, Principal Technical Marketing Engineer. Recorded live at Networking Field Day 39 in Silicon Valley on November 6, 2025. Watch the entire presentation at https://techfieldday.com/appearance/cisco-presents-at-networking-field-day-39/ or visit https://techfieldday.com/event/nfd39/ or https://Cisco.com for more information.
Transcript
Alright, good morning everybody. Uh, I'm Perge Gupta. I'm a principal engineer at Cisco.
Uh, I focus on AI infrastructure. Uh, and today I wanna talk to you about the AI cluster network design and operations. Uh, you know, right after the session, I'm getting in a plane to go to Melbourne, Australia.
We are hosting Cisco Live, even there, uh, the APAC version of it. And there I have a four hour session going in details on the same topic. So what I've done is summarize it into one hour and try to share with you, but ask me any questions.
I don't think to all of you have to tell that. Feel free to ask questions whenever you like, right? So look, one of the things is in this role in Cisco, I have the privilege of working with all kinds of customers across the world, right?
So on one end we have hyperscalers who have tens or even hundreds of thousands of GPUs. There are conversations of going into millions of GPUs. And then on the other end, we have enterprises who may have like hundreds or even thousands.
A lot of enterprises are starting to adopt this whole AI infrastructure. And in between we have neo clouds, you know, companies who are offering GPU as a service. And then we have sovereign clouds, like countries keeping data within their geographical boundaries and still offering this massive computational power of the GPU.
Right? Now the requirements are different and that's why I wanna bring things into perspective that not just the number of GPUs, also the use cases are different. Like hyperscalers may train their models from foundation, but uh, the enterprises may not do that.
They may just do fine tuning or inferencing also the locality. And there are terms like scale up. So we talk about scale up, which is all the GPUs within the same high bandwidth domain.
In simple terms, think about all the GPUs that are within the physical enclosure of a server. And then there are things like scale out where we take multiple servers and connect them each other and then scale across. You know, if you have heard about recent announcements of companies building gigawatt watts of capacity of GP infrastructure at one location, power is not enough.
So they want to have multiple of these locations and then connect them with each other. And that's what scale across is now enterprises, at least still now, you know, talking to other customers have not heard that yet, right? So scale up and scale out is the most common environment.
The other thing is most of these, uh, you know, hyperscalers have massive r and d budgets. You see they are writing research papers, uh, not just the research papers, they're also writing their custom software, right? And the opportunity size.
I think yesterday we were talking that what exactly is the opportunity size? Although we have few hyperscalers in the world, the enterprises are in thousands. And the reason I'm bringing things into perspective, because you, you're gonna see in a few minutes when I talk about the features that, the innovations that Cisco has brought to the table, the key reason is what use case it's solving.
Because the environments are different for hyperscalers versus enterprises and therefore the solution must be different too, right? I'll give a classic example where some of the congestion control algorithms in the network, the way hyperscalers have made it work, may not directly work for enterprises because there's a lot of knobs, a lot of fine tuning enterprises may not spend that kind of time there, right? So after setting the background, what I wanna do, go back to the very basics and start with what exactly is an AI cluster right?
Now when we talk about AI clusters, the core of it is the massive computational capacity of the GPUs, right? And these dus maybe within the servers right now, I'll take the example throughout this whole session, the example of the servers that have eight GPUs within them. There are also different environments, but let's, for the sake of argument, we are gonna stick with that.
Now, if you look inside these servers, and I'm taking example of two of Cisco's UCS servers. There are, you know, other vendors too doing similar kind of servers. But on the left you see, you know, the previous generation of server it has eight GPUs and then GPUs talk to other GPUs and different servers using eight 400 gig next, right?
And then these GPUs talk to the external world using 200 gig links or maybe 400 gig links, but that's already last year. Now we have the next generation, like the, the Blackwell series from Nvidia B 300 GPUs that use connect X eight nicks and there are eight of them each. Nick has the capacity of 800 gigs.
So imagine eight, multiply by eight, that's 6,400 gigs of capacity per server only for GP to GPU communication. And by the way, this is not theory. All these GPUs when they communicate, they do communicate at line rate.
And I thought that it's only a theory until I saw that happening in my own labs. And that's when I believe yes, this is real, right? So because of that, there will be some unique challenges that we'll talk and like I'll explain how Cisco is addressing those challenges.
So in the AI cluster, of course when we have GPUs, we want this GPU to make accessible to users like all of us, right? So that's where there is a network called, we call it as a front end network. Then of course when these GPUs work, they need data to churn on, right?
Without data, you know, what, what are we gonna compute on, right? So there's a storage network. Now if you go back maybe last year or two years back, I went on Cisco live stage and said that uh, best practices you create isolated, dedicated network for your front end and storage.
Guess what? It's not happening anymore, right? So we learn, we change our recommendations.
So now what what I'm seeing happening is like most of the time people are merging these networks. So there's a converge frontend and storage network. Of course when we talk about storage, there must be storage appliances out there, right?
So there's a high performance storage and then depending on the type of platform, there might be a backend storage network as well, right? Depending on the type of storage appliance. Vast is one classic example.
So Cisco has done partnership with vast uh, platforms and the vast platforms require a dedicated backend storage networks. I have a question for you on that. 'cause you know, thinking back to 15, 20 years ago when storage network convergence was a huge thing, we, you know, went, debated it and we went back and forth and some storage networks went converged, they go over ip, some still remain separate.
How is this different than that choice? Is that, do you see the separation of networks? Just why, why is it going converged versus separated?
Is it, is it cost? Is it efficiency? What's, what's driving it?
It's primarily driven by cost, but like the key technical reason is when people deploy. Because look, we don't wanna merge storage traffic with the rest of the traffic. That was the basic premise of having separate.
Right? After running these environments for a while we realize, you know what, there's not much traffic on the front end network, so I'll rather dedicate my storage traffic to that network too. It's 400 gig capacity.
Okay. People learned and then you know, they said that we can save this cost and we did that. Okay.
Right. Uh, the previous example that you took about the history, the only differences I think historically we had different like fiber channel protocol, FCOE, all those things, you know, have a role to play in convergence, storage, traffic with the other usual traffic here it's all IP based. Okay.
Right. Is it, are you, because I mean because this is new, are you going forward with IPV six only or are you still using ipv four for this? What does that both Options work?
I'll cover the, uh, I have, I mean the headers save you Free your slides, okay. Yeah. But short answer is both we before V six?
Yes. Okay. I have a question.
Um, when will you have a backend storage network versus not having one and having just the high performance storage for the front end? It depends on the storage platform. So as I said, vast requires a backend uh, storage network.
There are other platforms like DDN or wca, they do not require, it just depends on their architecture, right? The way that they design their whole product. And the way this is shown, it's as if it's going from the front end network to the front end, high performance storage and then to the back end.
Is that the actual flow? I will cover that. That is the flow.
Short answer is I have a detailed slide explaining the traffic flow. Okay, great. Thanks.
Right? Okay. High performance storage.
But look, we don't require high performance storage all the time. What happens, you know, take logs, what happens? There are some images, right?
For that we, standard storage is enough, which is more affordable. And then for all these GPUs to talk to each other, some of the jobs might be so large, a single GP is not enough. Even a single server is not enough.
So we want all these different servers to talk to each other. That's why there's a inter GPU back in network comes. This is dedicated for GPU to GPU communication only.
And then you can imagine there are a lot of devices here. So you wanna manage all of them, right? Using out ofAnd connectivity.
So there's a management network, right? So there are multiple networks involved and this is what we call it as an AI cluster. But remember there has to be peripheral around this AI cluster.
We don't offer just AI cluster to end users, right? There will be typical data center fabric and within the data center fabric there will be other kind of services like uh, you know, organization wide data. There might be some apps that have agents running within them and then there might be some centralized services, right?
The reason I'm bringing to perspective, because a lot of times when we talk about ai, a lot of us in the network industry, we just focus on backend network, GP two, GP communication. But you'll see that how much innovation have done in other parts of the network too, because those are the problems that are yet to be solved. Is this architecture primarily for when you're doing model training or is it also for inferencing?
Same, same model training inferencing, uh, the architecture remains most mostly the same. Right? Uh, uh, I will briefly talk about the use cases.
There are deployments where people say, Hey, I just wanna have an inferencing only cloud, right? We don't wanna focus on model training. Then probably you can, you know, debate about the InterG backend network.
Do you really need or not? Right? But most of the architecture remains the same, right?
Forgot Sam, Mitch Ashley with fu You always decide which one goes first. Oh, Rita Younger, very quickly that on that graphic, the inference, uh, the inner energy PU backend, that is for the training, correct. Where the inferencing is going to be your data center fabric, et Cetera.
Yeah. There is also distributed inferencing happening, right? Interesting.
So any distributor job, that's what I'm saying, distributed job, primarily training, but there is distributed inferencing too. Okay. So JD here, um, modern, you know, just built network, um, with all the latest and greatest, where are the bottlenecks on this?
We'll cover that. I just wanted to set the stage right because as I said, if I say AI network, people think it's back in network. Mm-hmm.
Right? It's not just about always GP to GPU. There's a lot of, uh, challenges with the storage traffic too.
So we'll talk about that. And for that you ask for a traffic. This is a very oversimplified traffic flow.
So imagine like users like us when we try to access an AI app. It's not like I'm going to directly land on a GPU server. My will probably go to an app and then there would be have to be some other services involved.
Like let's say, um, identity services, like authorization services, maybe billing services. So there would be some other TRA traffic going on in a data center. And only after that the request will land into one of the GPU nodes via the front end network.
Right now, depending on what the job it is, the GPUs may decide to read and write data to the high performance storage and based on the storage platform. Again, that may generate some traffic on the storage backend network. Remember, the storage backend network is different from the inter GPU backend network, right?
Also, when the jobs are running, you wanna take some logs, you wanna take some backups, so you will end up accessing or the rather, these GPU nodes will end up accessing these, these standard storage. The standard storage may also try to read some data from the organization wide storage. If that data is not into hot data as such, or tier one data as such, if the job is distributed, maybe it's training or inferencing, that will generate traffic on your inter GPU backend network.
And finally the request will go back to the user. It may go directly back or it may get routed by the application again, right? Depending on how things are configured.
Is this all making sense? Because I will take the examples from this traffic flow that every entry and exit points, the features that we have developed, the simplicity that we are providing to serve the use cases here.