Building AI Pods with Nexus Hyperfabric from Cisco
This presentation introduces Cisco Nexus Hyperfabric, a cloud-managed platform that simplifies the deployment and ongoing management of AI infrastructure. It addresses the growing need for repeatable, scalable, and operationally efficient networks specifically for enterprise AI clusters. Cisco emphasizes that while hyperscalers build immense AI factories, a significant and growing market exists for smaller, enterprise-level AI deployments, often below 256 nodes, which they term “AI Clusters for the Rest of Us.”
The shift to these smaller, on-premises AI clusters is driven by several factors, including the increasing size and sensitivity of data (e.g., healthcare, intellectual property), making cloud undesirable, a trend of workloads returning from the cloud, and the need for project- or application-specific infrastructure rather than shared general-purpose IT. The rapidly evolving AI technology also means enterprises prefer incremental build-outs rather than massive, infrequent investments, allowing them to leverage newer generations of hardware more frequently. However, designing and deploying these dense, complex, lossless Ethernet networks is challenging and time-consuming for traditional network practitioners, often involving weeks of design, lengthy procurement, and meticulous cabling.
Cisco Nexus Hyperfabric addresses these challenges by delivering a Meraki-like SaaS experience for data center network deployment. It offers pre-designed, NVIDIA ERA-compliant templates for AI clusters that automate the generation of a complete bill of materials, including optics and cables. This drastically reduces design time and eliminates manual errors, accelerating the “time to first token” for AI projects. Hyperfabric also streamlines day-one operations with step-by-step cabling instructions and real-time validation via server-side agents, ensuring correct physical connectivity. Beyond deployment, it provides end-to-end network visibility, proactive monitoring of components such as optics, and integrates advanced Ethernet features, including lossless capabilities (PFC, ECN) and adaptive routing, to optimize performance for demanding AI workloads.
Presented by Dan Backman, Distinguished Technical Marketing Engineer, Cisco, and Alex Burger, Principal Product Management Engineer, Cisco. Recorded live at AI Infrastructure Field Day in Santa Clara on January 28th, 2026. Watch the entire presentation at https://techfieldday.com/appearance/cisco-data-center-networking-presents-at-ai-infrastructure-field-day/ or visit https://techfieldday.com/event/aiifd4/ or https://www.cisco.com/ for more information.
Transcript
My name is Dan Backman. I work, uh, with Cisco on the Hyper Fabric team. Uh, so I've been involved with Hyper Fabric for about three years now since we originally, uh, specked out and built out the product.
And we're happy to talk about what we're doing with ai. I'm joined by my partner in crime here, Alex Berger. Hello.
And Alex has, uh, been working on a lot of the AI deployments with our early beta customers, and we're happy to kind of dig into exactly some of the things that we're building out and why. Um, I'm not sure you can put up a logo like that without explaining it. Fair.
There's also little stickers back there that we're happy to share as well. Moral Imperative. So originally Hyper Fabric was called Project Tortuga.
Mm-hmm. Uh, and Project Tortuga, uh, involved like the overall product creation and sort of expansion into AI as well. Um, the question came up is a Tortuga like Spanish for turtle, or is a Tortuga, like Pirates of the Caribbean?
Ah, so we decided why not both. So ever since then, a Pirate Turtle has been our theme. Uh, so you'll actually see, um, uh, just if you want a little bit more information, we did a tech field day on this about a year ago, uh, in Amsterdam.
Uh, and at that point we focused a little bit more on what is hyper fabric and how it works. If you want to know any more details about that, there's a full session. Uh, I'll give you the very quick TLDR and what hyper fabric is.
So hyper fabric, you can think of it as a Meraki like experience for deploying data center networks. So what we focused on is it's a different operational model and a different workflow. The idea is you can stand up a fabric, pre-design it, and cloud plug in switches.
They connect to cloud and they will dynamically provision the entire fabric. In fact, as soon as you plug in the switches, we will dynamically build the actual network model. And it's actually a full EVPN VXLAN fabric with a substantial amount of capabilities that serves both enterprise customers as well as AI scenarios.
As we go through this, just, uh, the, the key thing to think about is it's a different way of deploying and managing these clusters. Now, we did wanna spend a little bit of time talking about AI clusters in the enterprise. And the important part of this discussion is really we are seeing a lot of different AI cluster deployments.
And one of the things we wanted to share with you is everything that you see, a lot of the marketing material we're talking about huge AI factories. We're talking about, uh, about hundreds of megawatts. I'm still waiting for us to get to one point 21 gigawatts of power.
We'll get there soon. Um, but in many cases, something this big usually starts off with a question of which nuclear power plant are you co-located next to what body of water you're gonna sink the heat into? And where are you building the new building for all these racks?
Because none of the requirements fit into a traditional data center. So there's a lot of work happening on this, and it's really important to keep up on where that technology is going. But one of the things we've been focusing on a lot is we are starting to see enterprises deploy more and more AI clusters.
And I know, uh, I, I know at least one of you is talking about running several clusters in sort of the 256 node or less area. And that's really where we're seeing a lot of the growth here. And we see that there's an opportunity.
There are differences, though. These are smaller fabrics, they still need the ability to grow. But there's a really interesting thing that we're seeing in that customers are building out multiples of these smaller fabrics, not always the really big fabric.
And we'll talk about some of the reasons behind this. So we actually entitled this AI clusters for the rest of us because there are people out there who are building AI factories and honestly, you're not gonna listen to us to tell you how to build them. But what we are seeing is that there's a lot of our customers that are putting their toe in the water and having to build some incredibly complex network designs where we have the opportunity to actually help them deploy these that actually speed the workflow.
And we'll talk about why we built that. So let me also get in front of the other obvious questions. So Megan had a fantastic presentation talking about how we build and maintain large AI clusters.
One of the things you see is there's a lot of details and around understanding RAs, understanding the right class of service. These are all the details that you want. The ability, if you want to control these, fine tune the networks.
We need the technology to allow you to scale up and also fine tune how these networks build. But we also have a lot of customers that time is of the essence. They need to deploy these quickly, fast and repeatably.
And that's really the difference of why we build hyper fabric. If you want deep low level CLI config control, we have tools that allow you to do that. And we as Cisco have the benefit of being big enough that we have different customers of different sizes.
We need, we also have to give right size solutions for what they build, which is why we're talking about hyper fabric here. Um, the other important thing that that you're gonna see is these enterprise AI clusters are looking at what some of the bigger technology is doing. They're slowly coming along.
Like right now, most of these clusters are fitting inside of ex existing data centers. If you can find the right amount of power, most of them are still air cooled, but we know even those clusters are gonna start moving to liquid cooling and start to need some new power, uh, build outs. The other thing that we see, especially on the small site in the enterprise, people want ethernet.
We've already proven that ethernet can give the, give the actual performance that you need to support the actual backend network connectivity. Uh, so we know that ethernet is the right tool for the job, but it's also the tool that enterprises understand. So one of the other major things that we've seen is a big focus on that.
Mm-hmm. So what have we really learned about these AI clusters? Well, there's a few interesting things that we didn't really expect at first.
One, we are seeing in general, not just in ai, but in data center in general, a lot of workloads coming back from cloud, but when they're coming back from cloud, they're coming back to specific locations for specific reasons. And that has a big impact in how people are building these AI clusters. We're also seeing that enterprises tend to do these in a project or application specific build out.
It's not just building out a single shared infrastructure. 'cause the money comes with the project, not generally from a big IT budget. Third, the other thing that we're seeing is obviously the technology is evolving really quick.
And this is actually the other part that the other reason why we seek enterprises doing so many more of these deployments, because obviously if you get done with one, by the time you deploy the next one, the technology has changed. And you actually have to revisit even things like the network architecture and the actual hardware architecture. And in the end, if you haven't done this before, everybody in this room understands a lot of the complexities of an AI cluster.
But if you are an AI or if you're a standard network practitioner, there's a lot to know about these architectures. These are one of the most complex, dense networks You've seen Ray Leese, Silverton Consulting. Do you know why they're rehoming those applications?
Yeah, we'll talk about that right here actually. Um, so good question. So one and uh, and I'm gonna invite in.
Alex worked with a lot of our customers who are doing this. Um, there's primarily two reasons. It's either size or sensitivity of the data.
So Alex, you've worked with a couple customers. What was causing them to actually bring these workloads for AI on cloud or back from cloud rather? Yeah.
So in general, we've, a couple of the customers we've been talking with and working with, um, some of them have incredibly large data sets. Like, uh, one of the healthcare research companies that we've been working with has like, in some cases, like a petabyte worth of data per like, you know, data set that they actually need to process. And trying to get that into the cloud can be really difficult, especially if like that imaging is done locally.
Um, in some other cases we've seen customers have, you know, sensitive information, maybe it's intellectual property, and they've been wanting to try to keep that as localized as possible. Um, those are the two core reasons. So usually size a data set, and then a lot of times the type of data, um, is becoming less and less ideal to have hosted And distributed inference as well.
If you're doing like inference on video, you wanna be able to move that inference directly into the, uh, source of the video instead of having to back haul it. Yeah. Yeah.
Fred Van, her and Hyphens Consulting, that's exactly what we were going to say is it seemed like it was heavily focused on training, right. Where large data sets are coming back. And you kind of already started answering it on the, on the inference side.
So even on the inference side, you see this coming back. I mean, the public clouds are very popular with elasticity. Oh, they Are.
But it's a question to where's your data, right? So in many cases what we are seeing is there are people who want to have a little bit more control over this. And instead of deploying sort of a public cloud instance or trying to pipe that video back, you may be able to, you may, there's, everybody always talks about latency.
There's not a lot of cases that we've found where latency really matters that much. But if you're trying to make decisions on video, uh, on video or other data feeds, that's a place where latency really matters. So that's a place where we've seen people try to move it closer.
Right. And then, and then maybe an additional question, when they think in small clusters, do you see them building small clusters that do both training and inference? Or do you see those clusters being separated?
I think we, we've seen them target a little bit of both. What have you seen and who are the and the folks we're working with? Yeah.
So in a lot of cases, like training jobs take a long period of time and can fan across in a significant amount of the, the resources available. Uh, so it depends on kind of, it's an, it depends answer, but a lot of cases they're separate so that you can keep things that need that run quickly in one place and then longer term, um, you know, operations that can kind of sit in their own bubble If Right. Because the design is different.
Right. One more batch, the other one is latency. And so you can't really have both, right?
Yeah. You can try to optimize some things, especially in the network side. You can deliver a network optimized for both.
But what we're, one of the big things that we're seeing though, is that you're seeing more and more of these clusters get built because the elephant in the room is, right now, most of our enterprises don't exactly know what they need their clusters for. There's a lot of exploration happening. So the other piece that we're seeing is they're gonna build out a cluster.
And the reason for building out the first cluster is like our CI, our CIO said we needed to do it, what do we do with it? And then as people start using it, they start figuring out what they want to do with it. And then they start looking at scale.
We're seeing these this incremental build out. And especially for the case that you're looking at, that's where they would say, oh, well I'm already, I already have this cluster full of training. I need a new one so that I can start to do inference.
So it actually comes back to the fact that we're not seeing these single big build outs. We're seeing lots of little ones. And I think the other really important part, Alex mentioned this from one of the customers that we work with, that's a research institution.
I never would've thought of this. If you have all of your GPU resources that are being rented in cloud, it's actually discouraging research. If all of your researchers need to worry about how much they're paying per token, they're gonna start holding back from the exploration that you need to do to actually figure out how to develop the technology.
So especially for research institutions, the reason you wanna bring it in is you CapEx it once and you get to run it as long as you want. You can lifecycle it for a long time, hand it down to graduate students, and it really allows much more research to happen than you ever would if you're actually paying, uh, at, uh, for the usage on those, uh, clusters. Mm-hmm.
Pete? Well, um, this is very interesting to me because it explicitly, uh, recognizes something that I think has been coming more and more conscious over the last couple months, which is the whole cost driver. Maybe the giants with their AI clouds is not what everybody needs, and particularly the security confidentiality of the data has all of a sudden, it seems to me anyway, come onto the horizon as a real big driver.
And what a couple of us have talked about offline is just, uh, the problem of the huge data and getting near your data and all the different models where you might have edge data mm-hmm. As opposed to this petabyte in one place. And that starts getting pretty challenging.
Absolutely. And, and also, let's just be super clear for the audience. A lot of this is gonna happen in cloud.
Everybody starts at cloud 'cause this takes a lot of money to build one of these. So you're not gonna build these till, you know you need it. But again, I'm, I'm really glad to hear that.
'cause we are seeing the fact that people are really saying we really need to start insourcing a lot of this. Well, I have a pet hobby horse, which is, I'm sort of suspicious of large models, uh, run by the cloud providers. They're great for natural language type stuff.
Mm-hmm. But I'm wondering whether we're gonna see a proliferation of topically specific smaller models, because you can be much more efficient. Yeah.
You have the training cost and, but you also have a unique data set so you're not spinning cycles on, uh, the whole world, so to speak, or all of language. And it just seems like that might be a, a possible trend here. Absolutely.
Yeah. And I, we are very curious to see how this goes. We're network geeks, so it's really fascinating to see what the researchers and, and the customers are doing with these models.
Um, Hey, Dan, uh, Brad, Greg? Yes, sir. C strategies.
I'm a network geek too. So, uh, one of the questions that, that, uh, I get asked a lot. I'm, I've worked with a couple of startups mm-hmm.
Uh, realistically, you know, when, when we talk about building these EastWest clusters and stretching them mm-hmm. What latency wise, I mean, how much can you stretch a cluster? Because one of the, you know, and we're starting to see it more in the US but the, the startup I was talking about specific, specifically in the uk, they're actually just out, uh, buying real estate, right?
Mm-hmm. And they actually wanna stay below a certain, uh, power threshold, I think 500 k for some regulatory things in the uk I'm not familiar with, familiar with. But, um, their solution is, you know, customer calls up and says, I want, you know, a, uh, I want five gig of capacity, right?
Or whatever. Um, they say, okay, they go find 10 of these buildings that are, that are out there and starts stitching together. So real realistically, how big of a ring, how, how big of a, you know, a wan space can you realistically expect to do EastWest training for GPUs?
I'm gonna give you an answer you don't like, depends. That's not my expertise. Alright.
I don't know where stands I I can, I can tell you what I've seen though. Yeah. Um, and in general, there's sort of two ways to deal with it.
And there's actually three things to think about. Um, number one, we've seen from some of the really large build outs, either the AI factories or some private people that are doing interesting stuff with them. Um, what you tend to see is no matter what, when you build out one of these clusters, if you go for sort of a multi scalable unit design, by definition, none of those fit inside a one building.
Yeah. So you're gonna start to get, uh, get into scenarios where, okay, okay, these pods are in this hall, these pods are in this hall. You have to build a tremendous amount of fiber and interconnect between those to maintain that subscription level.
But at that case, you're looking at like fractions of milliseconds in terms of latency, but you're, but you still have links that are longer than the others, right? Yeah. And that starts to look like a super spine deployment.
So at that point, what people start to do is start to consider affinity for locality for these GPUs when they start doing these jobs. But that's one way to do it. And especially in the really big build outs, what you see a lot of times those are also multi-tenant.
So the, the actual workload management systems that are doing tenant allocation will wanna understand the affinity of where those GPUs are located and then try to do best fit for those multi-tenants. Um, on the other side, there's other, there's a couple other things to consider. One of the big ones is how do you actually start to peer these fabrics together?
Yeah. That's one of the things that we've been spending most of our time looking at. So for us, we know that we can build fabrics very, very large.
We can go to super spine, but at some point you need to peer these fabrics either northbound with existing networks, so you can get multiple tenants in. You can integrate that with an external network, or you may have different data center build outs. And in our case we're, we, wait, we went straight to EVPN VX land.
'cause it gives us tr tremendous amount of flexibility in the underlay network model without a really big penalty in terms of overhead. But at that point we can actually start to leverage things like border gateway so we can natively peer fabrics even in different locations together over a third party network. Yeah.
Kind of a stretch super spine if you think about it. Yep. So we are, we're actually putting a lot of investment in building multi-site capability just so that we can actually start to grow these fabrics.
So just on the network side, that's how we are looking at it. Yeah. I think, you know, we played a big game of Tetris with within the data center.
Now we gotta play that big game of Tetris, these fabrics all over The place. I think that's absolutely valid. And I, I'm, I'm waiting intently to hear what the actual latency boundaries are.
'cause I know even in even traditional virtualization, people spout a bunch of figures about vMotion latency and I've seen it go for hundreds of milliseconds, so. Right. Thank you.
Uh, no worries. Thank you for the question. And, uh, Andy Banta, uh, just a, a question going the other direction.
Uh, you're, you're talking about how you can go larger and go into, you know, multi-data center. Uh, can you talk a little bit about, um, where you're headed for the density story for, um, attempting to make things more compact and, uh, less power hungry? Um, I can touch on that a little bit.
It's a little bit that's more of a generic Cisco question. There is this constant work on how big, how, how do you improve switch radis? How do you get more bandwidth?
Uh, I can tell you that definitely Cisco is continuously working on how do we get better silicon? How do we put together better systems? What we've generally found is scaling up these networks are generally best done in these discrete devices.
So we're focusing mostly on fabric based technologies where you can easily start to linear or to, uh, horizontally scale those out. But we, we have a long roadmap of increasing, uh, overall size of silicon. We're leveraging spectrum four as well as we talked about in the 9,100 series.
So what you'll see is there's gonna be continuous evolution there, but what we see is that's gonna keep on happening. But the operating model is where we see a lot of the pain. Okay.
Well then to ask the, another question based on, on, you know, spreading these things out, uh, are you starting to build, uh, ultra ethernet considerations into any of your topologies or, um, structures or infrastructure? Uh, so absolutely. I mean, ultra ethernet is a big moving target with lots of pieces.
So one of the things that we are doing is we're actually tackling a lot of discrete technologies inside. So we're starting off with sort of the obvious lossless capabilities, P-F-C-E-C-N, uh, lossless capabilities, uh, WDRR waiting so that we can protect things like, uh, ECN packets back there. There's things like that that we're building in.
We're also, we will talk about a little bit later, well, so working closely with NVIDIA for their adaptive routing support, which gives us really very simple, uh, and highly effective load balancing across these fabrics. So there's a lot of pieces that are very similar to ultra ethernet. Um, as that technology evolves, we will intercept that as well.
We believe ethernet is the right answer. There's a bunch of things like packet trimming and things like that that are happening. Assume that all of that is on our radar on the development scope as well.
Right. And it's, it seems that, uh, there would be some opportunity to actually, uh, slim down the amount of, um, network that you need between various different buildings and whatever and data center if you actually do use some of the technology available. So I would, uh, my question is, uh, when you're building out these architectures, are you actually thinking about building some of these capabilities into the, the architecture to take advantage of, uh, like packet spraying, uh, you know, out of order delivery, that type of thing?
So that's actually something that we could already do today, uh, by leveraging the work that we're doing with adaptive routing with nvidia. So that actually does leverage packet spray. It does that.
It is actually tolerant to out of order packets. So that's stuff that's already happening. It does actually happen end to end in conjunction with, with the, uh, actual nicks in the server.
So there's uh, some Nvidia technology in there as well, but we're already working on that. We're also looking at adding packet trimming for improving a lot of these, uh, scenarios. In the end, one of the biggest things we can do though is increase the bandwidth because, uh, no amount of qua gives you more bandwidth.
Okay. Um, so with this, I'm gonna speed through the next ones 'cause we covered these a little bit, but I think this is really important from what we're seeing. The other key thing that we see, if the, if the hypothesis is we're seeing enterprises build a lot more smaller fabrics.
The other reason is this, every year there's new hotness that gets announced that it's about to come up in, in a couple months we're gonna hear about a new generation of asics. Mm-hmm. So what we tend to see is because there's this constant turnover, what we're seeing is our enterprises say, well, maybe I don't want to invest all my money in 2026.
I wanna save some so I can build a new cluster in 2027 to take advantage of that new technology coming out. So I think this is actually one of the things that we're seeing is that at least the customers that we're dealing with are starting to take a very lifecycle view for these clusters. Um, I won't touch much on this, but the other thing is, if you are new to AI clusters, especially if you're a network guy, these eras are very powerful and declarative, but they're not always written for network guys.
Uh, there is a fair amount of detail that you have to suss out from these and a, a fair amount of, uh, work that you have to do to really figure out how you want to build and scale these apologies. And one of the things that we see is if you're starting to do a lot of these, every one of these cluster build outs is a pretty significant investment. So if we can help our customers, just, it sounds kinda lame, but you gotta figure out what the right answer is before you click the order button.
If we can help our customers with that to help them deploy that, we think that there's a lot of value there. And then sort of the last piece, I think, uh, Megan covered this really well, but uh, there it is. It was a little slow.
Uh, uh, what are the, one of the key things here is the network really is the glue between these. And one of the things you find is on deployment side, there is a fair amount of detail in the topologies that you're deploying. You have to plug everything in correctly.
Um, you do have to make sure that RDMA is working everywhere. Uh, these are not simple network topologies and as any network guy knows, it's always layer one asterisk unless it's DNS. Right?
But there's a lot of layer one in these. Um, and this happens both day, day, day zero in design. It happens day one in deployment.
If you're plugging in hundreds of cables, that's a non-trivial task. Mm-hmm. And troubleshooting that and getting that right, this is the part that people really kind of glance over when they look at this.
I like this diagram. I think, Alex, you did this one, this is fantastic. Um, this gives you an idea if you're a network guy, wow, that's laggy.
Uh, you're about to see a diagram with lots and lots of wires on it. Uh, but this is really kind of what you're in. Uh, what you're in for is a sm even a small cluster can have hundreds of interconnects.
And if it's done on fiber, then it's literally double that in terms of transceivers. But anyway, this is basically, we like to show this to customers because this is what you're in for. You gotta build this.
And it's not, you gotta design it, you gotta order the components, you have to install it. Then once it's up and running you have to keep it working. So, can I ask something really quick?
Oh yeah, please go ahead. Regina Rosenthal from Digital Sunshine Solutions. And you, I love that you're saying all this and you're so worried about the customers.
What are some of the most common things they ask? I, I see, like for me as a, from a product marketing view, I see a lot of, um, talking and how how do we compare? Like there, there are already the knowledge they have about networking.
'cause a lot of this looks like it's layer one. It's like how do we compare to call it an AI cluster? What does that really mean?
What are the key things that like trip them up that are different from normal networking? Um, I will give an answer to this and Alex is gonna give a better answer to this. So in general, when I talk to customers about this, it's um, unlike a typical network, it's a highly optimized system.
So first it's a system. So everything really does work end to end. And most of the troubleshooting that you're gonna get into has to do with application behavior.
And you have to kind of suss out what the problems are underneath if you're not getting that right. So I think the amount of signal that you get as a network engineer is lower in one of these clusters. 'cause you're gonna get a call that says, my job completion time went from, you know, two weeks down or down to four weeks.
What did, what, what went wrong? And it could be GPU that had a problem. It could be a server that had a problem, it could be a software issue, it could be a network problem, it could be drops in our DMA.
And it's up to you as a network person to have to figure it out. Because the unfortunate truth is just like it's always layer one or DNS. Exactly.
It's always the network first until you can prove that it isn't. So on our side, what we are trying to do is look at for that experience for, for our users, because the designs are very, very detailed and have to get done just right. You are basically, in many cases spatially engineering traffic in the topology and then matching that to the right class of service config.
These are the things that most people have done a little bit in their CCIE lab. But generally in most enterprise networks, bandwidth takes care of a lot of cost problems. And you tend not to run into as many of these flow overlaps.
Like network guys aren't, in many cases, aren't ready for the fact that you can have a fat 400 gig flow from one GPU to another that could actually go for a long time. Mm-hmm. And if you're load balancing on 400 gig links, you only have one chance to get that right.
Once you have two of those that land on the same link, no amount of clause is gonna double your bandwidth. So these are some of the discussions that we would have. Alex, you have these discussions with customers and I said too much.
Um, what are, what are you hearing from them? Well, one of the, one of the first things I heard, which is kind of fun is most people are used to lossy traffic. Like, you know, TCP makes it really easy to have even a poorly designed network look good.
Um, with lossless traffic, any amount of inconsistency can cause Cascading effects to the topology. Especially around like, you know, running workloads. Like we were doing some, uh, we've been doing these EFTs with customers and working with them and our CX organization was doing some load testing and it was kind of fun.
You'll find out really quickly, um, or not really quickly within a few minutes of running a test that everything looked really good and then it nose dives. If you have a miss cable, maybe an optic failing or even then if you've improperly mapped like a, you know, DSCP value for lossless traffic. And so the biggest difference I think is just that things are like, it's really important to get things right and know exactly how things are deployed.
Um, and there's not a lot of tolerance for any amount of like kind of inconsistency config wise. So, And the answer to your question would actually be, this is the first thing I would talk to somebody about is this slide. So we put this together for this discussion, but we've kind of, we've intuitively known this, but we realized like we actually wanted to graphically show this.
Okay, very famous guy gets up on stage a few miles away and introduces a new set of GPUs. Start your CIO says I want a cluster. Okay, so it's on you.
What do you gotta do? There's a design phase. You have to look at the eras, figure out what your overall design is and you have to know exactly what your design is before you order something.
If you haven't done this before, this is not a day long project. This could be weeks or longer. You may want to consult with people, you may wanna understand that you're interpreting these eras correctly.
Okay. And at that point, now you start to order the gear. This is high performance stuff.
This doesn't show up overnight. There's lead time on getting the gear. I was gonna Say, I think you need a bigger gap between order and install.
I was trying to be as conservative as I can. I was trying to keep this inside of a year. And most people what we're seeing, if you put all these together, it's six plus months for all these phases to get done before you get to handoff.
And that's if things go well. Mm-hmm. And little stuff has huge impacts.
Alex is gonna show you a demo where we do things like cable plans and figure out transceivers. It sounds really silly, but that's the thing that causes delay at that install time. If you've ordered a whole bunch of transceivers and you suddenly realize these have MPO plugs, but all your structured cabling has lc, how long does that take you to fix?
Like getting in front of this stuff and realizing that you're on a timetable matters because in the end, the time before handoff, this is your time to first token or time to first value. And you as a practitioner are going to be judged on that. Okay, well that's just the first piece.
Once you've gotten that, now your cluster's in production, yay, everybody's happy. Okay, now your researchers are starting to worry about or starting to work on it. And now when you talk to your accountants, then they're trying to start to think about the depreciation cycle.
'cause you bought this gear, so it's gonna depreciate probably over three years. But first of all, how much of that depreciation cycle was in that time to first token that's if you take forever to deploy this cluster, you're losing the intrinsic value of those very expensive GPUs that you bought. Well, everybody who's deploying one of these clusters has one goal and that is get to handoff before the next generation gets announced.
Right? This is, this is one of the things you wanna be careful of because if it takes you a year to deploy this, there's already a new generation of asics out. Mm-hmm.
And by the way, as you continue to go through this, this is gonna keep on happening. The one thing that doesn't really change here though is that you have as an operations team, an ongoing burden of running this cluster. These clusters are not set it and forget it.
There's enough components that are running very, very hot. It's not quite as bad, but it's, if you remember the, if if, if you go to the computer to history museum, they'll tell you about the old mainframe days where they had people that were swapping out tubes. Every day you're gonna spend a non-zero amount of time troubleshooting transceivers that died or GPUs that died.
There is a fair amount of work that actually continuously happens. So as these continue to lifecycle, this does, this problem doesn't go away. So what we're seeing is, if you look at everything we've talked about, we're seeing more and more of these clusters.
But because of all these factors, what you're seeing is more and more of these build outs. So a lot of the work that you see is starting to multiply for our customers. And that's really where we see this come in.
And I'm gonna hand it over to Alex to show you some of the things we've done in hyper fabric to address that before you Go down that path. Ray ese Silver Drink Consulting is a hyper fabric a professional solution, professional services solution, is it a SaaS solution? Is it just how your GUI works?
Oh, it's, it's a great, great question. Hyper Fabric is a SaaS solution. It is a cloud controller.
It is a user workflow that Alex is gonna walk you through. It builds the network, it extends visibility into the servers themselves. So we have network layer end-to-end visibility inside the cluster.
And it allows you to actually go through day zero to day in and lifecycle manage the cluster in one product. So let me get in front of the other question. We talked about it a little bit earlier, like Cisco has different products for these things.
We have hyper fabric. If you're gonna go through this a lot and you need a guaranteed outcome, we've built something that'll help you lifecycle and get these deployed right the first time. This is something that if you know what you want, we can help speed it up.
We also have products where if you're trying to fine tune and get detailed configuration control, that's where we have products that allow you to actually get that level of control. But at the same time, there's sort of more to do on that deployment. So I know, I'm sure that question was gonna come up, so I wanted to put that out there.
But Alex, what don't you? Yes. Yes.
So let's say you had, I don't know, a 96 GPU 12 server environment. How long would hyper fabric take to get that configured, I guess is really the answer? Yeah, Let's do a demo.
I was gonna say also if we, um, just from some of the work we've been doing with like our own internal IT teams, um, if we look at a 256 GPU deployment, uh, they were able to go from obviously having the hardware but having everything deployed in about a week, maybe plus a couple days. And that includes like all the cabling, um, burn and getting things like ready to, So the order configuration was done long before that week began. Well, Yeah.
So I'll, I'll walk you through like what we, what we've done to try to speed up things. So it's very cookie cutter and easy to get going, but in that case, like it, yeah, I'll, I'll walk you through it. It, but we've tried to do things to make sure that it's not like a six month project to get to get things out the door.
So as Dan was mentioning, um, this is a SaaS service. Um, everything in hyper fabric starts with what we call a blueprint. And a blueprint is going to be, um, everything from the design to the actual operating environment that you'll manage and build configurations with.
com if you have a Cisco login if, and you can actually go and build blueprints and start working with it. You don't have to buy anything. And the main reason behind that was we wanted to give customers, partners and folks the ability to start getting, um, you know, in the actual dashboard itself, but also be able to use this as a tool for building out designs and then turning those into orders.
And so you can actually go and for instance, I can click add new fabric and if I click new from template, we have just common data center fabrics, but we've also got this tab here for AI clusters. The idea behind this was we worked with NVIDIA and we have like an enterprise reference architecture for hyper fabric. We've also got others for, uh, nexus dashboard and the rest of our data center portfolio.
But we worked with Nvidia to make sure we could build out and have templates that match the NVIDIA ERA compliant designs with the appropriate hardware speeds, et cetera. And so we've got a few different options. Um, as Dan was mentioning this is we've tried to do something that's repeatable, consistent, and so we have anything from like a small cluster, which in this case, you know, is about 4G PU servers or an NVIDIA scale unit.
Um, but if I scroll down here, um, you know, one of our larger clusters is like a 256 GPU, um, template. And so I'll click here, select, once I click select, we give you the ability to make modifications so it's not like a template you're locked in, can't do anything different. Um, you could get in here and you could actually know increase the number of switches in the backend.
For instance, we could modify, you know, maybe you want port side intake for air. Turns out a lot of people like having their switches all mounted the same direction. Um, some don't though.
So we had to give some optionality there. And then we even get into things like, uh, you know, the GPU servers. So today we have support for the our HGXH 200, Cisco 8 85 chassis.
We are gonna be adding some others like the smaller form factor 8 45 with the RTX, uh, GPUs. But we give you the ability to change quantities, change values, as well as we have a storage, um, provider. We're working with vast.
And so we have our C two 20 fives running vast storage. And so you can actually then build out also the storage environment. And so if we don't wanna make any modifications, we could easily go down here and you'll notice this nice little matrix of connectivity.
So if I, for instance, wanted to modify connections between different groups of devices, I could come in here and we've already predetermined the optics that we wanted to use. And in this case, these are 800 gig to 800 gig nicks that we're gonna be interconnecting. But if you needed to, and let's say we had another 800 gig optic, you could come in here and make a modification and save that.
And then we're actually gonna give you the ability to export this to a bill of materials. And so let's just say that everything here looks good. I'm gonna click save fabric blueprint.
And right now what we're doing is we're actually calculating that, uh, a lot of this really useful information that me as a person that doesn't like to build a bill of materials and I don't like excel enough to build cable plans. Um, really appreciate. And so I'm gonna run cabling.
Did you have a comment? Yeah, No. As, as Alex is doing this, the key here is what's really happening Is we're taking this template, but we're actually turning it into a data structure that is the configuration for the entire network.
There is no difference between the blueprint and a config. You could even have things like BGP peering, actual network definitions, multi-site definitions, all as part of that. It's actually generating the full network all at once and the cabling is done procedurally that work that you have to do to figure out that port on that switch goes to that port on that switch.
The reason we're running cabling is every one of those cabling groups, we are procedurally building those testing and actually putting in testing that the right transceivers will fit inside of that and doing all that work that you would normally do at the design time. So at that design time, we've done this in what, 10 seconds or so? Yes sir.
So 10 seconds to do, uh, three days worth of cabling, At least design plan. We haven't plugged it in yet. Plan, I should say three days worth of, uh, beating your head against the table due to Excel, um, for a cable plan.
Um, sorry. I worked with, uh, a couple different groups that were like, no, no, no, I need this in Excel format. And I'm like, oh, come on.
Really? Right. But do you print out the labels?
We've had that request. I'm not surprised. Did you tell me quantity and length of cables to order?
Hmm. Like specifically like is it gonna be at the top end versus the bottom end? Hmm, Absolutely.
Yeah. So one of the things I wanted to point out, and then I'll show you, uh, like the kind of some of the things we've, uh, added on here, but in this case, this is just the backend fabric connectivity for the GPUs for the east west traffic. And since we ran that auto cabling, we're actually gonna give you a step by step list of all of the different port to port connections that need to be present, including the rail group assignments.
But before we get there, let's say this design looks great, yes, let's spend a lot of money, um, we can go to this deployment tab, and this is where we've actually built out a list of every single PI and quantity that you need to actually order. And so you, we can scroll down here and you'll notice this thing scrolls for a while, which I had to build one of these by hand before we had this done. Um, I don't, I don't enjoy that at all.
It took a lot of work and you know, you can easily fat finger like an optic and then have something that is completely different than what you expected. Um, but not only do we give you this list here, we actually give you the ability to request an estimate ID and that'll go right into CCW. And the nice thing there is then I know I didn't make any mistakes.
I didn't for the audience what's CCW, sorry, Cisco Commerce Workspace, which if you're a customer you may not ever see CCW, but this is a big thing for partners and for Cisco to quote out and then provide, you know, order ability. Thank you for that. Mm-hmm.
I always forget that acronym is not necessarily well known. Um, so anyway, we we use this, uh, as a way to help make sure that you hopefully won't make any mistakes in what you order. And especially if you just take this right into that, um, estimate.
We've already vetted all of these optics, all these cables will all interoperate and work correctly. Yes. Wow.
I got a minute. Sweet. All right, gimme one second here.
So I just wanted to give an example of an environment that's operational because I wanted to talk through a couple things that are important. We are a network management platform, but with these AI pods we also need to be able to provide visibility into the server side. And so with a hyper fabric AI pod deployment, we also have agents that actually live on those storage appliances and on the GPU nodes so that we can see connectivity and validate cabling.
And so if we go to like the onsite interface, this is actually a, um, a link you can share and allow for like L one text to follow through and see and get feedback that they're plugging the right thing into the right peer using the right optics. Things are online, all the hardware's present. And if for instance, we have an issue, we're gonna highlight a nice little red box here saying like, Hey, check this optic and make sure that it is the correct optic.
And so we give you a very procedural process to deploy those fabrics. Even if you don't understand like how the network works as an L one tech, you could follow through this entire thing and then, you know, have things cabled correctly. Yes.
A feature like, uh, the capability to, to uh, flash the next port that you're supposed to be plugging into. We don't yet. That's a good idea though.
Put that in my pocket. Okay. I appreciate that.
No, um, we don't necessarily have anything that's gonna flash the interface, but we will tell you if you did plug it in the wrong one. Um, 'cause that will give you immediate feedback if that interface came up with the right optic right inserted. Okay.
But I mean, that means that you actually have to be looking at your screen while you're plugging stuff in rather than, So it could be, so it is mobile, um, available. And let me, I know I don't have a ton of time here. Your text can actually bring up a page on their mobile phone that shows them real time what steps they have to do.
And when they plug in a cable, it goes green on their phone. Yeah. So at most you could have like your phone and be plugging things in and just checking as you're plugging them in the optic.
Um, let me pop back. So imagine that you're working with smart hands at a data center. You literally take that URL, you put that in the ticket and say, send in your text, open up this URL log in and they will give a set of steps for everything you need to plug in.
And it starts with literally plug in the servers, plug in the switches, plug them in together, it'll walk through building the network, it'll walk through getting the servers up and running and it'll walk through every one of those physical layer connectivity, uh, every one of those physical layer connections. And when everything is green, they actually see that it's green and they can close the ticket. I don't think it'd be handy if you actually flashed the, the port.
That's a, that's a, that's a very good suggestion as well. Yeah, the last thing I was gonna say, 'cause I know I'm over and I don't wanna, you know, make anyone mad. Um, I did mention that the, uh, you know, we have agents running and like, this is an example of one of the vast nodes in the storage cluster.
Um, you'll notice here that we've got, you know, information on connectivity, we're gonna continue to point out like what's connected to what, but we also try to pull in information that's relevant to network people. Because you know, like for myself, I, it is kind of handy to be able to see, you know, what this host sees from a route perspective. So, um, maybe there's something misconfigured on the actual server.
Um, we also, and this is a huge one, is pluggable statistics. As Dan mentioned, there's any issues in the topology. You've got tons of optics, some of 'em are gonna fail.
It's guaranteed. So we proactively track and we also alert on, um, you know, optic statistics so that you can get in here. And if you're troubleshooting a server that seems to have flaky connectivity, you'd easily get in and get your digital optical monitoring data even from the server side.
And so it kind of extends down into the actual, you know, compute end of the spectrum there. So Have you considered going one level up into nickel libraries and making sure that those connect and talk? That's a wonderful idea.
Um, please not yet. Okay. So, um, thank you for this.
Uh, we're, we wanna be c we wanna be conscious of time. Really appreciate the time, really appreciate the feedback. Um, again, uh, if you wanna know more about hyper fabric, we do have the other session or feel free to reach out to any of us and we're happy to dive into any of this.
So really enjoyed talking to you guys. Thank you. Thank you.