Unleash AI Potential with Cisco’s Scalable GPU Server Solutions
Discover Cisco’s advanced AI servers designed to handle demanding workloads with dense GPU capabilities and scalable, future-ready architecture. Built on the latest NVIDIA platforms, these solutions offer exceptional performance and flexibility, supporting up to eight GPUs for AI model training, fine-tuning, and inferencing. Join our session to learn how Cisco’s AI technology empowers enterprises to accelerate innovation and transform data-driven operations across diverse industries.
Presented by Jason McGee, Senior Director, Compute Strategy, and Nishant Shrivastava, Engineering Product Manager, AI Servers. Recorded live at Tech Field Day Extra at Cisco Live EMEA 2025 in Amsterdam, Netherlands on February 11, 2025. Watch the entire presentation at hhttps://techfieldday.com/appearance/cisco-presents-day-1-at-tech-field-day-extra-at-cisco-live-emea-2025/ or visit https://techfieldday.com/event/clemea25/ or https://Cisco.com/ for more information.
Transcript
Last week, I walked into the office of a CIO at a large railway company, former military guy, gray hair, big beard, generally kind of an angry individual. He looked at me and he said, infrastructure for AI is kind of like riding a unicycle and a marathon. And I said, Gary, what are you talking about?
And he said, we've realized that it's a very long race, but we're gonna need to balance and adjust to line of business every single step of the way. I looked at him and I said, it's okay. We got you.
We've got a solution for you. And that's what Nsan and I wanna talk to you about today. So first of all, thank you for being here.
We appreciate the opportunity. Thank you for listening. Questions, comments, we've, it's been been fun to sit in the back of the room and listen today.
Um, keep 'em coming. That's, uh, we, we appreciate that. Um, Nant is our senior engineering product manager for our AI servers.
It's a pleasure to have him with us today. We're, we're very fortunate to do that. We, my name's Jason McGee.
I'm part of our office of the CTO for our computing business unit. And, uh, I think I have just a in incredible job at Cisco get to come in and talk to customers a lot about where we're at from an AI perspective, and there's all sorts of models that are out there. We break it down.
There's really five big areas that we are focused on. What makes Cisco unique is the ability to tie all of those layers together. La, our last tech field days, Siva and Tushar came in and talked about some of our solutions, but the breadth and depth of Cisco is what allows us to take advantage of the high powered AI servers that you heard about during the keynote and the ones that we'll hear about a little later on today.
Those solutions can be broken down into some big buckets. Our Cisco validate designs where we, where we take our best practices, our lab testing, our certifications, performance testing, and we roll that into a user consumable document. They've been around for years.
We do 'em across the company. They're generally several hundred page documents, but the really cool thing is we've made them as code. So they're consumable through get repositories, their customers are able to get those CVDs and deploy that validated configuration or blueprint based upon the knowledge base of Cisco as a whole.
But we looked at that and we said, you know what? We could probably make that easier for our customers. So we took and we put a product wrapper around our Cisco validated designs for ai.
And you may have heard this last fall, we introduced our AI pods. And what AI pods are is they're simply a purchasable Cisco validated design, including the networking, the computing components, the software components, all those layers that we showed on the previous slide. The servers that Nishant and I are gonna talk about today are what powers AI pods use the information from our CVDs.
And if you want that available as a service like Dan just talked about, we have hyper fabric. And so that's that glue, that's the solutions that pulls AI together for us. But it all starts with the fabric and hyper fabric ai, that portfolio is an incredible piece of that.
So we're gonna spend the rest of the time that we have together half an hour or so with some time for some q and a worldly focusing in on that next layer up is what does the compute form factor look like? How do we power the computing engine that's driving the workload, that's generating all the traffic that's populating our networks? How do we get the work done?
When we look at across our, our customers, we've seen the hyperscalers, some of the largest, the large model trainers, really be focused on the ver the large language models. When we get into enterprise customers, we're finding that your mileage is gonna vary as you move across the A landscape. Some of them are doing training, what you see on the far left here all the way over, they're adding their own data into it through retrieve, log, augment a generation.
They're doing light tuning, they're doing retraining, some of these existing models, moving out into the inferencing space. We are seeing massive growth in the inferencing space over the next several years. Uh, and we think that's gonna dwarf the training space.
As the industry moves toward more specialized models, we're seeing the, the large, large models generate into smaller specialized models that are industry specific or use case specific. And that's the what we're building the platforms for to take advantage of, you know, where, where the world is going from that AI continuum. If you look at the compute portfolio that we've been selling for over 15 years, a phenomenal amount of capabilities in the AI space, but it really fits in to that inferencing light training, tuning part of the, of the model.
And as a, as a quick glimpse, if you look at what's there, if you map the GPUs to the, the server platforms that we have available today, there's a lot there. Don't try to consume it, but the takeaway is that we can handle, we think about 60% of the AI workloads that are running an enterprise today with the GPUs that are in our blade servers and our rack mount servers. Um, quick refresher, we call our current modular or blade based systems, Xer.
We have a rack mount line of servers that we call C-series just so that you and I know as you move across the Cisco landscape, we've got all sorts of nomenclature and, and landscape, but the Xer of the are the blades. And that's what you see on the left hand side. The existing one, you kind of, the typical one U2 rack servers are what you see on the right hand side with the variety of GPUs from different vendors that are available in them.
What we're gonna be talking about the rest of the day is gonna fall into that right hand side, and it's gonna fill out some of that continuum that we looked at earlier. We heard loud and clear from our customers, from our partners, even Cisco IT internally, our, well, you know, one of our, one of our biggest customers is that we needed something to fit into that training space. And so we've introduced a large training server, dense GPS GPU server using the high speed fabrics for GP GPU to GPU communication.
So if I'm running the NVIDIA GPUs inside the system, then we use the NVIDIA in NV link fabric for that fast fabric internal GPU to GPU communication. If I put in the A MD GPUs into it, then we use the infinity fabric from a MD and na, Sean's gonna show you what those architectures look like as we get into that. And then very exciting, you see with the, the yellow blob onto it, what G two announced during the keynote today, what we call our C eight forty five a server, is our dense GPU server, but designed for flexible workloads.
A lot of our customers don't need the full 8 85 8 GPUs, high-end GPUs. They're either running an HPC environment today and they kind of wanna crawl, walk, run into the AI space, or, you know, they're getting into some of the smaller language models. They're beginning to get into training and they need something they can stepwise grow into.
What we are hearing repeatedly is that in the enterprise space, a lot of the software teams are not ready to implement ai. They're gathering their data, they're cleaning their data and and, and getting that ready to go. And a lot of times they need to, they don't need to start with eight GPUs.
They can start with two, they can grow to 4, 6, 8 depending on what they need. And so that's the box that we're very excited and introduce today and we'll dig into a little bit, little bit more details around that. No matter what these servers are, if I can't operate them in a simple, easy manner, then it's kind of pointless.
We're gonna, we're gonna make life harder for the end users. And so what we've done with all the compute portfolios is we look at our traditional compute, whether that's the rack servers, whether that's the blade servers, we've managed those servers locally for years. We call that very creatively our unified computing system manager.
Well, all of those servers in those legacy environments take take advantage of that large number of GPUs that we put up earlier. We can also take and plug those existing servers into our, as a service management architecture and what we call intercept, that gives us capabilities to do some of the AIOps functionality for our servers. So we can look at proactive tack, we can dismiss parts before the, the customer even knows that they fail.
Um, we can do over the horizon notifications for our customers, kind of help them see around the corners, uh, whether that's a security advisory, whether that's a vulnerability, whether that's a failing hardware component in the system. Using the metrics and the monitoring that we have on the, on the back end of that system. We're collecting a little over 13,000 data points a second on that platform.
An amazing amount of telemetry that we're using in a, uh, AI ops functionality to, to help out our customers going forward. What we're doing, and we're gonna talk about today is the AI servers that you see in the middle here will also connect those servers to the same operations platform. And the takeaway is that when we move AI into the enterprise space, the governance, the operation, the tooling, the APIs that are already being used in enterprise space can be used for ai.
I don't need a separate silo, I don't need a separate set of stack tooling training. I can use what I have in house today. So the sovereignty, the governance, the best practice that our, IT has put in place we can take advantage of with that existing management operations infrastructure.
So with that, we're, we're gonna dig deep into the 8 85 first and then we'll jump into the, the newer box that was introduced during the keynote today. So naan, I'll turn it over to you. Alright, Uh, so I'm going to go into the details of, uh, UCSC 85 8.
Uh, it's a dense GPU server, uh, which was announced, uh, and late last year. Uh, it was our newest edition to the Cisco UCS product family. Uh, and it's designed for some of the most compute intensive and data intensive, uh, workloads, including, uh, LLM training, fine tuning, uh, deep learning.
Uh, and depending on the size of the model and, uh, number of concurrent users you want to support, you can actually even use it for a r and forensic kind of, uh, use cases. Uh, I mentioned earlier, it's a dense GPU server. It can support, uh, a of NVIDIA's, A GXH 100 or H 200 GPUs, uh, as well as AMD's, uh, M MI 300 x GPUs.
And on the CPU side, it's uh, a dual socket server, uh, with support for, uh, AMD's four generation and fifth generation. Uh, CPUs here is a bit of, uh, technical details. Uh, it's, uh, it's an eight ru, uh, air cooled server, uh, and supporting at the mission earlier, both fourth and fifth generation A MD, uh, epic, uh, uh, uh, CPUs.
But instead of, uh, supporting the entire set of SKUs from these generations, we have picked certain SKUs, uh, of CPUs which are more optimized for AI kind of workloads. So there the focus is more on the, the clock rate. 75 gigahertz.
So those are the CPUs optimized for, uh, AI workloads. It has 24 DDR five, uh, dim slots, uh, which can operate up to, uh, 6,000 mega transfer per second. Uh, standard storage mirror boot drives, uh, 16 NVME for, uh, local data storage.
Could be checkpointing or, uh, any other purpose. Uh, GPUs already mentioned, uh, H 100, H 200 from Edia SXM pump factor and AMD's MI 300 x, which is OAM pump factor, uh, in terms of network cards. So it does have quite a few PCI slots in there.
So these servers are sort of designed to be, uh, placed, deployed in clusters. The compute fabric is extremely important for these. So we have one nick per GPU for that EastWest connection to deploy it in a cluster.
And then we have up to five nicks, which can be used for, uh, the front end network or the north south connectivity. I'll go to more details on that. Uh, what we are supporting there right now is, uh, the Connect X seven or B 31 40 h, the 400 gig nicks or, or super Knicks for the East West connectivity.
And two by 200 gig, uh, uh, Nicks and, uh, dpu, uh, for North South. And then of course there is an OCP card, which can be used for host management. Uh, in terms of cooling, there are quite a few fans.
There are 12 fans up front, and there are four fans inside to Cool the SSTs. And then, uh, power Supply, it's a power hungry server. Uh, there are six, uh, three kilowatt, uh, PSUs operating in four plus two redundancy mode for, uh, the GPUs.
7 kilowatt PSUs for the CPU sled components. I, I have a question regarding this. Yeah, our hungry means also the GPUs will produce more heat, I assume.
Yeah, I see this on the gaming computer from my kids. Yeah, they are always running hot. So what do you see overall on the, let's say, cooling requirements compared to enormous server without all these GPU density?
Is it double or do you have kind of a relation how much more heat they will produce? Uh, So, so when you say in comparison to a regular server, uh, I mean that's like subjective, but these are designed to be kept in regular data centers. It's air cooled and the amount of heat dissipation, uh, in the number of winds, it basically you can put in a regular data center.
Right now, it doesn't require any special cooling. So It is not that, uh, let's say magnitude's higher, it'll still work where let's say systems without GPU have have worked. That would be correct, yes.
Okay. Yeah, Yeah, you don't need any special requirements include these, I mean, there are, uh, newer GPUs coming, uh, on NVD and MDs Roadmap that would probably require, uh, some level of cooling. So, but that, that's coming, that's next, uh, on the roadmap.
Okay. And, and so here is an exploded view of the server as, and as you can, it, it, it's a modular design. On the top left, uh, you can see the GPU U sled, which is four RU.
Right below that is the CPU U sled. And then below that, uh, a layer of, uh, uh, PSUs front, uh, there are 12 fans, uh, and you, but at the bottom you can see there, those are the, uh, drive slots. And I, of course, I'll into more, uh, details on the specifics of these components.
Uh, there's a rear view of that. And then the middle there, there you can see those, uh, eight PCI slots for the EastWest connectivity. A closer look at the rear view of the server.
So as I said earlier, GPU tray on drop, uh, CCP tray in the middle, and then the PSUs. Now taking a closer look at the, look at the GPU tray, uh, it, it, as I mentioned earlier, it has, uh, support for 700 watt H 100 H 200 SXM GPUs from Nvidia and seven 50 Ward, MI threes GPUs from a MD. And if you look at, uh, the, the sort of like a logical blog diagram of the GP board here, this is for nvidia, you can see that there are eight GPUs, which are connected to four NV switches on that board through multiple envy links.
What it allows is actually is non-blocking communication, high speed communication between any pair of GPUs here on the board. And you can achieve up to 900 gigabytes per second by direction bandwidth and to provide connectivity to the outside world, uh, when, again, when you're putting it on a cluster. Uh, these GPUs are connected to an basically sort of like their own, uh, uh, sort of like CX sevens or super nicks via a set of PCI switches, as you can see here.
So that's Nvidia GP board design. Uh, Quick question. Yeah, those super nicks are those blue field three?
These are BF three B 31 40 H. Okay, so this is spectrum X? No, this is, uh, this is just the GP board actually.
Okay. Uh, there is no spectrum or, or any external switch on this. All of the, so up until this part, this actually sits on the GP board, and this is part of the CPU sled.
Uh, the PCI switch and whatever you see above that, that's part of the CPU sled. There is no external component. This is, all of it is on the server itself.
Okay. Are you getting to the Infin band question Eventually, because, well, you knew that's where I was going. Yeah.
I mean, You know, the blue filter three, the Connect X sevens support either or, right? All Of our reference architectures are ethernet based. That's where we see the in industry going.
But recognize we do have some customers, particularly the HPC, that have an INFIN band fabric. So we will support infin Band connectivity on this. It's not our primary use case or where we see the future going, but as a transition period, it's something that is supported with this box today.
Well, I, I mean, I don't wanna speak out of turn, but I think Nvidia realizes that too, which is why they put so much development into Spectrum X recently, is because they realized that InfiniBand basically has a ceiling and feature complete at this point, and they're wanting to get people to migrate off of it as much as possible, because especially if you're gonna start, uh, deploying this in a cloud scenario, clouds are not gonna deploy InfiniBand if they aren't their have it because that's a cost that they don't wanna incur. Exactly. And we looked at a hard stance and we said, let's just make it easy and acknowledge that there's some InfiniBand out or there's a lot of InfiniBand out there.
So we, you know, we'll support if somebody wants to connect it that way and transition over with Nvidia and the rest of the industry. Fantastic. Okay.
Thank you. And on the next slide, I, here I have the A MD, uh, GPU board design. So the, the top half of the, as you can see, is actually pretty similar to the, what you saw on the previous slide, but the bottom, that's the, uh, A-M-D-G-P-U connectivity.
So there are eight, uh, OEM form factor, uh, GPUs from A-M-D-M-I 300 X in the scenario, and they're all connected via this full mesh kind of network. Uh, and this basically ensures 128 gigabytes per second, uh, bi second bandwidth. And this, of course, is based on the infinity fabric from Infinity, uh, from EMD, Uh, Closer look at the CPU tray, uh, uh, support for, as I mentioned earlier, and AMDs, uh, fourth and fifth gen, uh, CPUs.
And we have picked, uh, uh, 95 54 from the Fortune and 95 75 f, uh, CPUs from the fifth gen, uh, uh, from the fifth gen of, uh, a MD cpu. Uh, and then for certain configurations where the customer does not require, uh, as much compute power, uh, or as many cores, uh, we also support 4 95 35. Uh, and then on the dim side, we support three sizes, currently 64 96, uh, 1 28 GB dems.
Depending on, uh, what the requirements are and what GPUs they're using, uh, they can pick, uh, appropriate, uh, dim. Now looking at more of like, um, uh, closer look at at, at the, especially the networking components here at one, you can see those eight PCI slots for, uh, the compute fabric or the backend connectivity. And as I mentioned earlier, we currently support PX sevens and, uh, uh, B 30 and 40 H super mix at number two.
You can see, uh, it's a B 30 to 20 DPU from Nvidia. We also have support for, uh, CX seven if you don't require, uh, DPU functionality or any offloading. And then, uh, but if you want DPU, but then 400 GPS connectivity, then there is support for B 32 forties.
Also at number three, uh, you can see the two PSUs. 7 kilowatt and one plus one redundancy mode. These power, uh, primarily the c pled components as well as, uh, the drives.
Number four is the D-C-S-C-M, which is data center secure control model, and has, basically, it houses A BMC. And then, uh, there is an RJ 45 port for out ofAnd management and the mini display port and a few USP ports, uh, for, uh, KVM connectivity. Uh, just one question we had before the presentation for hyper sheet.
So on two, you could run then EBPF. Yeah. If we have a DPU networking card, that would be supported.
Yeah. Yes. Okay, wonderful.
That's the super neck is the DPU, the branding thing. So we have both super next and dp, dpu. I mean, I think Nvidia uses the term, Yeah, Nvidia prefers the term super neck because they're cool, because in, because a MD calls them dpu, and then Intel calls them IPU for some reason.
Yeah. And then, uh, finally at number five, we have, uh, a couple of RRG 45 ports. That's an Intel card, OCP card, which can be used for host management here.
So quite a few ports in terms of how it, the, the network connectivity, we basically see four, uh, types of networks on these servers. There is of course management network, which is out band management, as I mentioned on the previous slide. There's front end network, which can be used to connect to the, the broader, uh, data center.
It can be used for getting access to certain data for inferencing storage network, of course, to connect to the storage. Uh, and then the most critical one, I think in the scenario would be the, the inter GPU backend network, also known as compute fabric. And as you can see, like eight nicks, uh, uh, which will support this.
And to create a small cluster, all you basically need is one, uh, 64, uh, port, uh, switch. Uh, in the scenario it's, uh, nexus 93 64 D DX two A, and, uh, you can create a really small cluster of just eight servers using it. Uh, but if you want to scale it, let's say a larger cluster of 32 nodes, then of course you need, uh, fine leaf kind of architecture.
Uh, the important thi thing here is that, uh, all of this network has to be, uh, non-blocking, um, non-rated. So basically, if you are connecting a 400 GPS, uh, nick, that, uh, leaf switch connectivity to the spine, switch from the, uh, leaf also has to be 400 GPS or higher. You got support and or gigabit nicks and Not, not, not yet, not in this one yet.
Uh, as we move to the future generations of this, we will have support for it power and cooling. Uh, again, it requires a lot of power to power all those GPUs and, uh, CPU and everything. So, uh, it has six PSUs primarily for, uh, the GPUs at the bottom there.
You can see those are three, uh, thousand tt, uh, 80 plus ccp, uh, PSUs running in four plus two redundancy mode. And let's say if, uh, uh, three of them actually, uh, become inactive, then of course the server still operates. But the GPU performance will be capped at, uh, 60%, uh, across all of them.
5 kilowatts per server. That is the maximum possible use power function that we, uh, expect. So that's basically primarily used for sort of like a provisioning purpose.
But in normal scenarios, I don't think you will see more of like more than eight to nine kilowatt. Uh, do you have like a recommended density that you're planning for in a rack so that you don't melt it through the floor? Uh, It depends on basically what power supply, uh, customer has.
But there is a customer that's trying to put four of these in one rack. So I think they have, uh, 60 kilowatt power supply in that rack. And so they will put, uh, four of these, so around 48 to 50 kilowatt will be for, for these servers.
And then there is top of the rack switch and other things in there Currently air cooled. Correct. It's All air cooled Plans to make it water cooled.
Uh, That will be a future platform. Okay. The only reason I ask is Nvidia has wells are water cool because they're hot?
Yeah, yeah. Is the storage, um, accessible from all the server? Um, so server one can access the storage from server two as well.
So it's Not, not the local storage, uh, if that's the question. Mm-hmm. Uh, that's all local, but yeah.
Uh, you can have a distributor storage, uh, software storage, uh, connected to these, but that's a different architecture, of course. Uh, there are, Okay. So do you do that yourself or do you use other, uh, solutions For We are looking at certain partners, uh, to support it.
I can't think of one. I think you could. All Right.
So hardware configurations. So these are, uh, what we are doing here is we have created these sort of like optimized fixed configurations for these servers. So, uh, let's say if you're doing, uh, training, uh, and then you want to deploy it in a cloud-like environment, then you will have certain specific requirements you might want, uh, B 31 40 H as well as a B 30 20.
But if you're doing inferencing, you might not even require those EastWest next, right? So for, for, for different deployment scenarios, we have created these fixed configurations, which are optimized for, again, as I said, like those specific scenarios. So we have about 16 fixed configurations here, uh, using H 200.
And then there is for H 100, fewer for H 100, of course, because, you know, we have H hundred now. And then, uh, uh, for MI 300, about five or six of those. And, uh, this is also used of course, in hyper fabric cluster, uh, for ai, uh, as we mentioned, again, so I think the key differentiator for the servers are these three things.
It, it delivers unmatched performance. It has this high speed GP track connect on the, on the server itself. The board, uh, that has NV link and we switches and scalability, uh, those eight nicks, which allow it to be, become part of a really large cluster.
And that helps us deliver, uh, serve basically all these use cases, uh, ai, uh, starting from gen AI training model training, all the way to deep learning, reinforcement learning, and even some large model inferencing, uh, HPC, quite a few of our customers actually deploying these in HPC uh, uh, environments. And then you can also, uh, do some realtime data processing, which is through, uh, you know, by accelerating through, uh, GPUs. Uh, so again, with Jason, as mentioned earlier, not every customer requires such a powerful server.
Uh, for some of them they just starting, uh, on their AI journey. They require fewer GPUs, uh, smaller density. So for that, we announced, uh, C 88 4 5 A today, and Jason share more detailss on that.
Wonderful. Thanks Dhan. Thanks.
Appreciate the question. Anything else on the 8 85, the big box before we move on? Okay.
It, it's big, it's powerful. Uh, it, it'll, it, it will do a tremendous amount, Just my head around 60 kilowatts in a single wreck like that could be, or just four servers, hazardous environment. It's a high voltage coming through.
Don't, don't look up tech specs too closely. 'cause Nvidia can get 60 kilowatts and a half rack. Yeah.
And even, even that, I mean, yeah. And one of those network diagrams was 32 of those servers put together. Um, I think, you know, one of the big takeaways, you know, I started with solutions and I, and I really want to emphasize that, that the validated design on how you put it together, the AI pod where I can order is one, or I think the demo for hyper fabric is incredible because Dan started out with, you know, user guidance, what cable to plug into what port.
If I'm paying that amount of money for those high-end GPUs, I wanna maximize performance. Um, I mentioned Gary, the CIO that I talked to last week, we had a, a, another, uh, CIO that mentioned to to us. He said, he said, my job is not to build the AI applications.
It's not to transform my company. My job is to keep the GPUs as busy as possible to fill them with data. He's like, that's where our investment's at, and I need to ensure that I'm doing that.
And that's where those guided configurations, the hyper fabric AI can come in and really benefit around that. Okay. But the conversation we're having is a lot of customers are not ready for that massive, massive box.
And, you know, we really think the 8 45 is gonna fit a wonderful niche architecture that's out there. And I say niche, that's a, I shouldn't use that term, but being able to start with two GPUs grow in increments of 2, 2, 4, 6, 8, uh, of that and have a variety of GPUs. I can use the H 100, the H 200, kind of the top end from NVIDIA today, or I can use, uh, the L 40 GPUs for, um, that are, that are just fine for some of the workloads that are needed out there.
When you get an outreach for something like this, does that create like a professional services engagement so that you can talk the customer through that kind of discussion? Or is this something that you're really only contracting through other partners that have experience doing this deployment? Both.
So yeah, we're doing a lot of that. When I mentioned the AI pod that comes with the services wrapper around it, that that's a piece of that and That much money better. It's kind of interesting to know, especially 'cause of everything that's going on in, in, in Europe and, um, probably, uh, others will follow as well soon with ESG and everything that's going on.
Yes. Making sure that you are going over the regulations that going on with ai. Exactly.
Yeah. It's, and even as we get into the switches, the optic density, you know, it's, it's gonna go beyond just the GPU, the server. Our entire data centers are, you know, are being affected.
Um, and we hear it every day, but when you actually see it on paper with the specs, it's, it's mind boggling. Um, you know, what, what we can do or, you know, what's out there and the amount of work that can be done from, from this, um, with the, I'm gonna have a hard time not using the code name for this box, but the eight four, since it's just brand new today, but the, the 8 45 box, it is a a a two C-P-U-A-M-D epic based system, uh, and I said it can, it can grow, grow in the GPUs. Where I do want to emphasize on this is that we're using the 8 85 we talked about before, that's an HGX reference architecture from nvidia.
This box is the MGX reference architecture from nvidia. Um, we looked at that, um, and there's some things with the off the shelf MGX specs that most of the MGX servers in the world, um, are doing that we thought we could do better. There's some simple cable management, cable routing items that we improved that drastically reduced the amount of airflow, improved serviceability of, of the box.
Uh, one of the things that we did was that most of the time on MGX architecture, all the power supplies are on one side of the machine that prevents the box from passing a drop test because it, you know, goes on one side. We evenly distribute the power supplies, uh, horizontally across the bottom of the box. So if we have to do something like a drop test for reliability, again, those things that are just standard for our enterprise compute customers that they're, you know, they're used to using it and have to have, you know, in that.
Um, one of the other things is that, uh, typically it takes about 38 to 40 screws to remove the GPUs from an MGX design, this box one screw. So if I need to replace A GPU, we've made serviceability easy on it, uh, for the customer. So, you know, really trying to bake in some of those things that our Cisco customers are used to using in, in those systems.
You are only supporting the A MD CPUs. You are not going to support the intel CPUs. Th this box is a MD CPUs only.
Um, we will support the MI series of GPUs in it from, uh, you know, and then, you know, going forward I think we'll see a, a broader variety of the vendor CPUs and the AI specific servers. But just point in time where we're at today, um, it's what makes the most sense for, you know, for this box. Um, the 8 85 that we looked at that is more of an appliance model where you buy a fixed config, this box, you have a, have a lot more flexibility, uh, in, in how you can, you know, buy it out as far as the, the a MD CPUs, the storage, uh, GPUs, and then this, the super nicks and the, uh, north south nicks that we have, uh, from Nvidia as well that connect X sevens.
So we are, I said, announce the box today, we anticipate April timeframe is when we'll begin shipping this box out to our servers or to our customers, but taking orders today, and we do have it on the show floor, uh, if anybody wants to take a walk down and go take a look at it. Um, key thing is it is obviously optimized for ai. We have the capability to do some incredible workloads on it, but it does a lot of other things as well.
So if I need to scale up into AI, or I need to have a dual function server, this Fox has the capability to do that. And I can't stress enough the importance of the consistent management. So it plugs into the, as a service based management platform that we have, which you can run OnPrem or can be consumed from the cloud of Cisco, whichever is easiest for the, for the customer, but that inter site platform allows that consistent management of it.
So it fits in, into the overall, uh, what, what, What do you mean with a dual purpose server? So if, if I am, maybe I have a traditional large database workload, HPC environment, um, I also have the ability to run AI workloads on that box, I could transform over the lifecycle of that box, you know, or I could, you know, actually have multiple applications running on if I needed to. Yeah, I would, I would see that happening for, um, PDI environments, hospitals doing right.
Uh, a lot of these things. So, um, yeah, That's at our, at our recent customer advisory board last fall, we heard that loud and clear. We, we kind of came all in on, hey, here's this eight GPU server, and they were kinda like, hold off.
Most of us are not quite ready for it yet. Um, we need something that, and we have a perception because AI is new and sexy. Everybody's talking about it, that it's a thing by itself, but there's never an AI system or an AI application that doesn't have a lot of other traditional servers wrapped around it.
So if, you know, especially if I'm putting my, if I have a off the shelf model that I'm plugging my own data into with retrieval augmented generation, I've got a vector day sitting out there that's on a server somewhere. I've got traditional storage, probably coupled along with the, um, you know, vast or ddns or CCAs or others of the world that are, that are wrapping that AI specific storage around. So it's, it's an ecosystem.
And so that's where that dual purpose comes in, Especially the university, uh, hospitals doing all the hospital stuff during the day, during the research stuff. Research doing the doing. Yeah.
So yes. Pretty, Yeah. Or just normal old fashioned machine learning.
Yeah, Absolutely. Right. We don't see it in practice a lot, but one of the unique things about the unified computing system, and I'm sure you're all, you know, very familiar with it, but we have the concept of a service profile.
So when, when we deploy the, the blade servers, our traditional rack servers out, we go in and we carve out what that server looks like in software, how do I want it to be configured? What buy do I want on that box? Uh, and in our blade based system, we've had some customers actually convert that profile day and night where I can run an application during the day and research during the night, something like that.
And so that's something that, you know, we'll be able to do in the future on, on these boxes as well. Um, kind of a corner case, but it, but it's out there. But I think, you know, the HPC space is a big place.
Some of our traditional database, uh, our traditional workloads, very large databases, those type environments, uh, it's gonna be a be a really good fit and we're gonna see our customers grow into it. And these will be part of the AI p validated design. So they're not part of the AI pods today, but that is coming when, So we're not gonna start shipping these until April.
And so, you know, we anticipate the validated designs, the AI pods and all of that will be available around that same timeframe. So, you know, we're working to get all that together. Okay.
Question? No. Okay.
I saw the mic. I was, um, as we can just kind of wrap up and look at these layers. So we talked about the Cisco fabric, hyper fabric.
We didn't go a lot into storage that's there, but we know that's a very big piece. We feel that we've got a full portfolio at that compute layer, whether it is running inferencing on the CPU itself. We have a radiology demo running downstairs in the Intel booth on doing, uh, AI inferencing, looking at pneumonia images with just the Intel CPU no GPU, all the way up to the large eight GPU systems that we have.
And, you know, with that, uh, C 8 45, you know, fitting in here with the, uh, you know, two all the way up to eight GPU systems that are there. Um, and then being able to wrap the solutions, the, the validated designs, the AI pods, the hyper fabric around that. So the Nvidia, N-V-A-I-E, the PyTorch framework, whatever the customer needs to run these top layers, comes along with that all in one package.
Okay. So it is a massive amount of complexity. Um, but it's an exciting place.
It's a really, really neat place to be, and we're excited to have the portfolio built out. Okay.