Broadcom Tomahawk 6 Scaling AI Networks with the World’s First 102.4 Tbps Ethernet Switch
Tomahawk 6 is the world’s first 102.4 Tbps Ethernet switch, designed to meet the demands of massive AI infrastructure. In this session, we’ll show how it enables both scale-up and scale-out networking, with unmatched bandwidth, energy efficiency, and congestion control. We’ll also debunk common myths about AI networking and explain why Ethernet has become the fabric of choice for the world’s largest GPU clusters.
Pete DelVecchio introduced Broadcom’s Tomahawk 6, a 102.4 Tbps Ethernet switch designed for both scale-up and scale-out AI networking. Tomahawk 6 doubles the bandwidth and SERDES speed of its predecessor, Tomahawk 5, and incorporates features to enhance load balancing and congestion control. DelVecchio emphasized that Ethernet has become the dominant choice for AI scale-out networks and is gaining traction in scale-up environments due to its open ecosystem and the ability to partition large clusters for different customers.
Tomahawk 6 comes in two versions: one with 512 lanes of 200G SERDES and another with 1,024 lanes of 100G SERDES. The presenter highlighted that Tomahawk 6 is built on a multi-die implementation, with a central core for packet processing and chiplets for I/O. The chip is designed for both scale up and scale out applications. For scale-up applications, Tomahawk 6 can support 512 XPUs in a single-hop network.
The presentation also touched upon power efficiency, emphasizing that Tomahawk 6 enables two-tier network designs, which significantly reduce the number of optics required compared to three-tier networks, leading to lower power consumption and reduced latency. DelVecchio also discussed advanced features like cognitive routing with global load balancing, telemetry, and diagnostics for proactive link management. Broadcom emphasized that Tomahawk 6 is an open, interoperable solution that works with any endpoint and offers flexibility in telemetry, load balancing, and congestion control.
Presented by Pete Del Vecchio, Data Center Switch Product Management, Broadcom. Recorded live on September 10, 2025, at AI Infrastructure Field Day 3 in Santa Clara, California. Watch the entire presentation at https://techfieldday.com/appearance/broadcom-presents-at-ai-infrastructure-field-day-3/or visit https://www.broadcom.com/products/ethernet-connectivity/switching/strataxgs/bcm78910-series or https://techfieldday.com/event/aiifd3/ for more information.
Transcript
My name is Pete Delvecchio. I'm the product manager for the Tomahawk Witch family at Broadcom. And we'll have three presentations today on three chips we announced over the last few months.
I'll start us off with Tomahawk six. Uh, Robin Grimley will cover Tomahawk Ultra and Henry Wu will take us home with Jericho four. So to first set the context, make sure everyone is, you know, familiar with all the terminology we'll be using.
Um, just wanna make sure people are familiar with what's known as like scale up, scale out, and sometimes being referred to also as scale across networking. So scale up networking today is usually within one rack. I mean, it used to be like one box with like, you know, four to eight GPUs.
Today it scales to one rack. We see in the future that that scale up domain will be increasing significantly. We see customers asking for scale up connectivity up to maybe a thousand or more xps.
But current state of the art is, you know, currently one rack scale out would be connecting those scale up racks together. So you could have scale out within the data center, so across racks. And then what's referred to as scale out or scale across between different data centers.
So we have custom chips that are designed for these specific applications. So the way these different chips that we announced over the last few months fit in is that we have OCK Ultra. Now this focuses on high performance computing and AI scale up Ock six that I'll be going into on depth focuses on AI scale up and also scale out.
And then Jericho four would be bridging and scale out and scale across. One of the reasons that these, you know, span different domains is that you see that, you know, it's not one size fit all that different hyperscalers, different types of xus, different workloads, they have different architectures, you know, different communication patterns. And also AI training and Inference has different needs for the network.
And because of that we do allow flexible deployments. We have say, you know, Tomahawk Ultra, so on the bottom left there you see Tomahawk Ultra and Tomahawk six could be used for scale up once you get to scale out. As mentioned before, that could be Tomahawk six, or depending on the endpoints, the traffic patterns, it could be a Jericho four that you see on the right hand side.
And then Jericho four would be something that this will allow the connectivity between data centers and potentially between regions for larger scale out clusters. Oh, I mean, Jericho four doesn't have the bandwidth of the Tomahawk six. Mm-hmm.
So you could see that you, you wanna go back to the prior slide? I'm sorry. Sure.
That, um, the Tom Hawk six would be piped into the Jericho four. I mean, I'm trying to understand how it would work. Yeah, so normally what you'd have is within a building, you know, generally you want to have, you know, a wide and flat cluster.
So they, and you know, most cus companies, most customers don't want to go to go to more than two tiers in the network within a building. So you'd have that two tier network within buildings and then you'd have something connecting those two tier clusters together using a Jericho four. Okay.
Yeah, if you do, and I'll get into this some more on OX six, but there are a lot of benefits in keeping that cluster as wide and flat as possible. Mm-hmm. Um, if you go to multiple tiers, it really makes the congestion control the load balancing a lot more difficult.
And you also wind up with a lot more switches, a lot more connectivity between the tiers. I did wanna mention before getting into Tom six that we do have with Broadcom a full stack of, you know, AI products. So, you know, provides switches.
So Tom Ultra, Tom six Jericho, we add the endpoint. So we have Nicks, um, you know, Broadcom is well known that we do a lot of custom xus. So for our customers that are doing custom xus, we can provide Nick chips.
So they would focus on the processing elements. We could do all the networking. They could instantiate either as some IP on their chip or as a full chip.
And then if a customer wants to roll their own ip, their own ethernet connectivity, we provide a framework specifications for how you can use ethernet efficiently for scale up. Uh, we have physical layer products. So we have re timers, we have DSPs for optical modules, and we have optics.
So we have co packaged optics, we have multimode and single mode optics. Now on the CPO side, uh, you know, we've announced that for Tomahawk six, this will have a CPO version. We've actually been doing co packaged optics since Tomahawk four.
And I think the first thing I'll pass around is for show and tell here. So this, I'd be interested what, like, you know, something with CPO looks like. So this is a TOMA five with co packaged optics.
Um, so this is shipping to multiple hyperscalers. There'll be some very interesting results that some hyperscalers will present at ECOG later this month. But this is basically the TOMA five, the same dye that we use for the electrically connected version with co packaged optics included on the same package.
I'll pass that around in case anyone wants to take a look at that. I wanna go back. You said XPU, I didn't catch that.
Yeah, so basically, and so GPU is kind of a relic of people using, you know, a graphics processing unit for AI training or inference. So XPU is just kind of a generic term saying it's not like Star. Yeah, it's kind of like a star.
Yeah. PU Exactly. Yeah, exactly.
F-P-U-N-P-U-L-P-U-G-P-U-C-P-U, Uh, PPU Star. Mm-hmm. Not about star.
Thanks. So Wild card P. Okay.
Gotcha. Okay. So, uh, getting into Tomak six.
So this is a following Tomak five. Um, TOMAK five is very widely deployed in hyperscale scale out networks. This is actually the chip that really drove the transition from InfiniBand to ethernet or AI scale out.
If you went back a few years, what you saw the, there's a debate about, you know, that for the scale out networks, do you use InfiniBand or is it ethernet? It's pretty much a given that for the general data center networks, the wider network is gonna be ethernet. But the narrative a few years back is that to get the highest performance, it's gotta be InfiniBand.
Right? What you've seen over the last couple years is that's completely transitioned. I think there's one hyperscaler that's announced, they have large clusters within InfiniBand.
Everything else is on ethernet at this point. So that debate about scale out is over this, you know, it's clearly ethernet at this point. Um, as we'll get into, there's a debate about, okay, what do you do for scale up that last top, that last part of the network.
You know, we also feel that ethernet is the appropriate choice there. So, you know, teak five has proven itself in the largest GPU clusters. 2 T chip on the market.
So it's really the most efficient as far as power, as far as the latency, um, you know, fixed performance. And, you know, this is widely deployed throughout the world. For GPU clusters, teak six steps us up in every dimension.
So, you know, we have double the bandwidth, you know, double the 30 speed. Uh, we also add a lot of features that will improve the efficiency of these networks, reflecting the different endpoints and the different workloads you have for ai. So TOMA six is not just about speeds and it's doubling the bandwidth.
It's also about adding a lot of efficiency for load balancing congestion control. Okay. So getting into, uh, scale up and scale up, you know, as I talked I mentioned before, there is this discussion about, you know, what do you use for scale up networking?
We're seeing a lot of traction for ethernet. Uh, a couple of the reasons is that you have, it's a, you know, it's an open ecosystem. You know, a lot of suppliers for ethernet.
One thing that's probably not obvious too is that a lot of these hyperscalers, they have these very large clusters that they can partition between different customers. So, you know, if they have a hundred thousand GPUs, one customer coming to them is probably not gonna rent that whole cluster. So they might want to partition that differently.
So they have ethernet for a common interface between scale up and scale out. They can actually partition that cluster differently for different workloads and really what a customer wants to rent from them. Another big benefit is using ethernet.
Here you have a common technology stack. So the same debug tools, the same networking tools, the same optics, you know, DAC connectivity, you can use that for scale up and also scale out. On an earlier slide, you mentioned Tom Hawk six was also used for HPC.
Where do you see that? Um, so that would be Tomahawk Ultra, but Tomahawk Ultra. Um, so Robin will get into that.
But yeah, the in initial design point for TOMA Ultra was HPC. And then we saw later on that scale up connectivity was becoming more important. We actually added some features to that.
So it's appropriate for HPC and scale up, but Robin will go into that in in depth That sort of a different, um, using that, you know, networking profile for HPC. Okay. Yeah.
Yeah, yeah, Definitely. Yeah. And Robin will get into like some of the exact characteristics.
What's different from, say, scale up or general scale out or data center networks versus HPC? Where does ultra ethernet fit in all this? Um, well actually it's right here.
So I mean, ultra ethernet was formed about a little bit over three years ago now. Um, and the purview of ultra ethernet is, you know, to advance ethernet for all of HPC and all of AI networking. 0 spec, you know, is known to focus on scale out, but there are some technologies there that are also applicable to scale up.
So this, this kinda gets into this slide here that, you know, this is the way we see the ethernet, you know, scale up landscape is that a lot of folks want to use the ethernet phi, they wanna use the ethernet data link layer, but they wanna add some innovation above that. So we had earlier this year announced through, so scale up ethernet. And that is one thing that you could have as like an upper layer or like a complete framework for using ethernet for scale up.
You know, we see a lot of hyperscalers, a lot of customers want to be able to innovate at that transaction layer or that, you know, that, um, you know, the up the layers above the data link layer. So we provide Sue as an example. Uh, we do see that, I mean, it's been announced that some companies are looking at do and say UA link over ethernet and some other people want to have their own custom accelerator implementations over ethernet.
0 spec. And those can also be used for scale up. So all of our switches are ultra ethernet compliant.
And then depending on what features you wanna leverage from that spec, um, you can pull those out or just, you know, use that for your network. You don't have to use all the features, for example, L-L-R-C-V-S-C or optional depending on what sort of congestion control what, how you wanna manage the, the link layer. I thought alter ethernet had some application specific functionality as well.
I mean, you, if you were doing AI training versus mm-hmm. You know, uh, scientific analysis of, Well, there three different profiles in T ethernet. There's an HPC profile, there's an AI base and an AI full.
So yeah, depending on if it's HPC, there's gonna be a subset of the functionality. Um, is That at the data link level? Uh, only or is Or is it?
No, it's, I mean, the thing is, so ultra E Ethan ISS actually kind of a full stack. I mean, it's, it's pretty expansive. So it's got all the way from, you know, link layer.
It's got file layer, link layer, it's got transport, it's got software. Um, so there's some, you know, there are like four horizontal working groups. We got five layer, link layer transport and software.
Then we have some vertical working groups that try to like integrate those and show for certain applications like storage, you know, we'll say storage. This is how you'd apply that drawing from the different layers. Okay.
So yeah, ultra ethernet, I mean, it, it is pretty important. Um, you know, we were one of the founders of ultra ethernet and you know, it's, we got 130 plus member companies, so like 4,000 contributors right now. So it really is kind of a driving force for, for ai.
Yeah. And the switch layer, there aren't a whole lot of changes that you need to add things to your switches to make sure that they're not going to block ultra ethernet protocols. Mm-hmm.
But add, but add a switch layer at link layer, there isn't a whole lot of changes to ultra ethernet. Mm-hmm. Yeah.
And a lot of it's optional too. So basically, you know, there are a few features within ultra ethernet. So for example, it's gotta be ethernet, um, you know, the way you do ECN marking, there's, you know, just, there are a few, a handful of things if you say that if you're ultra ethernet compliant, here are the things you need to have.
But then there are a bunch of optional features depending on is it HPC, what type of AI workload you have. Right. So Robin will get into this.
So we'll get into some more about, um, you know, just this framework for ethernet to scale up and also talk some more about Sue or scale up ethernet. And then, uh, just while you have a pause here. Sure.
Going back to this chip, we've been passing around uhhuh, what is this? So that is Tomahawk five Bailey. Okay.
So Tomahawk five. Um, with that five, five nanometer generation of products that we started, we made the SER really flexible that you can connect it directly to dac. So you know, pass copper, you could have chipped modules, you could have retime optics and it can drive co packaged optics directly.
Yeah. So that silicon is in the middle is the exact silicon that we, you know, shipped originally for AU five for the electrically connected version. Mm-hmm.
And then we put optical engines around the periphery. Yeah. So this is basically, you know, all the optics aside from the la so the laser is gonna be external.
2 T system. Yeah. Cool.
Is the switch in my hand on a Yeah. Switch with the optics Switch and the chip? Mm-hmm.
And Bailey is a BA. How's that? B-A-L-I-L-E-Y.
Is that for somebody's name or is that a cartoon? It's point. It's a point.
It's A-B-A-I-L-L-Y. Okay. Of course.
Yeah. Yeah. So what you'll see is that, um, so there was a, you know, so each of the generations, so you see for example, this TH five, uh, Bailey will have a TAMO six Davidson, and Davidson is an optical engine.
Um, it's so named because we have these optical engines and they can be applied to multiple products. So th five Bailey has the Bailey optical engine on ITT six Davison will have the Davidson optical engine, but that same optical engine could be used by, for example, if someone has an XPU they're doing with our custom asics division, they can have that same optic embedded on their chip. Okay.
Uh, let's see. Males speed this up a little bit. Interest of time.
So just um, on the scale up side, um, you know, the way, one way to look at this is that scale up is, in general, you can look at it as like memory sharing. That you're gonna be taking, you know, some of the bandwidth on A an X today, you know, next generation they might have up to like eight HBM stacks. So up to a hundred terabytes per second of bandwidth to the hbms.
So really between eight, you know, X Ps you're doing some sharing, you're doing some like, you know, memory, coherency memory transactions. So you could look at what you have for scale of bandwidth being a portion of that bandwidth. So today what you'll see is that companies are talking about maybe 10 terabytes per second for one XPU or the scale of bandwidth.
Mm-hmm. There's one hyperscaler at OFC earlier this year that said in quote, the next few years per XPU, they wanna see 50 to a hundred per second of scale of bandwidth. So 50 to a hundred terabytes coming outta one shift just for the scale of interface.
So, you know, a hundred terabits per second for a TIMELOCK six might seem like a lot, but you know, we could throw an infinite amount of bandwidth at AI and it would be consumed immediately. Mm-hmm. Um, and getting to that point, you know, where you see it today is that, you know, xus, you might have like 10 K XUS per cluster.
It's kind of well known with people are getting to like that a hundred K level, moving to a million plus XUS in one scale out cluster. If you look at the bandwidth scale up today is about 10 terabytes. Now as discussed, that's moving to 50 to a hundred terabytes per second per XDU.
6 plus. So if you look at where things are going, we're talking about, you know, a hundred million terabits per second total bandwidth within an AI cluster. So as mentioned before, you know, one chip a hundred terabits seems like a lot, but you know, we're just seeing that bandwidth being consumed immediately.
I mean, literally if we had an infinite amount of bandwidth it would be consumed. Okay. So getting into haw six.
So as mentioned before, world's first 102 tur per second switch tipp, this is double the bandwidth of any other ethernet switch. Tipp, um, can be used for scale up, scale out, scale out out to more than a hundred million X vs. Uh, we do have two versions of TOMA six.
So this is one thing I think surprised a lot of folks, is that people were expecting that we would go to a 200 gig ER with five 12 lanes. Mm-hmm. Um, TOMA four was the first switch chip on the market.
That kinda broke the barrier going from 2 56 to five 12 lanes. And AK six, we have a version that's got a thousand lanes. So this is, I'll pass this around.
This is a AK six with five 12 by two gig er, and this is a AK six with a hundred gig er. So this size package, this size chip has basically never been seen before. Mm-hmm.
Because this is the first chip on the market that has more than this five 12 interface. There's actually 1,020 400 gig thirties coming outta this chip. Uh, what you'll see is that it's a multi chip implementation or multi D implementation.
So the central core has all the packet processing, it has all the traffic management, and there are triplets around it that provide the different IO either a hundred gig or 200 gig. So what's interesting about that is that we have some customers that today with their xus with their optics, they're currently running at a hundred gig PAM four. Mm-hmm.
They know they're gonna upgrade to 200 gig PAM four at some point in the future. So it's actually, it's exact same software, the exact same SDK that you would use for these two chips. And it's literally just a config file change.
If you've developed your software for OX six with a hundred gig, you wanna upgrade your network to 200 gig. Really, it's literally, it's the exact same SDK that also applies For the CPO, the co packaged optics version that it is the exact same silicon and it's literally just a config file change to say, now I want to use this switch with CPO, this is how I set up my thirties for that type of connectivity. This is the first, I've seen two words on this slide that you have instead, training and inference.
Mm-hmm. The, the, the, um, data traffic patterns mm-hmm. Uh, vary enormously Yeah.
Mm-hmm. In AI throughout the lifecycle. Mm-hmm.
Uh, training is the whole different, I mean, I really think of that as probably more of a scale up problem than a scale out problem, but also, um, the, the, you know, read versus write and inverse out and all of that sort of thing. Just extremely different. Yeah.
And then inference because of the tremendous amount of experimentation going on right now mm-hmm. Um, particularly in public sector and, and, and, uh, large enterprise. Yeah.
Um, you know, are you targeting certain customer types, certain use cases here for these? 'cause I, I think that, um, as impressive as this is mm-hmm. It, it's gonna have a relatively, um, while, while a large total addressable market mm-hmm.
A, a more limited number of, of scenarios where it's really needed, Um, I'm not sure. I mean, what we see for Timelock six is that, um, you know, it bridges scale up and scale out. Mm-hmm.
You know, inference would normally focus on scale up. Yeah. But you do see a lot of instances even with, you know, inference would Be scale up, I would think the Other way.
Well, I mean, 'cause inference is gonna tend to be, you know, a smaller number of GPUs or xps. Mm-hmm. Um, so I mean the, the, the normal thought was that, you know, you could basically do inference on like small number of GPUs maybe within one, you know, one rack.
Um, we are seeing a lot of cases, however, where that scale up is actually expanded across multiple racks. Mm-hmm. Especially for like mixture of experts where they're gonna have experts.
It could be a large number of experts and distributed amongst multiple racks. But I say inference is no longer like, you know, four GPUs or eight GPUs or one rack. I mean there are a lot of cases where we're seeing multiple racks being involved in inference.
So those inference workloads do vary fairly dramatically. And that's one of the reasons that we have, you know, both Tomahawk six. So if you wanted to have, you know, we'll get into, we can connect five 12 xbs in a single hop for a scale up network with Tomahawk six.
Um, Tomahawk six also has 200 Giger. So if you had a scale up network that needed either 200 Giger or that fivefold rate X Tomahawk six fits into it very well, it doesn't have quite the, the low latency of Tomahawk Ultra. So if you're doing inference and that low latency is really important, then Timeout Ultra would probably be the better choice for That.
And for training, I would think Ultra would be a better choice because, uh, um, you can really extend training cycles out. Like, you know, uh, a, a good deal if you have a latency issue. Well, I think with, um, so with the training, what happens normally is the data transfers tend to be a lot larger.
Mm-hmm. So you don't have as many transactions. The packet sizes are larger, so that flow through latency.
I mean, it's still fairly, fairly short latency, it's a pipeline device. Um, it's not, you know, it's quite what we have to ultra, but the training, first off, you tend to have very large clusters. So you need to have high rate X, you need to have, you know, multiple tiers normally.
Um, and you know, the, to Ultra wouldn't have the features to support that. 'cause we really wanted to focus on this single pops kind of, yeah. Scale up connectivity.
All right. But also you do find that, you know, because especially on the training side, you know, instead of like these memory transactions that you'll normally have on the scale up network, it would be, you know, large packets, like very large transfers be, you know, megabytes of transfers at a time. And so that instantaneous latency through the device is a little bit less important.
Yeah. Can you talk a little bit about the power, uh, efficiency? I mean, it seems like a lot of these large clusters nowadays, uh, GPUs obviously are consuming a lot of power.
Mm-hmm. But optics and networking and all that hitting. Yeah.
Okay. So let me hop to one side because I've got a lot of content here. I think I've got eight minutes left.
But let me to, to address that point. What I'll do is hop to one point here that addresses that. Mm-hmm.
Um, this is one of the, so Tom six we discussed. It can handle, you know, it's, it's for like scale up and scale out on the scale outside. One of the big benefits you're gonna get here is that, you know, it's very common these days.
People would have a 200 gig ethernet connect, you know, connection per XPU. Um, if you wanted to have more bandwidth, you would have multiple rails in parallel. But what you're trying to do is, you know, have, as you know, wide and flat of a network as possible.
There are a lot of issues going to a three tier network, and I'll get to that on the next slide. But this is one thing. So OX six, we can connect 128 K, so actually a hundred, one 31,000, you know, XUS in a two-tier network.
And that two tier versus three tier might not seem like a big, big deal. But this is where, what it boils down to is that if you have a two tier versus a three tier, three tier will have two thirds more optics. So that immediately gets through the power efficiency.
That if you look at the networks today for the AI clusters, up to 70% of the networking power can be just in the optic. So the fact they have two thirds more optics with a three tier, you know, means you're gonna wind up with about two x the power for the network. Um, you'll have a lower, you know, higher latency if you have a three tier, because now you have potentially five hops through the network instead of three hops higher reliability.
So you know, much fewer switches, you know, less optics. And then one thing that's probably not immediately obvious is that you do get significantly higher performance. Um, and I think there's actually, I'm trying to recall, there's a presentation pretty recently by, uh, Microsoft where, which conference it was within the last month or so.
But you know, what you see is Microsoft a lot of a hyperscalers, a lot of customers are adamant that they don't want to go to more than two tiers in the networking because it makes the congestion control, it makes the load balancing a lot more difficult. So, you know, going to two tier versus three tier is kind of a big deal both as far as power efficiency, as far as network efficiency. Mm-hmm.
So, uh, do you have power specs on each one of you? Um, we do under NDA. I mean, what I could tell you, it's like much less than one watt per gigabit per second.
So, um, but you know, basically, um, you know, we, That that's a lot of watts. Mm-hmm. Yeah.
But it's a lot of bandwidth too. Okay. And, okay.
And, uh, do these have the capability to drive, uh, either passive variety, uh, copper rather than optic? Mm-hmm. Yeah, so basically the, the, so the 30, so the same thing we did with Timelock five, we did with time haw six, all the Ss are designed to drive, you know, passive copper to have, you know, chip to module for retime optics to drive, you know, event like optics like LPO or LRO and also CPO.
But yeah, same, exact same as security can be used for That. Okay. For the scale up applications, do you see people using lots of copper rather than optic?
Um, not so much on the scale up. Scale out, certainly on the scale up, you know, it's very common these days. I meant if I said scale out, I, I meant scale up.
Yeah. Within Iraq. Yeah, Within Iraq it's all about, you know, passive copper connectivity.
Okay. So that's actually, I think one big benefit of working with Broadcom is like, we're pretty well known to have, you know, the best thirties on the, on the planet. Now we could say is that for, I think I actually have a slide on this.
So we just talked about the type of copper and type of connectivity. So as we discussed, we got, you know, D connectivity, pluggable optics, co packaged optics, um, 45 plus DB at 200 gig. So insertion loss, um, TOMA five generation.
We did talk about the fact that we can support like a four meter D cable from switch to switch. Uh, we can't actually say exactly what that meter reach is for 200 gig is what we're seeing is that all these systems are highly, highly engineered. 2 T kinda to five generation, basically it was like everyone stamping down a pizza box.
It was air cooled. You had front panel pluggables with a D cable. I mean, these days for scale up, everyone's doing these GPU racks where you got your DPU blades on the top and the bottom, you got the networking in the middle and everything is passively connected.
Mm-hmm. So, and just the different configurations, what they have as far as the number of xus, the distances are very different, but you know, this is industry leading as far as like 45 DB plus. So yeah, you can have passive copper connectivity, certainly within one rack.
I think you go outside of rack at that point, you'd probably need to have some kind of active copper, you know, either a EC or a CC or, or Optics. Mm-hmm. Okay.
Thanks. Important. Okay.
So we talked about, um, I'll just briefly cover. So, um, for a scale of connectivity, you know, to six can support, you know, five 12 x in a single hop network. So this is more than seven times the scale of any other competing technology.
And you know, again, we see that these scale up networks, you know, it used to be a couple years ago it was, you know, maybe eight GPUs. Today it's like 64 or 72 GPUs in a scale up network. But we're seeing people definitely wanna go to like hundreds 56, 5, 12, and we have requests for like one K or more XUS and scale up.
So just to be really clear here mm-hmm. Um, when you're talking about connecting these xus mm-hmm. Let's say you're using ultra ethernet, that would be 512 UA link connection.
Well, so UA link is a different standard. So I mean, it's basically, it's a, you know, they, they use mostly the ethernet FI with some modifications, but you know, we believe firmly that, you know, ethernet is the way you should be going for scale up. So UA link is something that, you know, we think is, is a, it's an enough standard, it's an alternative, but we think there are a lot of benefits to using ethernet.
Mm-hmm. Okay. And then for ultra ethernet, there are certain parts of the spec.
0 spec is kind of here every, here's everything you need for a scale out network. However, there are some technologies in there in particular with link layer retry, creative based flow control that can also be used for scale up. And that's one thing I think Robin will get into is, you know, what is it that differentiates, say a general data center network versus a scale out versus a scale up network.
Okay. We talked about one tier. Okay.
So, um, as mentioned before, you know, haw six is not just about speeds and feeds. There are a lot of features we have in here that improve the, the load balancing the traffic management. Oh, uh, Just, just one question.
Sure. If it's 512 GPUs per haw, six mm-hmm. What's a two tier network look like in this configuration?
Um, so two tier would be normally for a scale up network. People today are just using single hop and they're really trying to move, you know, stay within a single hop. They don't wanna go to multiple tier.
So that would be more on the scale outside and scale out. What you'd see is that, you know, you see clusters today, maybe tens of thousands of XUS within a building, maybe going up to a hundred k slightly more within a building. After that you run outta power.
So that's why you'd have like a scale, a two tier scale out network. You run out of power, not bandwidth. You run outta power, you just, you can't power the xus, they just can't get enough power in one building.
A Hundred thousand From a networking perspective, the single tier is basically, um, switching in the rack. Mm-hmm. So it's, it's, it's intra rack communications.
Yep. Mm-hmm. And then, so from that, from that jump from, from a single tier two, two tier mm-hmm.
Now we're talking about connecting multiple racks together. Yep. Exactly.
With uh, uh, I'm presuming it's a spine leaf configuration. Yeah, yeah. Normally, yeah.
So, and usually it's kind of, you know, a symmetric, you know, full bi sectional bandwidth. Okay. When you hit that, we can't bring enough power limit in, so it's not, what many reaction are we talking about?
Well, so it depends on how it, it depends on the data center configuration, but you can figure that today it's, what you'll actually see is that for some of the XP, the GPU systems that are being deployed, it's not the number of racks because each one of those racks has consumed so much power so that they actually leaves some racks empty. So it's just the total amount of power that goes into the data center that today it's on the order of like tens of thousands, maybe like 30,000, you know, XUS within a data center. Mm-hmm.
Maybe for these large data centers getting on the order of like a hundred thousand xus. Mm-hmm. But after that, they literally, they just can't supply the power into that building.
And how many xps are we talking about per rack? Um, so today it tends to be on the order of, you know, the high end tends to be like 36 to 72, maybe, you know, somewhere in that like tens, like low tens. Okay.
You know, certainly less than a hundred within a rack. Okay. And that's probably where you're gonna stay for the foreseeable future.
Just the power density within one, within one rack just becomes too high. Okay. Job B.
Okay. Um, one thing quickly I didn't wanna talk about is cognitive routing. One thing that's actually created or gotten a lot of attention is this thing we call global load balancing.
Mm-hmm. Um, so the idea here is that, you know, with say Tomahawk two, we started off with dynamic load balancing and this gets intelligence from a switch. It looks at all the ports, what's the load in on the ports, what is the queue depth for those ports.
Mm-hmm. And then we'll determine what's the best outgoing link based on those local metrics. With global load balancing, we actually get congestion information throughout the network that's fed back, you know, from switch to switch where we can determine the best global path through the network.
So that helps out a lot as far as just avoiding long-term congestion. And also as far as avoiding like link failures. So with this global load balancing, for example, we can react to link failures, which become a huge si huge problem with these large clusters 10,000 times more rapidly than if you had a centralized system with standard ethernet.
Mm-hmm. Uh, let's see. Um, we do also have a lot of features as far as telemetry and diagnostics.
So with our deep insight, we have a lot of features that, um, again, it's, it's not one size fits all as far as the congestion control that's on the endpoints, the characteristics of the endpoints, um, you know, the traffic patterns. So there's a lot of telemetry as far as, you know, fast congestion notification that goes back to the sender. Um, cig.
So this congestion signaling was being standardized within IETF. It's now being standardized within ultra ethernet that is supported by TOMA six. And then we have a lot of diagnostics at the physical link layer, so we can actually see if physical links are starting to fail, if they're getting kind of flaky.
And then you can proactively replace those links. Yeah. Let's see, we talked about that.
Okay. It's kinda like six summary. So you, again, twice the bandwidth of any silicon.
So you know, huge benefits as far as the size of the scale up domain as far as the, the fact that you can use the two tiers instead of three tiers for these very large XP clusters. Um, and I think we talked about a second point there. So, um, and I think one important point, point is that this is open, this is interoperable, you know, this is, this switch here can work with any endpoint.
So, you know, if you have a nick from different vendors, if you have an embedded inter, you know, ethernet interface on your XPU, this will work with that. You know, we have a lot of features, a lot of flexibility as far as telemetry, as far as load balancing and congestion control. Um, overall highest power efficiency, you know, what we've seen is that for the switch systems from previous generations, some of our competitors, just that one system will be 40 plus percent higher power.
So, you know, just the, it's not just the fact that we, you know, we're first to market, but this is actually, you know, by far the most efficient silicon. And also the fact that we have, you know, point solutions for scale up, for scale out for scale across means we can really optimize the silicon for those use cases. Right.
Which one of these is the six? Um, so these are both tamock six, this has 30 giger that has a hundred gig Thirties. Okay.
So the five, six bigger six. Little six, Yeah. Okay.
Thanks.