Broadcom Tomahawk Ultra Low latency High performance and Reliable Ethernet for HPC and AI
Tomahawk Ultra shatters the myths about Ethernet’s ability to address high-performance networking. In this session we will show how we added features for lossless networking, reduced latency and increased performance – all while maintaining compatibility with Ethernet.
The presentation introduces the Broadcom Tomahawk Ultra, a 51.2 terabit per second switch chip designed to bring high-performance Ethernet to markets traditionally dominated by InfiniBand, specifically HPC and AI. Addressing perceived limitations of Ethernet such as high latency, small frame size constraints, packet overhead, and lossy nature, the Tomahawk Ultra is a clean-slate design focused on ultra-low latency, high packet rates, and reliability. The chip is pin-compatible with Tomahawk 5, enabling quick adoption by OEMs and ODMs, and it’s currently shipping to partners who are building boxes with it.
Key features of the Tomahawk Ultra include a 250 nanosecond ball-to-ball (first bit in to first bit out) latency, high packet-per-second processing optimized for small message sizes common in HPC and AI inferencing, and support for in-network collectives (INC) to offload computation from XPUs during AI training. The chip also incorporates an optimized header format to reduce packet overhead in managed networks and advanced reliability features like link-layer retry (LLR) and credit-based flow control (CBFC) for lossless networking. Topology aware routing, enabling optimized packet paths in complex HPC networks, is also implemented.
The speaker emphasized that the Tomahawk Ultra aims to provide an open and standards-based approach to high-performance networking, adhering to Ethernet standards for compatibility and ease of management. It utilizes standard Ethernet tools for configuration and monitoring, with features like LLR automatically negotiating between the switch and endpoints. Broadcom has contributed the Scale Up Ethernet (SUE) specification to OCP to encourage an open ecosystem. The Tomahawk Ultra is positioned as an end-to-end solution for high performance, offering an alternative to technologies like NVLink in scale-up architectures while ensuring compatibility and openness.
Presented by Robin Grindley, Data Center Switch Product Management, Broadcom. Recorded live on September 10, 2025, at AI Infrastructure Field Day 3 in Santa Clara, California. Watch the entire presentation at https://techfieldday.com/appearance/broadcom-presents-at-ai-infrastructure-field-day-3/or visit https://www.broadcom.com/products/ethernet-connectivity/switching/strataxgs/bcm78910-series or https://techfieldday.com/event/aiifd3/ for more information.
Transcript
Everybody. I'm Robin Grindley. I am also like Pete a, uh, product manager in the data center Switch Group, CSG Business Unit.
And, uh, I'm the manager for the Tomahawk Ultra. So this is, uh, kind of a new segment for us. First chip in a new segment, really focused on extremely high performance ethernet, um, making it reliable, making it ultra low latency.
And I'll show you what it does and go into a little bit of a discussion about how we got here. So we started off in 2022 and we had a vision. We said, okay, let's, what is the market segment that we haven't really addressed with ethernet?
You know, ethernet was now dominant in everything from wifi all the way up to hyperscalers in their cloud data centers. We said, high performance computing still doesn't have or enjoy the joy of ethernet, so we gotta bring it to 'em. And, you know, InfiniBand was the dominant standard, is still a dominant standard.
And ironically enough, when HPC started, it started with ethernet years and years and years ago, and then realized, you know, ethernet really does, doesn't cut the mustard for high performance computing, which was unsurprising. 'cause ethernet is and was a commodity technology. You know, it was focused on low power, low cost, not highest performance, but just getting all of the enterprises hooked up, right?
And then later on moving into data center. So it was designed for a different type of problem, different type of application. So it was unsurprising that it didn't really address high performance.
So we said, okay, what are the things that really limit it for high performance? And what can we do to try and make ethernet available for high performance and, uh, give people an alternative to a single source technology like InfiniBand? So there were a bunch of misconceptions, right?
So our, our design team said, well, okay, what do, what do we think? What's the limitations of ethernet that people perceive? Well, it's high latency, okay?
They can't support, uh, very small frame sizes. You know, ethernet frame sizes tended to be larger, right? If you had very small packets, 64 bytes is the smallest ethernet frame size, you just can't crank that many through a switch chip.
It's got some limitations there. It has a high packet overhead. You know, you're doing an ethernet header, you're doing a TCP IP header, you're slapping on UDP.
So there's a lot of extra overhead on each packet that you send. And that's, you know, you look at the HPC world and they're hyperfocused on getting as much payload in the packet as possible and minimizing that overhead. Ethernet is just a lossy by nature, right?
And it can't really provide low end to end latency. So the design teams look at this and said, okay, well none of these are fundamentally an ethernet problem. It's just, like I said, ethernet was designed to be a certain product for a certain market segment, and there's no law that says we can't redesign the chip.
So in 2022, this project started, the original focus was on HPC, make ethernet work for HPC, and the team realized that this would have to be a clean slate design. So this was not taking an ethernet switch chip that we already had. We had, you know, tridents and Tomahawks and Jerichos and saying, okay, couple minor tweaks and we'll crank out a chip.
Mm-hmm. The team had to go back and say, alright, we can do this, but we're going to have to look deep in the pipeline and the way what we do packet buffering, and we have to actually re-architect things, keep it as ethernet. That's the overall fundamental requirement.
It has to stay ethernet, but everything else, um, we can address with, uh, re-architecting things. So the goal here was really to get ethernet up to this level. And we actually achieved this with Kama Haw culture.
And I'll go into some of the ways that we did this. 2 terabit per second switch chip. This is it.
Um, it is monolithic, which means there's no chip plates here. You, you mentioned Pete, if you saw those going around, you'll notice that on this central dye, you'll see on those other chips that there are little, little lines. That's where the different chips go.
And then the, the cap goes on top. This is all one monolithic dye, and this is necessary to get low latency as well as low power and get the efficiency out of it. Who Would make, who would make these manufacture these?
Are there a lot of different, uh, manufacturers who could put together something like this? The, the chip itself? Yes.
Um, yeah. Any, anybody who does ethernet switching, you know, large silicon ethernet switching certainly has the capability. Like I said, there's, there's no, you know, we're not violating laws of physics here.
We're just looking at the way that we do ethernet switching and say, okay, let's go in there and make sure we optimize it for low latency for high packet rates and some of the other features that we'll show you. So anybody could do this, but we started in 2022 and has taken us three years to get here. So it's not a trivial effort.
Right. So are you fing these yourself? Well, they're FBB at TSMC.
Okay. Yeah. I mean, we design, we fab at TSMC.
Mm-hmm. And these have already shipped out to about half a dozen OEMs and ODMs, and they're building boxes. 2 terabits per second.
So any OEM or ODM who has a Tomahawk five box can actually drop this chip right onto the same PCB and start shipping a product. So they're gonna come out pretty quickly now. Um, so this particular implementation, this is 512 ser of a hundred gig, uh, PAM four.
Okay. So you mentioned, Pete mentioned earlier with the tomahawk six, you see two variants. There's a hundred gig pam, four variant, and there's a 200 gig PAM four variant.
So the first instance of this is a hundred gig PAM four. Okay. What Is PAM four?
I'm sorry. Modulation on the wire pulse add pulse amplitude modulation. Okay.
Uh, with four bits per, um, per symbol. Does this imply we may be seeing a 200 gig in the near future? We have the technology, We don't, we don't pronounce things.
So, um, I'll just say that we have, we certainly have the technology. So, uh, performance, so if you look at historical switch chips, tomahawk is kind of our, our best in class mm-hmm. With latency, um, and was in the six to 700 nanosecond range, which is pretty good, well under microsecond.
Um, but for high performance computing, you need to get a lot lower. So we've actually achieved just under 250 nanoseconds of this is ball to ball. So the first bit into, first bit out on a packet coming through this chip, and this is, uh, independent, right?
So this is a fixed latency that you can rely on. It really doesn't matter what other stuff is going through the chip. Um, if it's all going through point to point, all crossing through the chip, you'll see 250 nanoseconds across all of these packets.
Okay. Um, the packets per second. So this is kind of the horsepower of the chip.
How many packets can come in and how many can you process and send out? So if you look at, um, hyperscalers clouds, they tend to have on average larger packet sizes, 400 bytes, 500 bytes. You could design the horsepower inside the chip to handle small packets, 64 byte packets.
But if you did that, the cloud guys would say, well, that's cool, but I don't need it. All that means is that you're wasting silicon area and, uh, power when I'm never gonna use it. So all of our chips are really designed specific to a market segment.
So if you look at the other chips, they were designed with lower horsepower. It's like, you know, if I'm not doing a Formula One race, I don't have to put a Formula One engine in this car. That's not what it's used for.
It's going to Costco. Okay, well this one isn't going to Costco. This one's going to the Formula One race.
So you can see that this is 77. Can I stop you real quick and ask the question? Yep.
So, just on, on that, um, explanation, it's great we're using the race car to continue with picture theme, but from, from a customer, um, actually application for yours, are you looking to, uh, are you targeting customers that are primarily doing training, that are primarily doing inference? Are you targeting like, what size customer are you you targeting? And like, I'd kind of be interested in that if you could talk about it.
Okay. So the, the packets per second here means you can handle small message sizes, which is something in the HPC world that they've always need. That's a fundamental requirement for their shared MEM and MPI type applications in the AI world.
So I mentioned we started in 2022 with this project, right? And in the fall of 2022 chat GPT was announced. So chat, GPT comes out.
Now we have ai AI kind of settled on this scale out and scale up, and everything was on training initially, right? Mm-hmm. For ai, it turned out that the AI scale up that we'll get into a little bit more, um, was very similar to HPC requirements.
In fact, we looked at this chip and said, well, we are already really good for a AI scale up. Um, we need to add one or two more features, and we'll kind of get into those in a minute. Okay.
2 T you remember Pete was talking about the rad, right? You want to go collapse the network and, and go as big as you can for these large clusters. 2 T has a 256 rad.
And it is really optimized for scale up. So within a rack or a couple of racks and scale up is typically where you focus on inferencing. So inferencing is very high performance.
You're, you're focused on cranking out as many tokens per second as you can, right? And it tends to be, um, kind of optimized models that are focused on one specific area. And in that case, ultra high performance, low latency, small message sizes, high packets per second is the goal.
So AI inferencing is a sweet spot, and HPC is a sweet spot for this chip could be used for scale out. Um, if you had a smaller cluster size, uh, there's nothing that stops it from doing that. But typically what we see is for the large clusters, it's the tomahawk sixes for scale out, and then the Tomahawk Ultra would be within one or two or four racks for scale up.
And what kind of customers are you targeting? Are you targeting kind of the, the fortune five hundreds, fortune fifties? Like what, what kind, what is your customer base Like?
Well, so right now, scale up, uh, as Pete mentioned, is mostly a, a customized solution, right? So you, you have people who make custom racks where they have, uh, XPU blades. They have these scale up blades in the middle, and then they have custom copper cablings in a back plane.
So these are all highly engineered bespoke solutions made by a MD Helio, super micro, all of these companies that are making kind of a, a, a rack as a basic unit that has all of the scale up in it. And then it has top of rack switches that do the scale out. So initially it's for those kind of customers, the whoever needs it, whether it's, you know, pharmaceuticals or government or DOE research labs, they would go and buy these racks and systems from vendors.
So, so we make the switch chips and we sell them to these customers who make boxes and then make rack solutions. Mm-hmm. Got it.
Okay. And, And question, thanks for that. Maybe you have it further in this slide as it builds out.
But is some of the architecture a change on the different memory types and how much memory you've put into the chip to achieve this? Not specifically the memory types, just the way they get used, just the way they Get used, Right? Right.
I mean, it still has, you know, SRMs and TCAs, all of the standard stuff, but it's the, the way it gets used in the chip that is really architected for this specific use case. So performance is one area efficiency and just another, so again, we talked about, you know, having these huge header sizes, okay, an HPC or a scale up network is a customized network. It's a highly managed network.
It's not a general data center where anything could be swapped out, plugged in. You don't know what server you're connecting to. All these X views are the same, all of them are genius.
They are all using the same interface speeds, right? This is just the way that the network is architected. And in that case, you have complete control over every element that is connecting to those switches.
So now I don't need an IPV six address to have these two XUS talk to each other. I know there's only sixty four, a hundred twenty eight, two fifty six of 'em, whatever it is. And so I can just say all I need is an address, 10 bits of address to address up to a thousand xus, and then a little bit of extra header.
But I need to keep it ethernet. So it has to have the same sa da format with the CRC at the end. It still has to be ethernet 'cause it should go into any ethernet switch.
Mm-hmm. And it should get, uh, switched the right way. So within the constraint of keeping ethernet, we can optimize these headers.
And so that is supported in this chip specifically for scale up ai scale up, there's this notion of in-network collectives, um, that has been called Sharp in other implementations. And what this does is during ai, um, there are these collective rounds, you know, all the XUS do their thing, then they all have to share data. Sometimes it's an all reduce where they share all the data and then they have all do a computation to reduce that data and then go onto the next round of the calculation.
So instead of having all the xus send, all of everyone sends their piece of the, the overall memory space to each other, they all redo the same calculation to reduce it and then go, the switch is sitting in the middle of all of this, it can catch the data on the fly and say, Hey, why don't I just do the calculation in the network itself? Do all of that calculation, avoid having everybody sending data to everybody else, and then just send the result, broadcast it out to everybody. So you're offloading some computation.
You'd rather have your XUS doing real work, right? This is kind of busy work accounting work that is necessary to get to the next stage. Why don't you offload it?
Also, really, you're saving communication. You're saving a lot of bandwidth. All of these XPOs are sending data back and forth when they don't need to.
We can have the switch chip to it. So these are these in-network collectives, You're sharing, you're sharing the workload across multiple GPUs and they're trying to update each one's, uh, version of the, the current network or, or whatever layers they're working on. And you're doing some of the actual algorithm activity com, computational activity in the network switch rather than having the GPUs do it.
Is that what I'm trying to understand? Right? And how does that fit with like Cuda and all the other software?
That's, That's, it's al it's already in there. So they, they have, they have a layer in Cuda where you set up a, a collective, like a Nickel and PCL L, right? Mm-hmm.
What that does on their hardware is it goes up and it sets this in-network collective in the hardware, in the switches. Okay. And then it tells the X ps, okay, you, you all are participating in this.
You all have to talk to this switch to get this collective done. Exactly. The same thing happens here.
All of the xus just are going to a Tomahawk Ultra switch and saying, okay, you are doing an all reduce, you're doing an all gather. These are the XUS that'll be sending you data. They send it, it does a computation, it sends everything back.
So it's exactly the same. In fact, the CCL layer, we have a plugin that just shoves underneath there and makes it look identical seamless, right? What's happening under the hood is the implementation where it gets implemented.
But the functionality of moving this collective into the network and saying, what is the collective, what are you doing that is all software layer in the C CCL that sets it All up and that's only part of Ultra? Or is that part of the, That's just an ultra right now, right? That's a new feature In Ultra, You would think that would be against the latency requirements.
It, uh, the computational activity would cost time. And It's actually better, right? Because otherwise Transferring as much data other, Well, otherwise what you have is all of these xus, they all have, there's nx they all have one of the overall data set, right?
They do their own thing on their own data set. Now everyone has to send a copy to all of the other N minus one X Ps. And every XPU has to do that.
So there's n times N minus one big transfers going through the networks. Mm-hmm. Then every XPU gets all n chunks of data.
Now it has to do a computation on all of that data. And they're all doing the same computation, right? So now you're wasting GPU cycles all doing the same thing until, and then they have to wait until everybody is finished.
And then there's a synchronization where they all have to say, yeah, yeah, I'm done. I'm done, I'm done. Okay, now we all go on to the next phase.
So instead of doing that, right, there's a lot of latency waiting for that synchronization and computation. If they're all sending their chunk of the data anyway, capture it here centrally, right? So it doesn't have to go to these N minus one, it goes to just one, right?
It computes on the fly as the stuff comes in. You don't have to wait for it all to come in. It does a computation and then there's a broadcast.
I'm already here at the central point in the network. Boom. I send it out to everybody and we're done.
So it's actually, it improves for things like MOE and some of the training and it improves latency quite a bit within the Rack. Within a scale up domain. Yeah.
Which could be multi rack. And then, uh, finally, in terms of reliability, you know, HPCE and, and obviously this is a benefit for AI as well, has always looked at, uh, network as a lossless fabric. We, we have to assume that packets get through.
We can't, we can't wait, you know, have a timer at some transport layer that waits to see few microseconds or milliseconds later, oh, I never got an act for this thing. I have to resend it. That's just too long.
So we add features like link layer retry and credit based flow control. I'll explain these a little bit more in a, in a further section. These are new features in Tomahawk culture Ultra.
0 spec. Um, and these are optional features, right? You don't have to have these, um, in your switch or in your XPU endpoint, but if you have them, then it makes the, the network more reliable and you'll get higher performance.
'cause you don't have to wait for these transport or software layer timeouts. So those features are in the uc one spec. Yeah.
And that uc one spec applies to the, some of the other tomahawk shift, right? Are those, are those features only available in the Tomahawk Ultra Right now? The, what's the breakdown between them?
The LLR and CBFC are only in the Tomahawk Ultra right now. Other features the, like the ECN and some of the things that Pete mentioned that are also in the UEC spec are in all of our chips. Okay.
Gotcha. Okay. So it's basically breakdown of the UC spec then figure out which typical for those Features, right?
If you want the full suite right now with all of these extra features, Tomahawk is the one that's Out there. Okay. Okay.
Um, so again, we focused on HPC initially, so the ULTRALOW latency, um, the small package handling, um, we also added, if you look at HPC networks, you know, typical data center networks, cloud data center networks, they're very homogeneous, right? There may be multiple tiers, but, but every rack really just has end links going up, and they're all kind of the same, right? The, the cost of any of those links, it's the same.
It's gonna be, you know, three hops or five hops to get over to here in an HPC network. It, it's more complex than that. They have different things like Dragonfly and, and Taurus where there's this notion of there's a bunch of nodes that are collected or they're connected very, uh, low latency locally.
And then there are near nodes and there are far nodes. And if your routing can be aware of this fact and say, okay, I know this is, if I want to get over here, it's a far node, but I can hop through this other guy, then you can optimize things. So there are some other features like this topology aware, uh, routing that were added in, uh, Tomahawk Ultra, but still staying with open standards, right?
Getting the cost efficiency, the power efficiency Quest question on the, the topology aware. Um, so you're building, uh, state somewhere across the fabric, right? Right.
Is that a cluster of information that's held in its own kind of Right. When you're, when you're, when you're doing routing, you can look at things like, okay, I have a packet that comes in and I have multiple links I know it can go out on. Mm-hmm.
I have alternatives right now it's going out on this link, but I realize in real time this link is congested. Sure. I'm seeing a lot of queuing on that link.
So I'm gonna switch it over to this other link and send it over here. Right. So is that, is that being done per, uh, per tomahawk or is it being done across or shared information across multiple blocks?
Both. Both. Okay.
So it's being done in the tomahawk. You can do it on local information. Um, Pete mentioned this global load balancing, right?
Where you take into account remote information as well. Downstream, I may say, okay, this link is really congested, this one looks better. Mm-hmm.
But if I go through this one, it turns out that two hops later, I'm gonna hit another congested link. So There's a, there's a sharing in the control plane of between hawks of what's happening between Yes. There's a real time sharing.
So there's inside the chip, right? Which is obviously hard real time. I mean, it's gathering it at hardware speed, at line speed.
Mm-hmm. And then there's a control plane sharing, um, at, uh, amongst all of this, which chips as well as global. And that doesn't require anything beyond the tomahawk.
That's, that's kind of what I'm getting at, right? It doesn't require an extra server or state information elsewhere. It's all just direct switch to switch.
Right. Okay. Okay.
Um, so our goal here is really to open all of this stuff up, open standards and, and make sure that we have a full ecosystem. You know, we, Broadcom has always been about open standards. We don't mind competition.
We're pretty confident that we can execute on a, a very quick cycle and have the FAFSA ser best power, et cetera. So what we're showing here is that we've contributed this thing called, uh, scale up ethernet, this SUE, um, and this is just an open spec. We've actually contributed it to OCP and all it tells you is these are the features that you should put in to get the highest level of performance from this kind of open ethernet, uh, networking.
But you can do other things as long as it's ethernet, right? So it's an ethernet header and it has ethernet features like PFC and it's using an ethernet fi, then everything should plug together. There are features in UEC, like LR and CBC, these are optional.
You don't have 'em, it's okay. You'll still get, you know, you plug in, uh, an endpoint, a nick or an XPU that has no LLR, no CBFC, you'll still get ultra low latency. You'll still get small message size performance.
Um, so it's really kind of optional how, how much effort you wanna put into your XPU. But as long as you're adhering to these interface specs, it really doesn't matter. You'll see a discussion in the industry of UA link over ethernet where yeah, the transport layer is UA link, but you encapsulate it on ethernet and then fine, you send it through a Tomahawk Ultra, right?
Um, any x other XPU accelerator, you can make that transport whatever you want it to be. It really doesn't matter. We have a small ethernet header that'll get you where you need to go and you can transport it over these over ethernet.
And that's really the goal. Yeah. Up here.
And, and you're talking about the freedom to implement however you want to, uh, in your, in your high performance network. Uh, and, and you're talking and, and we've been having a discussion about, you know, about, um, not just local routing, but global routing and global and, and topology awareness. And, um, kind of the broader discussion is how, can you talk a little bit about how if okay, I'm gonna manage this network, uh, and I'm a network engineer, how do I get my hands on to actually configure the network, Right?
So the benefit, Yeah. How, how does that happen? All the standard tools, this is just an ethernet network.
It's what you're used to day in and day out. So all of the same tools that you've been using for configuration, for monitoring, they will all work and they should work in the same way. These things just look like standard ethernet Now.
So do I get into the, do I have to get into the CLI on these? Can I push automation to configure the way I want it to? You, you, how do I, you know, You should not, you should not need to do anything differently.
So for example, LLR, this is something that the XPU and the switch chip will in hardware, when you power them up, they will handshake to each other and say, can you do LLR? Yeah. Can you do LLR?
Yes. Okay. It turns on you don't need to do anything, and then it'll just run in the background.
Okay. Now, if you say, I would like to gather statistics about how that LLR is running, well, those are standard APIs, si and Sonic and whatever you need, you can, it's just another thing with some statistics in it if you want to gather that. But the provisioning of it and the management of it is really just happening automatically.
So there's really nothing new that you have to do here. Okay. And that is, that is a goal of the whole effort is to make it standard ethernet, standard software, standard provisioning, management, everything, the control plane, keep it the same, so it looks the same.
Mm-hmm. Right? And I, and I, I, I agree with that.
I, I like being able to have that kind of con continuity, but I, I I, I, I would push back a little bit, you know, even if we're just talking plain old vanilla, I'm gonna move email ethernet, you know, there's, there's still some amount of configuration that needs to have happen. Um, that somehow there has to be some way to communicate, uh, with the switch to, to, to that configuration, whether or not we're doing it, you know, human style or whether we're not, we're doing it by a script. Um, And all of those should still, you know, if you have CLI scripts, they should still work.
If you are doing it by SMP, you should still work. If you are using Sonic, we support Sonic on all of our chips. Mm-hmm.
So you, You're, you're, you're gonna, you're expecting that whatever control plane, um, Michel is using, um, she's gonna be able to see the switch and configure it like she would the last switch configured, right? Yeah. Yep.
That is the goal. Okay, good. Uh, Cisco iOS plans.
Yeah. Well, no, so, you know, That's gonna be my second question because, you know, I'm not even sure I know how to spell that. Um, Probably with a q maybe two Q's, there's a Q in there, it's like the hedgehog, um, YQ Yeah.
Yeah. And then add over the y These guys, um, is there some ability to say, I don't wanna, I don't wanna run Sonic, I wanna run something else. Yeah.
Mm-hmm. Yeah. I mean, so Sonic in, in, you know, uh, the Hyperscaler data centers, it was started by Microsoft, right?
Oh, sure. Yeah. It was, yeah.
So I mean, it's an effort to standardize the network operating system, but it doesn't matter from our standpoint, you know, we, the, these, we have legacy compatibility back for decades, right? So any of those interfaces that you were used to using, I, I'm not sure that Cisco would, would like to be, uh, referred to as their iOS, uh, referred to as, sorry. Yeah.
Uh, nobody wants to be able, It worked on a 25 0 1. So the goal here is that if you use this scale up ethernet, so you have X views here, and then going through this Tomahawk Ultra, that overall end to end, you can have an XPU send out data to another XPU and it'll be under 400 nanoseconds. So historically we're talking about well over a microsecond, one to two microseconds of latency mm-hmm.
To get this kind of transaction done. Um, in the new world of ai and certainly in HPC world, that was not acceptable. So we're focused on end-to-end latency here, not just a latency through the tomahawk culture, which is already very low.
But these endpoints need to have a very light interface that, uh, reduces the latency on the endpoints themselves. You can come up with the wrong answer faster. Yes, that's right.
Well, yeah, you get the wrong answer is not our problem. We're just a transport guys. Right.
Well least it's done past we're just UPS or FedEx. We'll get, we'll get your packets there on time. Um, so there you go.
If you look at scale up, right, the dominant one right now is NV link. And so o obviously if we're talking about scale up, we have to look at NV link. So in a single tier of scale up right now, current generation of NV link goes to up to 72 as we showed with the Tomahawk six and the Tomahawk culture ultra, you can get up to five 12.
So you've got a lot of roadmap there. Um, obviously the bandwidth is quite a bit higher. 8 terra.
And then we support multi rack. We support both scale up or scale out. It'll go either way.
It's really just ethernet, it's the way you use it and we're really adhering to open standards. Mm-hmm. So this is part of the goal with Tomahawk Ultra is to provide an open, completely open and standards based approach to scale up.
Do you Think you have the latency are, is your latency somewhere to MV link from an XPU to XPU perspective? Well, they don't, they don't publish numbers. So as far as we know, yes.
Uh, and, and we published, we released this chip in June and we have the actual lab measurements to show a sub two 50 nanoseconds. Um, I was kind of waiting personally for the announcement saying, oh yeah, but didn't see it. So, uh, as far as we know, we have different versions of this through.
Some people care about highest performance, some care about just maximum density of fitting into their XPU. So the Sue spec that you'll see on OCP will address both. Alright.
This in in-network collectives for the sake of time, I think we're kind of running shorts. Um, I'm not gonna go through this in too much detail, but I kind of described this verbally. You have the xps, they all send their data, their chunk of the data.
The INC engine then does all the calculations, figures out the final result and then broadcasts it out, right? So you're saving communication as well as, uh, offloading some of that computation from the XUS themselves. This optimized header, if you take a standard IPV four packet with UDP, um, including everything in the CRC, you're talking 46 bytes of overhead per packets because we can print this down for the reasons I mentioned before, that we know it's a managed network.
We can get this down to 10 bytes. And what this requires is the XPU has to set up that addressing scheme. Mm-hmm.
Tell our switch if it's using it. Um, and then we just route it normally. So essentially we ignore pieces of the ethernet address that we know aren't gonna be used for this addressing.
Mm-hmm. And we can, I can take questions about this afterwards if you, you're interested in the details, but it's still ethernet. So it's still ethernet.
You can put one of these packets from an XPU into any other ethernet switch and it should work just fine. There's a little bit of load shifting. I mean, this is happening a lot in ai, um, too.
'cause you just, there, there's a lot going on in in XPU, right? There might be two or three chips going on here. A DPU maybe, but also just maybe a, a, you know, a CPU headnote or something like that.
Right? That's, that's handling Yep. Some of this.
So there's a little bit on the system architect to uh, enable this. Right? Okay.
Yeah. But the, the switch chip can get an IP packet and the next one is this optimized packet and then another IP packet. This is all packet by packet on the fly.
Right. So the switch dip handles both. It knows what to look for.
It looks at, looks at something in the header. Does feel a little bit, a little bit like that. I mean, you know, I get tired of car metaphors, but they're just so useful.
Yeah. 'cause there's a lot about this that, that, that, you know, feels like F1, which is there's, or maybe even IndyCar, which is, you know, you're just going around a loop, right? There's only so much you need to do.
Yeah. Um, so, so in this case, you're asking the system vendors to sort of create that loop for you. Is, is the, uh, 250 nanosecond latency associated with the optimized using that header or standard ethernet Head handle?
No, completely independent. Okay. Yeah, this is purely about payload.
And is this still working within like standard mts or is there like a modified, you know, bumped up to make it super size, like what you would think about with like jumbo frames and stuff like that get Well, Uh, we still support jumbo frame. Mm-hmm. Nine, nine k.
Yeah. Okay. Uh, the LLR quickly, what's happening here is at the link layer right over the wire, an error happens Now on the receiving end, whether it's a switch or uh, an XPU, we have FEC forward error correction.
Yep. So forward error correction can connect, can correct automatically one bit, maybe two bit errors. Mm-hmm.
If it cannot, what typically happens is it just drops the entire packet. Yeah. So what happens here is if we find some uncorrectable error, we will actually immediately send a retry request saying, okay, you broke this packet into end subframes, um, the third and later subframes you have to retransmit.
Mm-hmm. So at the chip end, it has to have a small buffer to keep a packet. Yep.
And wait until it gets an acknowledgement saying, I got all of it, and then it releases that buffer and then it buffers the next one. Right. And if it gets a retry, it says, start free in and just resend free to the end.
Right. So this is hardware that's built into the chip itself. Do the link level retry, link layer retry, but it happens automatically.
And what ha what the benefit here is that the upper layers never see anything. This packet that would've been lost and triggered some TCP or whatever, timeout, um, just got through, never see it. Right.
So better reliability. Similarly for the credit based flow control, um, essentially you're saying that I don't wanna lose drop packets because there wasn't enough buffering in my switch chip. Okay.
So the switch chip, the haw Ultra will send credits back to the sender and it'll only send if I know there's enough space to receive that, that packet. And this happens link by link. Okay.
Mm-hmm. All right. So this is not an end-to-end system thing.
This is just the two ends. Say I can do it, you can do it. Okay.
We'll do this right. Once the packet gets to the next link, it does the same thing. It says, okay, do you have the space available to accept this packet?
And so this is all happening in hardware in real time, and you don't need any, uh, communication overhead for this. So in summary, um, we hit the reality with TOMA HA culture ultra. We managed to achieve all of the goals.
Um, it's really an end-to-end solution for high performance, but staying within the ethernet family. Um, and it is shipping now, so you should see, uh, start seeing boxes coming out probably by the end of this year or early next year. So we started off with the InfiniBand vision and HBC and now, uh, along the way we realized we can also address scale up and, uh, provide an open solution alternative to scale up.
So all of these misconceptions, we managed to address all of them.