Jericho4: Enabling Distributed AI Computing Across Data Centers with Broadcom
Jericho4 – Ethernet Fabric Router is a purpose-built platform for the next generation of distributed AI infrastructure. In this session, we will examine how Jericho4 pushes beyond traditional scaling limits, delivering unmatched bandwidth, integrated security, and true lossless performance—while interconnecting more than one million XPUs across multiple data centers.
The presentation discusses the Jericho 4 solution for scaling AI infrastructure across data centers. Current limitations in power and space capacity necessitate interconnecting smaller data centers via high-speed networks. Jericho 4 addresses the growing challenges of load balancing, congestion control, traffic management, and security at scale by offering four key features. First, it allows building a single system with 36K ports, acting as a single, non-blocking routing domain. Second, it provides high bandwidth with hyper ports (3.2T), a native solution for the large data flows characteristic of AI workloads. Third, its embedded deep buffer supports lossless RDMA interconnections over distances exceeding 100 kilometers. Finally, Jericho 4 has embedded security engines to enable security without impacting performance.
The Jericho 4 family offers various derivatives to suit different deployment scenarios, including modular and centralized systems. The architecture supports scaling as a single system through various form factors, from compact boxes to disaggregated chassis, and further scaling across a fabric. Hyper ports improve link utilization by avoiding hashing and collisions, leading to reduced training times. The deep buffer handles the bursty nature of AI workloads, minimizing congestion and ensuring lossless data transmission even over long distances. The embedded security engine addresses security concerns by enabling point-to-point MACsec and end-to-end IPsec with no performance impact.
Presented by Henry (Xiguang) Wu, StrataDNX Product Marketing, Broadcom. Recorded live on September 10, 2025, at AI Infrastructure Field Day 3 in Santa Clara, California. Watch the entire presentation at https://techfieldday.com/appearance/broadcom-presents-at-ai-infrastructure-field-day-3/or visit https://www.broadcom.com/ or https://techfieldday.com/event/aiifd3/ for more information.
Transcript
I'm Henry Wolf from the Code Switch Group at p robin. So today I'm going to talk about Jericho four for the scale across data center. So there are multiple ways we discussed to scale and build a large AI cluster.
So Scale Up is going to connect XPO within rack or among engine racks. And Scale Out is going to connect the racks port cluster within a data center building. Mm-hmm.
And scale across is to connect buildings, campers, e even regions together. And all of them is essential for the a scaling. So now let's take a look, look at the scale across looks like in the practice.
So this is a data center map in the United States, and actually we filtered with the power capacity 10 megawatts all about every circle represent a data center. The bigger the circle, the higher the power, uh, capacity. So instead of having a lot of big circle, as you can see here, in reality there's a lot of smaller circle and cluster together.
The reason is because the power and space capacity limitations for given facility. So the leading gen AI company has been grabbing the computation, uh, capacity. We are Aly capacity.
Wherever they can find and interconnect them through a high speed network, there must be a strong infrastructure to enable, uh, the distributed training. So what's the answer? The answer is ER code four, which can enable the XPU to unmatched the level.
So the latest, uh, AI model can easily have billions, even tradings of, of parameter and the demand enormous amount of data for the pre-training. So with a large, uh, AI cluster, you can apply all kind of parallel mechanism to greatly reduce the training time and with the writing of the reinforcement learning. So basically the different data center buildings, they can contribute and improve the model intelligence synchronously during the poster training.
But the, with such kind of scale, for example, one meeting XPO in one cluster with such kind of scale, the challenge grows significantly. Load balancing congestion, control, traffic management, and security all is becoming essential and critical. So now let's take a look at how JEAL four is going to address them.
So basically there are four pillar in the jeal. Answer number one is the scale. So basically with Jericho four, you can build up a single system with 36 K number of port.
So the single system means it's a non-blocking and single routing demand within just a behave as a single gigantic silicon. So all the complexity such as load balancing and the congestion control is going to be handled natively by hardware. And second is robot is really about the, uh, single fastest PO bandwidth.
We call it hyper po. 2 t, uh, fat pipe. We call it fat pipe.
This is two x, uh, uh, what's its latest ethernet port one point sixt. So basically inside a hyper po there's no hashing, there's no collision. So it is a native answer to the elephant flow from AI workload.
And the third one is really about the distance with, uh, embedded deep buffer from Jericho Jericho family. So we can easily address more than 100 kilometer lossless RDMA interconnection. I'm sorry, can you say that one more time?
You said a hundred kilometers? Yeah, more than 100 kilometer. Mm-hmm.
Okay. Yes. With a deep buffer.
So we are going to address that later. Thank you. Yeah.
And the last one is about the security. We have the embedded security engine, so therefore you can turn on security for all the front panel ports without any performance impact. So we have multiple derivative in this Jericho four family.
So we have Jericho four itself, which is aiming for the modular system. So we have half of the bandwidth for the front panel port and the other half to the fabric, uh, interconnection. And we also have the Kuran 40 device aiming for the centralized system.
So this is a 51 T centralized router. We have a 100 gig service version and also 200 gig service version. Mm-hmm.
So the, the device has been sampled to customers starting from the beginning of last month. So quick question, when you say, uh, 64 or uh, 36 K ports in a system, do you mean in the complete network? No, it's just a single system, single domain.
So you can, uh, logically this is just a gigantic citizen. I'm going to cover that in the following page. Okay.
But, Okay. I know it's hard to imagine. Yeah.
Question on the, uh, you, you mentioned within the past month, does that match up with any change in the object market that's happened? Like the reason that we can now do it between data centers matches with Yeah. Actually the, the third is we used in the, in this Jericho four family actually we used shared with the Tomahawk family.
Mm-hmm. So it's the same third core with uh, uh, such as to six, uh, as please show here. Mm-hmm.
So basically we see the demand from customer. They, they want to have 100 gig version in case, you know, some Nick endpoint. They only have support 100 gig and some demands with the 200 gig for the, to matching with the latest Nick capability.
And also with the 200 gig version, we support the one point sixt, uh, ethernet port. Mm-hmm. Now let's zoom into those four pillars one by one.
First of all, it's about how to scale as a single system. So basically we have multiple form factor. So you can start to build in Jericho four with a, a very compact piece of box.
It's a two IU piece of box 51 T, and it can be either 64, 800 gig or 32, 1 point 60. Mm-hmm. And for the middle, uh, middle side on the right hand side, basically customer can build up a modular chassis.
Therefore, uh, the redx will be much bigger. 6 or 1,800 GigE from a single modular chassis. And, uh, basically if you want to have a even bigger, uh, scaling, you can physically disaggregate the modular chassis, the line card and fabric card into multiple pizza box as showing this, uh, uh, middle picture.
So basically we have the line card device, we call it Jericho four, and we have the Ramon device for the fabric, uh, pistol box. And you can interconnect them through the standard ethernet optics. So basically it's get rid of the physical dimension of the, uh, modular chassis to enable up to 36 K number of port.
2 terabyte. So why we call it a single system? The reason is because no matter is the modular chassis or the d we call it the DSF in the middle and also the piece box, no matter which form factor, they basically the work together as a single gigantic silicon, gigantic silicon, since this is a non-blocking demand and single rotting demand.
So for example, in between the Jericho four and the Ramon, it has the native package spray, it has native spray, and also the end-to-end scheduling. So all those challenges we, uh, discussed, for example, the congestion control load balancing is going to be handled natively by hardware. And therefore the, the direct benefits is that, uh, it hide all the complexity by the hardware and you can attach all sort of endpoint.
It can be a very high performance ai nick. It can be a very sim simplified low cost Nick, and, And just to clarify, when you say a system, you mean a global system? I mean, a single device, a single device, it Acts as a single device.
Yeah, you're looking at a bunch of boxes. It acts as a single device. But this is, this is, are you Talking physically?
It's a, it's a system physically. Okay. But logically it's a single device.
Okay. It's a global system. Uh, potentially Yeah, Direct.
Direct, yeah. Between data centers and so forth. No, Actually you, you can use, for example, in the middle, the DSI, we call it distributed scheduled fabric.
Mm-hmm. So it can be a single system to connect up to 36 k and points in the AI cluster. It can be within a building or across a d uh, different building.
Okay. A system is a very nebulous term. So I was just trying to figure out the scale of what you're talking about.
Yeah. The scale, for example, for the regular, for example, for Pete, talked about about six, a single device is uh, 100 T, right? 100 t you have 1 28, 800 gig.
But for this single system, it has 36 K number of ports. It's like, it is like as if you had some monster node with Yes. With with, uh, yes.
Uh, 36,000 ports in it, it, the physical, uh, architecture of this is there, is there any, how involved is that? Is that, uh, is, is there, you know, um, a fair amount of work that needs to go into a physical architecture to enable the single system building to building rack to rack, whatever it is? Yeah.
Actually, no matter which form factor has been field proven. For example, the DSF we call the distributed scheduled fabric has been deployed by meta. And they publicly announced their deployments since last October OCP.
And in the recent, uh, public talk, they shared more information regarding their deployment scale and also the performance. For example, meta mentioned that they build up the, uh, multiple AI zone with ESF. Each AI zone, uh, has 18 K 800 GigE, and they interconnect multiple such kind of AI zone into a even bigger cluster.
So basically the piece box is very mature. The modular chassis, all of them has been more than decade. And DSF actually is started from probably around, uh, five years ago, five years ago, and has been field proven by the service provider and also the, uh, ai, uh, space.
Basically, the scaling doesn't stop as a single system. We want to, in a lot of cases, we want to scale across as a fabric. So basically no matter the, the form factor we talked, it's either a piece of box modular chassis or the DSF, they can be a basic building block and you can combine them flexibly to build up even bigger fabric.
So in this example, so we have two tier of the chassis to build up two close, uh, uh, fabric architecture. As each fab, uh, each chassis supports either, uh, 500, uh, one point 61 k, 200, uh, 800 gig or 2K, 400 gig. We can easily address 2,000,400 gig endpoint or half 1,000,800 gig endpoint.
So this is the basically, uh, simplifi the network in a great way. So instead of having multi real multi-plan, you can have a single real single plan for most of the AI scale requirements. And that's a non-blocking fat tree low network there.
Uh, actually inside each modular chassis, it is non-blocking. Mm-hmm. But in between where you need to run, uh, as a regular E-C-N-P-F-C congestion control mechanism, because in between of them, this is just a standard ethernet, right?
Yeah. Yep. I guess what I meant to ask maybe was full bi sectional bandwidth mm-hmm.
Through that. Okay. Yeah.
Yeah. That's great. So we continue.
Now let's take a look at how the hyper port, uh, works. So in a traditional way. So we aggregate multiple link members through ECMP or lag, and, uh, those flow is going to be hashed and pinged through one of the members.
The drawback is very clear. You have flow created and also the hashing polarization. So for example, in this case, if we have five flow, each is 500 gig.
So the fifth flow is going to overflow, and you are going to have 200 gig package drop. And another example is, uh, one point 60 elephant flow, uh, that will be going, is going to be paying. So one of the member and only half of the capacity going to pass through the, the other half is going to be dropped.
But with the hyper pos, we build up hyper PO based on the, uh, standard ethernet technology. 2 T pipe. Inside the hyper port.
There's no hashing, there's no collision. So every flow is going to take advantage of every member. So because of that, so on the transmission side, we do the package spray.
On the receiving side, we do the reordering. So basically we can make sure to deliver in order for every flow for the downstream devices. Mm-hmm.
So, uh, we observe the more than 70% link utilization improvement compared with the traditional way either ECMP or lag. So, which results with a much less training time. Now let's take a look at what is the distance implication.
So basically for the XPO in the scale out, all scale across network, they address and access the remote memory through the RDMA protocol. But the fundamental assumption for that is the network has to be lossless, basically, but the distance actually has a lot of implication over the lossless. For example, the light speed in fiber, the propagation delays five microsecond per kilometer and 100 kilometer, the propagation one way delay is, uh, half a millisecond.
So the round trip will be one millisecond with 800 gig traffic and points. So basically the wrong trip buffer demand will be 100 megabytes, and we assuming a 51 T device on both ends. 2 gigabytes.
So this amount of number actually is not possible for the on entry buffer. It must be addressed by the deep buffer. So the Jericho family actually has the deep buffer support decade ago.
So we started from the DDR based on technology and evolved to the HBM based on technology risk, uh, starting from GECO two. So this has been field proven in the, and a DC backbone, and it is the same technology and we just reapply and yielded for the AI scale across. So basically one of our customer is building the world largest RDMA deployment over 100 kilometer.
Now let's take a look at how the buffer actually works. So basically the, the, the nature of the AI workload is very bursty. There are many different reasons to create a burst needs, for example, for a single, uh, model.
For, for basically you have multiple different dimension, and those dimension is going to compete for the network, right? And the, in case you have a large cluster, you may run different talents or different, uh, different task. So those tasks or talents, they are going to collide and create a congestion.
So the, the buffer in the network switch must be designed to embrace and handle such kind of bursts. So inside Jericho four will have two different type of buffer. We have the TI buffer, OCB, and also the HBM deep buffer.
So the ti buffer is going to handle most of the case to guarantee the low transient latency, but in case just a congestion, the deep buffer kick in to absorb those, uh, burst needs to make sure are lossless. And on the right hand, we take a look at the logical view, uh, on the, on the buffer. So for example, when the congestion started, we want to kick in the EN as soon as possible.
Therefore we can use ECN to slow down the transmission. Mm-hmm. But be the moment you start ECN to the moment it takes effect on the standard, it takes at least one round trip time.
So the round trip type is dependent on the capacity and also the distance it scales linearly with the distance. The longer the distance, the, the, the bigger the, the, the buffer you need to host. So mm-hmm.
If I can ask, do you have, based on your experience in ai, do you have a guidance with, based on the, the distance of buffer size and everything else, what is a guideline for maximum distance between clusters or data centers? Yeah, Actually it's depends on the user, uh, user comment, but uh, from our buffer perspective mm-hmm. We can suppose to, uh, even more than 100 kilometers, no problem.
Right? Yeah. So basically we can support a few hundred kilometers, uh, very easily because we have very deep buffer mm-hmm.
In your Jericho four Henry, we're overtime. So I need you to wrap up as soon as you can. Sure.
And just to finish this page. Mm-hmm. So basically for the, and when the congestion persist, eventually we need to have the PFC turn along and to make sure the loss list, so the moment PFC is turned on to the moment it takes effect, it will, the, the buffer has to absorb all the on fly traffic.
This is also scalable with the distance. So the longer distance, the higher the ONF fly traffic you need to absorb. 2 gigabytes.
So no matter what. So you see the, you see the metrics, you see the mass. Mm-hmm.
And we, we need to make sure the buffer is deep enough to accommodate for longer, longer distance. Okay. So when the fiber goes beyond the building, so user will have much less control.
So it is critical to protect the, uh, valuable AI training data. So the Jericho four has embedded, um, security engine, it can enable point to point mac sac and uh, end-to-end map apsac. Mm-hmm.
So, uh, it has no impact on the performance of all the training. So to summarize, basically I think we, we have saved this page in probably piece slides, uh, outer for the HPC and the scale up and TOMA six for the up and scale out and Jericho four for the scale out and scale across. So the deployment can be very flexible.
So user will be able to use TOMA six for the scale out inside the bu, uh, building and scale across with Jericho four. Or you can use Jericho four all the way from scale out to scale across both of them has a lot of customer deployment. So in summary, so the distance actually you must be addressed with, uh, with multiple, uh, aspects.
So number one is the scale. So we have 36 K as a single system. Number two is, uh, you, you need to have answer for the elephant flow from AI workload, which is, uh, addressed by the hyper port.
Then we have the deep buffer to address the longer distance and the max stack EPIs stack engine for the security concern.