Maximizing the Performance of AI Backend Fabric with Keysight
This session provides an overview of the Keysight AI (KAI) Data Center Builder solution and how it supports each phase of AI data center design and deployment with actionable data to improve performance and increase the reliability of AI clusters. The presentation explains how KAI Data Center Builder helps streamline the design process, optimizes resource allocation, and enhances the overall efficiency and stability of AI infrastructures to achieve superior performance and reliability. The discussion focuses on the performance of backend fabrics within AI clusters, highlighting Keysight’s expertise in creating solutions for companies that build various products, including microchips, smartphones, routers, 5G towers, and AI data centers.
Alex Bortok, Lead Product Manager for AI Data Center Solutions at Keysight Technologies, introduces the KAI Data Center Builder product, which is designed to test backend networks GPUs use for data exchange, especially during model training and inference. He distinguishes it from front-end network testing solutions like Keysight Sideperf. The presentation covers the capabilities of the KAI Data Center Builder in emulating AI workloads, benchmarking network infrastructure performance, fine-tuning performance, reproducing production issues, and planning for new designs and topologies. The product emulates the behavior of real GPU servers, reducing the need to purchase them for lab testing.
The presentation also highlights two key KAI Data Center Builder applications: Collective Benchmarks and Workload Emulation. Collective Benchmarks allows users to zoom in on a single transaction between thousands of GPUs, which allows network and system parameters to be fine-tuned. Workload Emulation considers that GPUs not only move data but also compute, and that there is a dependency between how much data you compute and how much data you will move. It is noted that Kai Data Center Builder is a part of the larger Kai Solution Portfolio. The presentation concludes by leading into a KAI Data Center Builder demonstration.
Presented by Alex Bortok, Lead Product Manager, AI Data Center Solutions, Keysight Technologies. Recorded live in Santa Clara, California on April 25, 2025 as part of AI Infrastructure Field Day. Watch the entire presentation at https://techfieldday.com/appearance/keysight-presents-at-ai-infrastructure-field-day-2/, https://techfieldday.com/event/aiifd2/ or https://www.keysight.com/us/en/assets/3123-1809/solution-briefs/KAI-Data-Center-Builder.pdf for more information.
Transcript
Everyone is Alex, Alex Bortec, product manager at Keysight. Um, so I'm gonna talk about the performance of the backend fabric in the AI clusters today. Alright, um, quick overview.
What Keysight does, um, we create solutions for, uh, companies that build products. Um, it could be a microchip, uh, smartphone, uh, router, 5G, uh, tower or AI data center. We help them to move from idea to prototype, from prototype to a real product, and then understand how this product is work in production, accelerating their, uh, lifecycle of these, uh, products.
Um, now, um, in ai, uh, data centers, um, a lot of you already know there are at least two types of the networks. There is a front end, um, network and, um, my colleague Tom, uh, presented today, um, a solution called Keysight Cyber Perf that is well positioned to test, uh, frontend networks. Uh, and our customer cruiser told us this story, how they use that product.
Uh, I'm gonna talk about the, uh, a different product called Kai Data Center Builder. Um, and that product is used to test backend networks. The backend networks are those, uh, that GPUs use to exchange data, uh, mostly during, uh, model training, uh, but also nowadays, uh, uh, during the inference as well.
And, uh, so my, my, I'm gonna have two parts. So in the first part, I'm gonna introduce the Kai Data Center builder product, and then in the second half I'm going to show you a demo of that product. And, uh, I'm gonna end up with a real use case, uh, of fine tuning the performance of the Data Center network.
So, uh, stick around if you are watching live and if you are watching this video, make sure to go and, uh, look up for the second, uh, video part, which is gonna be the demo. Alright, so what the EQI data center builder does, it emulates AI workloads so that you can benchmark performance of the network infrastructure, um, figure out how you can fine tune the performance for them to be, perform better, um, reproduce issues you see in production, and also plan for new designs, new topologies, uh, new generations of those networks, and understand how they be, they would behave at scale. They do all of that, uh, by emulating behavior of the real GPU servers, so you don't have to buy those for your lab.
Okay. Um, so who, what kind of customers need that, uh, type of benchmarking? Um, so we have three very established categories of customers.
Um, uh, first of all, it's large scale AI operators, uh, as well as GPU cloud providers. The second is the network vendors that make routers and, uh, switches and all that kind of stuff. And then, of course, the silicon vendors that create, uh, you know, ASIC silicon and, uh, that goes into the switch or a nick or May or very often nowadays into A GPU as well.
Now, how does it apply, uh, to your company? Well, um, if you, for example, have a performance issue with the GPU Cloud, you might very well benefit from using our product to understand where these, uh, performance problems are coming from. Also, if you're building your own GPU cluster and it has scaled at least 120 80 GPUs, uh, you might, uh, you know, be in the position to already benefit from, uh, from using this, uh, this testing solution.
And then finally, if you're thinking of buying GPUs to have them in your lab just to run network experiments, just, you know, consider alternatives. Uh, you, you probably will find better use for those GPUs. All right, so let's talk about why, uh, using GPUs to test and benchmark is not the best idea.
Um, so obviously they cost a lot and hard to get, um, but now if you are an engineer who is focused on the infrastructure, especially if a network engineer, um, running end-to-end system with GPUs and, uh, software stuck on it comes with a lot of complexity. You have to learn a lot of things that you actually never learned before. So that stands between you and achieving your goal of, uh, you know, having network to perform the best.
And now even if you figure out how to use, you know, nickel tests not that hard. Maybe you figure out how to run ML perf much harder. But in any case, um, when you run them and then you collect, uh, data, uh, from different parts, you actually don't know how to correlate them.
So you see some performance, um, you know, values, uh, from the, from the MEL curve, uh, and then you see some metrics from the network or network card. You don't know what caused it, what, right? So this metric acceleration is a big, big problem.
Uh, if you are not using specialized tool. Now, um, what do we offer here? So the, we offer a couple of solutions.
First solution is a software from key side, which you can put on on real servers, uh, that would pretend to be a GPU, uh, in terms of data movements, right? It's not gonna pretend to be a GPU in terms of the extra training, the model, but it's gonna pretend to be a GPU in terms of data movements. So it'll work with the natural cards, uh, that, uh, perform hardware flow of, um, RDMA traffic, um, uh, also so-called rocky traffic, uh, that we were talking, uh, you know, earlier today about, uh, so that you can figure out, uh, how your network and the network cards perform altogether.
And you don't need a GPU on such a server, right? So it's very handy to have that in your lab. It's, uh, costs less.
And, uh, it, it covers this gap that I was describing where the metrics are not correlated. Now we actually know what's happening in the system. We know what to ask the network card to do.
So when, when it does the job, we can understand, uh, you know, uh, if, if the job was done correctly or not, and on every single natural card and every single data movement, uh, levels, right? We can also put that into a, a server where you have GPUs, so they production. So if you're renting GPUs from the cloud, it's fine.
Just put them over there and experiment, uh, to find out if, uh, if you can, you know where the problem is. Now, this solution is sometimes, you know, have gaps where, uh, for example, if you don't know where the problem is, is it a network card or a network? It can hide it because we are operating on the level of the software.
Um, and then also if you're planning to put a new generation of, uh, network fabric so that when the new generation of GPUs come ready, the network is already tested, right? It's hard to do because you don't have, um, these new servers with the high performance network cards. So what do you do in that case?
So for that case, we have a second solution that we use with exactly the same kind data center builder software. But instead of using real servers with the natural cards, we emulate those natural cards using our Keysight 81 traffic generators. So those traffic generators, they capable every single port of them can emulate one or more, um, Rocky Network cards.
Those are the main, uh, capable network cards so that we can stress out the network now pretending it has a bunch of, uh, you know, Nicks and, uh, GPUs behind them connected to it. And so that how you bridge that gap, uh, for testing the new generations of the networks before the gen new generations of the GPUs, uh, become available. And we are also able to see very low level details on every single packet that way.
So, uh, there is no this hidden area between the nick and the network. In that case, if, if you wanna test the, the, the network itself, this is the best solution. Alright?
Now, what does the user actually sees when, when, when they use, uh, this, uh, product? So we have two applications for 'em. Um, so the first is collective benchmarks.
Uh, it's been around for almost a year, so it's pretty much sure. Now there's plenty of customers who use that, and I'm gonna demo that product to you today, that application. And then we, earlier this month, we, uh, we announced the workload simulation, uh, product.
And, uh, so, uh, yeah, we had the first live demo of that, uh, during the OFC conference, uh, earlier this month. So how do they differ? So the collective benchmarks application, what does it do?
When you have an AI workload that is, for example, you are training the model, um, it uses collective operations. Uh, this, this, this is a terminology to move the data between the GPUs, right? And it does it move over and over again, many, many, many times, millions of times.
So if you wanna benchmark a collective operation, right? You are zooming in into a single transaction that happens between thousands of GPUs, right? And, uh, what your objective here is, is to, uh, make sure that all the parameters of the network and the system are fine tuned in a way that the network utilization can be as high as possible so that this collective preparation can finish as fast as it can, right?
So you are you, like, you, you narrowly zoom in into the performance of that, uh, data exchange between the GPUs workload. Emulation, on the other hand, takes into account the fact that all GPUs that don't only not only move data, uh, they also compute. That's the point, right?
And, um, there is a dependency between how much you compute and then how much you, how much data you're gonna move. And sometimes those, uh, go in parallel. And so in reality, uh, it's not necessarily important to always achieve the best possible performance for all types of collective operations because a lot of them are sort of hidden but behind compute.
So it doesn't really that important to get the very best performance. What you need to know is what's the critical path in your workload? What are the things that are actually slowing it down?
Mm-hmm. Right? And so if you optimize that, and essentially the, the, the key terminology there is ex exposed communication time is the, the time that GPUs really have to wait, uh, for the data movement to finish.
So the workload simulation allows you to identify that, uh, um, uh, you know, critical path and, uh, do steps to optimize, uh, uh, optimize it, right? Um, and then, uh, finish up that part of the presentation. Uh, Keysight, uh, uh, Kai Data Center builder is a part of the larger, uh, portfolio, uh, that IGN has, uh, you know, it's called Kai Solution Portfolio.
Um, as you can see, pons from solutions focused on, uh, testing elements that go into the servers, the compute part where we test memory, where we test P-C-I-E-C-X-L, where we help you design new silicon. Um, we also have solutions that test, uh, cables, transceivers, you know, interconnect at speeds of a hundred, a hundred, uh, KBS per second, one point 60 and further. Um, and then, uh, we, we spoke about the network part, right?
And then the power part, of course is very important. You want to understand how all the choices you can make in your, uh, design can influence the power consumption and, uh, in dissipation Under your power consumption. Do you also take into account cooling?
Yeah, of course, of course, Of course. I mean, then you can snoop power, right? You have to cool.
Essentially, yes. Uh, so that's the, uh, the thermal loading. Yep.
Mm-hmm. They still measure it in PUE, um, U-E-P-U-E. Now you gonna make me what it is?
Power Usage. Don't ask me about that too much. I'm a natural guy, right?
So, alright, so, okay, so what's coming next is the demo of the Kai Data Center builder. Um, so I'm gonna, uh, talk about our data center, uh, fabric test methodology. Essentially that's the way how you can test our product to get the best performance of your fabric.
And specifically, I'm gonna show you how to fine tune, uh, DCQC and congestion control. This is a voodoo magic for a lot of people. And so I have on this track for you, okay?