Validating Frontend Networks to Optimize and Secure Low- Latency LLM Data Flow with Keysight
As large language models scale, new challenges emerge – not only in maximizing GPU performance but also in validating the infrastructure that fuels the data pipeline used for training. On the front end, this includes securely ingesting user data from distributed cloud and customer environments into centralized AI data centers and ensuring high-speed, low-latency data transfer between VMs and hosts within the data centers. This session focuses on testing products that evaluate the performance, scalability, latency, reliability, and security of critical data pathways.
The presentation addresses the critical role of front-end networks in ensuring efficient data flow to data-hungry GPUs for LLM training. It highlights the two primary data movement patterns: north-south, involving data ingestion from sources like cloud providers and user environments, and east-west, which focuses on data transfer between virtual machines within the data center. Each of these patterns has unique testing demands, with east-west requiring ultra-low latency and high line rates, while north-south necessitates a balance of minimal latency and strong security to protect user data in transit.
Keysight’s solution, Cyperf, is a software-based traffic generator designed to emulate various application traffic types and measure key performance indicators like bandwidth, latency, connections per second, and security. Cyperf ensures that GPUs receive data with low latency, security, and at the promised rate of the network infrastructure. The presenter emphasized the importance of thoroughly testing all layers of the OSI model, especially the application layer, to avoid relying on end-users as beta testers. The presenter lauded Crusoe for proactively testing their front-end infrastructure to ensure optimal performance.
Presented by Amritam Putatunda Senior Product Manager, Application and Security Solutions, Keysight Technologies. Recorded live in Santa Clara, California on April 25, 2025 as part of AI Infrastructure Field Day. Watch the entire presentation at https://techfieldday.com/appearance/keysight-presents-at-ai-infrastructure-field-day-2/, https://techfieldday.com/event/aiifd2/ or https://www.keysight.com/us/en/products/network-test/cloud-test/cyperf.html for more information.
Transcript
So we want to validate the front-end networks and op and understand how we can optimize it for securing and low latency LLM data flow. Uh, but before we go to that, let's talk about what we do as in what we do in Keysight, right? So we are a emulation company.
If there is a certain thing that needs to be emulated to test anything in the network infrastructure space, for sure we have a product to emulate it, right? That's basically our purpose. We test things and because of testing there will be some optimizations.
And because of the optimizations, hopefully network infrastructures will operate better than they were before. And today we are talking about LLM ai, right? It's a tech field, AI tech field day.
And the majority of the sessions will be held by my very good friend Alex Bok, who will talk about the backend network and backend infrastructure testing that needs to be done for all these training workloads. But hey, before you need to go to backend, you need to go through me, which is the front end guys. So I'm going to talk about a little bit on the front end side of the AI and what we need to do on the LLM side of it to ensure that the front end is working as well as those backend GPU clusters, okay?
Now, in large language model, right? There is a certain level of data bottleneck problems, right? What, what happens with these GPUs is they do all these maths with these parameters and tokens and do all these functions, mathematical functions with the 10 thousands, 10,000 cores that each of these GPUs have.
But before they can do this, they need to have all these data, right? So that they can do the maths. And that's where the front end networking comes in handy.
You have the smarts that are trying to offload and make the CPUs and GPUs more efficient. You have all these network, uh, optimizations, encapsulations, and all these infrastructure pieces who is trying to give all the data continuously in a stream to this super data hangry GPU clusters, right? And that's what we are going to talk about in the front end networking.
And there are two segments of this data movement. One of them is the north south, right? This is where all these storages, all the data lakes, you know, the cloud flares, the air Azures, and everywhere your network data user's own data is stored.
That needs to be sent to the GPU cluster so that they can run the training. And once the GPU clusters have these data, they have certain VMs, right? There are certain virtual machines across these clusters.
They need, each of them need the data as well. So there is east west reversal of this same data that all these VMs are optimized, optimized and working in cohesion, right? So north, south and east west.
Now, both these data movements have their very unique testing demands, right? You need to ensure that these are optimized. For example, in east west is all about line rate.
You know, if, if it's a 400 gig Nick, it has to operate at 400 gig. If the Knicks have 14 million packets per second, we need to ensure that it can process that many packets per second, right? And everything should be in sometimes millisecond microsecond and even nanosecond latencies in the east west direction.
But in north-south, there is a little more leeway in terms of latency, not too much because if, if it's too much latencies than your rags, your inferencing and all these things will run much slower because you continuously fetch incremental data after you have fed the initial training data. So that's where in the north south you have to have minimal latency. Of course, you have to have that promised link that any of these cloud provider service providers are promising you, and you have to ensure that there is enough security because this data is user data that is going to be in motion over public internet.
That's the demand for North South. So you can see these two typologies here that we talk, that we are going to talk about. Now, I have done some fancy animations around it with the limited capabilities I have.
So what you can see is just to showcase what kind of voice video data transfer that happens over north-south, and some of it, again, when it's going through VM to vm, it's mostly database storage, some of these things, right? So just to show you the difference between the content variability that happens in a north-south versus that east west. And what we do in terms of the test tool is said to me to be coming to the product within five minutes.
I'm almost at it. You entered that in the rub contest. Uh, I I I did an ICAM seven.
Good response. So if you, if you're amazed by it, look at my next one. Uh, so what we do in terms of as a test tool is that we have something called cyf, some, if you have heard of it, it's a software based traffic generator that can be distributed anywhere.
It has all these fancy bells and whistles to generate all kinds of application traffic to emulate the voice video, um, you know, regular applications, your pictures, this video, whatever we are taking today, all these things, it can emulate some of these things and it can also emulate all this low, low latency traffic like, you know, UDP streaming and stuff like that. So with cy, what we are doing is we are measuring what kind of bandwidth, what kind of latencies, what kind of connections per second packets per second. What is the quality of services?
How many users, how many millions of IP addresses can access this traffic? And what kind of efficiency or the reduction of efficiency happens when you have underlying infrastructures that have these complexities like netting, proxying, encapsulations, encryptions. The whole idea of cyfe is to ensure that these data hungry GPUs, when they're trying to do maps, whatever data they need, they get it in a low latency time in a secured way at the promised rate of those network infrastructures.
But, you know, I think somebody said that we are the old people of tech, and I have enough, uh, gray hair now to also be in that, uh, you know, in that legion, whatever. So what we have seen during this 20 years in my, uh, career is that if this is the several layers of a SI, when we had first seen Stevens books and all back in the day when we didn't have chat, we to ask questions and ask what are the seven layers we had to learn our way through books. What we understood and what we have seen over the 20 years is that most customers, most users put a lot of emphasis in testing the lower layers.
These are the lights on and lights off. Mm-hmm. Everybody notices that the link is down.
Everybody notices a ping failing, right? Everybody notices if the switch is not sending traffic, if the routing algorithms are not working there. So people spend enough time and justification, it makes sense, right?
Because this is, this is where the core of it is, right? You first take care of heart and then you, you know, have your muscles. But when you move up the layer, unfortunately, that's when people start to take some shortcuts because they have already spent so much money ensuring these lower layers work, right?
So this is where they're looking at free tools, some kind of whatever optimized ways making shortcuts, right? But the point is, I will give them the credit. They still try to do tests, at least at the transport layer means P-C-P-U-D-P is working fine.
Mm-hmm. But many, many big companies, I'm going to not name them otherwise this will be my last conference ever, would always try to treat their users as their beta testers. Any of you who has used any kinds of cloud initially would know.
And this is where I would give a lot of crudo, uh, uh, kudos to Cruso who went ahead and said, I'm not going to let my users be my vita testers for my applications for their experiences. I'm going to do proactively some of this testing by myself and ensure that whatever infrastructures I'm giving to the customers, whatever that infrastructure is rated for, it's actually doing that exact job. And it's not, not going down when it comes to E-C-P-U-D-P and application layered reversal of data over this front end infrastructure.
This is a good segue to the next session, and I would request you to listen to that session or the video where Cruso is going to take you through what they did in terms of testing this frontend infrastructure for the optimum and highest performance.