Building Trust at Scale. How Crusoe Validates Network Infrastructure for AI Workloads with Keysight
In this session, Crusoe shares how they are actively testing frontend networks and inter-VM/host data transfers that feed their GPU clusters. By validating the performance, reliability, and scalability of its infrastructure early, Crusoe aims to identify and resolve issues internally, minimizing the chance that end customers will discover them first. This is a differentiator for them, which enables a more robust, production-ready AI platform. Crusoe is a vertically aligned AI infrastructure company powered by sustainable energy sources, including wind, solar, and geothermal. They build AI data centers, with a large project underway in Abilene, Texas.
Crusoe’s AI cloud platform offers infrastructure as a service, where customers consume GPU supercomputing via virtualized machines. They also provide managed AI solutions like AI as a service, inference, and workloads. Their mission is to build the world’s favorite AI cloud, purpose-built for AI, with enterprise-scale infrastructure. The company focuses on the design and engineering of data center networks, software-defined networking, and GPU-to-GPU fabrics, all optimized using NVIDIA reference architectures. They emphasize customer support, offering 24/7 assistance to address GPU systems’ complexities and potential issues.
Crusoe partners with Keysight to conduct rigorous testing to ensure optimal performance and stability, particularly focusing on stateful traffic and high connection rates. They simulate various workloads to stress the system and identify breaking points, provide deterministic performance, and prevent noisy neighbor issues in their multi-tenant environment. This proactive approach allows Crusoe to understand the system’s limits and provide transparent performance data to customers, ensuring a world-class service and preventing users from becoming beta testers. They use Cyperf as a traffic generator to understand the behavior of open-source OVS and NVIDIA’s stack to optimize testing. Plans include incorporating Blackwell platforms, advancing telemetry and monitoring, and focusing on storage optimization, scale, and security.
Presented by Gavin McKee, Cloud Network Infrastructure Architect AI/ML/HPC, Crusoe. Recorded live in Santa Clara, California on April 25, 2025 as part of AI Infrastructure Field Day. Watch the entire presentation at https://techfieldday.com/appearance/keysight-presents-at-ai-infrastructure-field-day-2/, https://techfieldday.com/event/aiifd2/ or https://www.keysight.com/us/en/assets/3125-1157/application-notes/High-Performance-Networking-Offloads-for-AI-ML-Focused-Cloud-Platforms.pdf for more information.
Transcript
I'm Gavin Key. I'm a principal engineer for network architecture at Cruso. I'm gonna talk to you about the pragmatic side of just what, um, just talked about.
Um, we are not an emulation company. We're a cloud. We're, uh, a vertically aligned, uh, sustainably powered AI infrastructure company, and we source, um, environmentally aligned power sources, um, whether that's wind, solar, um, geothermal, um, I'll talk about some of the projects, uh, to, to give you some context on these.
On top of that, we, we build, uh, AI data center infrastructure, and we have probably one of the biggest projects on the planet happening right now in, uh, in Abilene, Texas. Um, it's a gigawatt scale campus, and we have a world class digital infrastructure team who are leading that. And that group also, you know, they source the power and we take the compute to that power wherever it is in the world.
Um, on top of that, we build our AI cloud platform, and that's our infrastructure as a service. Um, so today customers will typically consume GPU Super compute, um, via virtualized machines. Um, and then on top of that, again, we build, um, managed AI solutions.
So AI as a service, uh, inference and workloads, um, would be one example. Um, the work I do at cruso has kind of been from the blank whiteboard, um, bringing cruso cloud to life, um, that has been, um, designed and engineering of the, the data center networks, um, the software defined networking that sits on top of that. And the, uh, GPU to GPU fabrics in finna band, et cetera.
All fully real optimized, um, NVIDIA reference architecture. Um, and our mission is really to build the world's favorite AI cloud. And, uh, we're purpose built for ai.
We have enterprise scale infrastructure, so you can get anything from one GPU to any number of thousands of GPUs in, in a single fabric. Um, we have a world class, um, customer support. It's one thing in this, in this problem space, it's so fundamentally important is a support because, um, you know, GPUs, these systems are massively complex.
Um, there's a lot going on. There's a lot can go wrong. You need to respond quickly because this is so capital intensive.
You know, customers trust us to both help them train and also inference. So if any infrastructure problems happen, we respond 24 7 3 6 5 as as quick as humanly possible. Um, and because we source power, we're vertically integrated, um, we feel we're more cost effective and sustainable as a platform.
Uh, talk about the, um, uh, sustainability and source power. Are you co-locating close to power? Is that your thing?
Colocating close the power systems? Yes. Or building the, the power systems themselves.
So, um, you know, a good example of a project, I'll introduce it in the next slide. We'll give you some examples of how we take the compute to stranded energy. And, um, you know, with these systems, they're like, so power intensive, you know, that you can't just go to a metropolitan area and take a typical data center facility.
Now you need in the tens of megawatts, if not hundreds of megawatts of power to build at scale. So I'll show you an example of, of two of these shortly. Excellent.
Um, I just wanna call out some increase cloud customers. We, we have some amazing customers on our platform. And, you know, one thing I'm particularly proud of is the fact that, you know, we are customers of our own customer.
So like, um, code are sort of rebranded to Windsurf. Most recently I used their AI coding assistant. Uh, it's, it's, it's awesome.
Um, but we have a range of customers doing absolutely amazing things in our, our platform. Uh, you know, Descartes build in, uh, world models. ai, um, go interact with Minecraft, uh, that's powered by, by Cruso on our platform.
And all of that inference and workload is done out of, uh, an environmentally aligned data center that we have in Iceland, um, using geothermal power. Um, and this, this just an example of like the first facility I ever worked on, um, to bring CR cloud to life was using, uh, digital flare. Uh, so we call it digital flow mitigation.
So we would build our own modular data centers and, um, we take them to remote oil fields and we build the turbine capacity to ex use that gas that's otherwise flared, um, into the atmosphere. Um, you know, you would see them on oil fields, just those big torches flying. Mm-hmm.
Instead of burning that, we combust that in, in generators and we power the GPU compute on site with that. So if you've seen where these places are, they're, it's, it's kind of amazing. It's just in the middle of nowhere.
Um, it's, uh, the operating conditions are really harsh during the winter. We have amazing oil fail operations teams who work in all environments, um, to ensure that we can keep that uptime on those sites. Um, you know, during massive storms, of course, if there's, uh, any disruption of power to the oil field operations themselves, you know, we can flip between different power sources on site and we can even gracefully take down a whole site if, if needs be, um, you know, renewable energy sources.
So something like I mentioned there, one of our, uh, facilities is in Iceland, so we use geothermal power there. Again, Iceland has an abundance of energy, so we take the compute to where that energy is and we build backbone infrastructure to connect our customers from either mainland Europe and the US into that, uh, facility. Quick question.
Yeah. Uh, Brian Martin signal 65. How far into that backbone construction do you get?
Um, are you actually doing undersea cabling? Are you leveraging existing cabling? Oh, we will definitely leverage, um, you know, third party providers that offer those, uh, subsea cable systems.
So we will go to, you know, major providers of that infrastructure, um, and, you know, by up to 400 gigabits per second of, uh, of, um, capacity on any of those. But we'll use multiple different providers for redundancy. Um, you know, what we'll talk about here very shortly is like the, uh, was was was talking specifically about how data hungry GPUs are.
Mm-hmm. You gotta feed them a massive amount of data and you gotta do it in a timely way. And data is distributed among many different cloud providers.
So if you need data to move from US east and your S3 bucket into a site like Iceland, we need to do that with high capacity, low latency links if possible. Yeah. It's, uh, I worked in a renewable energy company for a while, but, uh, you might wanna explain what digital flare mitigation is 'cause it's not something that everybody's gonna understand.
Yeah. I, I, I think, uh, yeah, it's a good question because I kind of like assume that everybody knows that this is a crucial branded thing. Digital flare mitigation is just like our term for, um, you know, using that, um, flare gas to power our compute.
Um, so instead of that, you know, burning that flare off into the atmosphere and all the toxins and all the bad things that come with that, um, just actually running that into our own generators and being our own power source on that site. It's, it has to do with like methane capture and Exactly. Yeah.
Yeah, yeah. Um, let's talk about accelerated networking on cruise cloud. So one of the things that, um, you know, I'm gonna talk about today, I'll, I'll, I'll not talk about backend networking here.
Which would you consider that GPU to GPU Alex as well is gonna cover that, but I'm gonna talk about accelerated networking on CR Cloud. So what it is that we're trying to do is, you know, it's a, it's a really true cloud experience. So we build a multi-tenant infrastructure.
So we have many different tenants running on top of the same physical infrastructure, and we want to accelerate, um, the customer workloads onto the hardware for the maximum amount of efficiency. The reason is that we want to give the customer the most amount of CPU capacity. We don't want to be using CPU for shifting network traffic around.
We want to do that in hardware. Um, so we implement a virtual private cloud networking. Um, and this is just a simple mental model that you can use for, for this slide.
Um, so each of these hosts down here have different tenants located on these, uh, compute nodes. And we basically, um, isolate those tenants over secure overlays. Um, and today we use, uh, geneveve as an overlay layer.
Um, and what that means is, like we, we have to, um, basically build a software pipeline that implements, um, you know, logical routers, logical switches, it implements things like DA and snat, um, load balancing and of course securities and, uh, access control lists and firewall rules. And as I said, this whole logical pipeline today is accelerated using the NVIDIA Connect X SMART Next. And more recently, we're, we're kinda working to bring the, the Bluefield three to life.
And then of course, on top of that, we have software defined storage for obviously moving data into these systems. Um, Kimberly Base, I'm the storage person here, or a couple others are What, what, what are you using for software defined storage? So we use a combination of things.
So, um, you know, for block storage, we will, we partner with light bits, um, for software defined storage, um, that would be for block, and then we use Vast for file systems and S3, et cetera. So we'll talk about data path acceleration, um, so in the, in the previous slides, but what we, we've seen is there's different contexts in which packets move in this infrastructure, it's either we can consider that, um, north south as it basically rees from the customer overlay or VPC networking, um, or East West, where it's encapsulated within the VPC networking across nodes and, and the infrastructure. Um, this is a mental model basically, of how this logical pipeline is implemented.
So customers have VPC networking, they have distributed logical routers, they have distributed logical switches, they have, uh, all their DA and SNET rules implemented in these logical router pipelines. There are firewall rules implemented in logical switching pipelines. Um, and these all have very, very specific requirements.
Um, so if we would need to accelerate this, our, our goal here is to reduce overall latency and CPU overhead while we pa uh, process a lot of packets through the system. We use stateless offloading for that east west traffic within that Geneve overlay, and we offload that onto the, the CX card. Um, we also have the stateful operations challenge.
So like I mentioned, as you egress from the VPC onto the physical network to get to services, whether that's, uh, S3 or it's just the general internet, whatever it is, you, you actually have to go through a connection track pipeline. So you're actually using an access, uh, control rule or you're using a NAT rule. Um, so you can be routed on that external network.
Um, so that's a really complex thing. Um, the performance impacts that we're really trying to mitigate here is like the high A CP utilization on bottlenecks as a result of that. Um, and again, the goal is to efficiently accelerate this onto the Nick hardware, and we're getting close to the point where we can really start to, to talk about how we have went on this journey from initially optimizing our pipeline to implement multi-tenancy VPC networking, to optimizing that further.
So, I mean, you can go and use, um, open V switch, um, and it's purely a software switch layer two very simple. You can accelerate that, um, onto an SWI on the nick, and you can do this all using a kernel pipeline where we have both a slow path and a fast path. We're gonna talk a bit this, I'll show a slide next that'll really illustrate this, but this is just, this slide shows you a journey that we went on in terms of like efficiently building our cloud platform on this front end networking.
So we found that, um, as we used OVS kernel, the slow path can become a bottleneck at very, very high connection rates. Why does that matter? Because you're moving from a world where when we started, everything was based on training.
There were no cool models that you could really use for inferencing. Um, neither are a lot. So the, the world is changing from just model training to actually serving these models through inferencing.
Um, so the workloads really, really are starting to change Question. Yeah. Um, the question about your, your journey here through Kernel to daca, um, did you also do some DPDK mode testing on OVS?
No, because connection tracking on O-V-S-D-B-D-K is just not fit for purpose. Gotcha. What we need.
That's a good question. Another question is, um, at this point is what's sort of your mix that you for cruso of tenants that are doing, uh, training versus, uh, inferencing? So there are multiple different types of customer.
Mm-hmm. Uh, some obviously will train and then they'll serve their own models. Other types of customers will resell, compute, and customers will come and inference on them.
And then there's of course, ourselves building our own inference in and sort of manage AI services on top. Um, so our own internal customers, um, but many customers will both train and inference because of course, if you've got all this expensive compute, you're really gonna put it to, to use and try and drive some revenue off it. Um, so OVS doca, what is doca?
It's NVIDIA's data center on a chip architecture. And Thank you. Um, yeah, it's, there's a lot of acronyms I know in this networking space.
Yeah. Um, so the journey that we went on was, okay, so how can we optimize for the hardware that we actually use? We don't build our own smart nicks.
We, were a very small team when we started this project, we're growing very rapidly. Um, but we, we leveraged like the help from NVIDIA quite a lot. And, uh, our attempt was to really improve the efficiency of the system and to facilitate these moves towards more inference and based based workloads.
Um, and our attempt here as we move from, you know, the OVS kernel, pure open source world to actually using more of NVIDIA's stack to optimize, um, was okay, how do we test that? How do we know based on our operational experience, um, that dokey is gonna serve the needs that we have? Um, so we're really attempting to get to initially like high connection per second rate of about 30,000 connections per second with 1 million in hardware.
Um, and ensure that under the most extreme load, that we are able to understand that, that the system will behave in a deterministic way. This slide shows the differences in, in these two models is like OVS hardware, TC offload to OVS, DOCA hardware offload. They're essentially very similar data paths.
Of course, you have slow path and you have fast path. Anything that's not, this blue line at the bottom here can be considered somewhat the slow path, right? It's exception packets.
It's the first packets in the flow. And we really want to test what the differences between these two, uh, two pipelines and what the impact of these is at scale. Um, the end result is we're trying to really do the same thing, but add very, very high loads.
We're trying to ensure that the, the determinism of the system remains. Um, so this is where I met my friend and, uh, colleagues. I would call 'em.
'cause we worked very, very closely with, with, uh, the Keysight team. You know, we looked at how are we going to actually test this to ensure that we can pull the system apart, that we can ensure not just the performance of the system, but the stability of the system under extreme load. Um, so we came up with, um, you know, uh, a test in setup using CY perf as a tool to generate stateful traffic through the system.
This is a contrived sort of example of how we tested this. So we come up with a topology that really exercises the code path, and then we use CY perf as the, the traffic generator to ensure that we can really exercise that, uh, that data path at scale. So, um, we, I, I would say probably the, the journey initially was like we, we didn't really know how to test.
We just knew that what we like, wanted to test, right? So part of the exercise and working with the, the Keysight team was really understanding the behavior of the product. And, you know, being able to actually generate stateful traffic is massively important to us.
Why? Because we don't know the customer workload ahead of time. So we have to build a contrived pipeline in order to stress the system as much as possible.
So we simply start with just some simple throughput testing and measuring the impact of that on the system. So with Cy perf, you know, we're initially just able to look at our virtualized infrastructure. I'll just skip back here one second just to call out here is what, what we have in, uh, virtual machines is, um, S-R-I-O-V.
So this is directly connecting us to the NIC hardware. And in order to achieve the data path implementation, you know, Linux has a a switch switch driver model, or we can implement the kernel bypass with NVIDIA's O-V-O-V-S doca framework. And, you know, when we wanna test the behavior of these two systems, we wanna run the exact same test, right?
And CY perf allows us to do that. But not only does it allow us to do that, it allows us to kinda modify, um, all the different parameters that you have. So, for example, we might test with very, very small packet sizes.
We might, uh, test with, uh, jumbo frames and kinda look at what the behavior in the system is, um, as we change those different workload types. So a throughput test, we just are looking for, you know, 400 gigabits per second bidirectional, uh, throughput. Um, just real quickly, I apologize if I missed it.
CPS, you? Yeah, I, I'm, I was wondering the same thing. So, Uh, you say, uh, perform similar at smaller CCPs CPS.
Ah, okay. Yeah. Sorry, say that again.
Connections per connections per connections per second. Yeah. So these, these are the results of like comparing the behavior of these two, um, data path acceleration pipelines.
So we have OVS, uh, doca hardware offload, and OVS kernel, uh, TC hardware offload. And you can see here like very, very similar performance, um, in terms of throughput, right? This is exactly what you expect because essentially you're offloading as much of the data path as possible onto the, the hardware.
And, you know, you just look here on the right hand side on the server core utilization. You know, we, we can measure the efficiency of the doca kernel bypass mechanism relative to the OVS kernel TC hardware offload. You can see that we're using more cores and OVS kernel mode and we're, we're super efficient on OVS doca, but that's not the whole story, right?
Um, we really need to test what happens with state full traffic. And this really can stress the system, whether it's, um, NVIDIA's implementation or it's the kernel implementation, um, with the OVS kernel module. Um, but in this very simple thing, you can see that we're, we're very, very close in performance still at about 25,000 connections per second.
This is where we're actually leveraging both NAS and potentially access control list rules. Um, but doca achieves this performance. Um, you know, initially this ramp up in connections per second, much quicker.
It takes OVS kernel a bit longer. If you're monitoring the performance of the underlying kernel, you're going to see like a lot of spinlock happening. There's a lot of work still being done on the kernel side with OVS doca and the kernel bypass is just a single CPU core that we reserve for this purpose.
And that's what gives us a lot of determinism. You know, in this example here, we're actually testing with very, very small payloads. So we're able to really stress the interaction between the, the hardware and the, Just a quick point on that, since the connections per second terminology came up, if you think of applications, right?
These issues will open it to many numbers, numbers of connections to serve their purposes. Now it is modern applications, right? They, they would continuously be chattier small packets, small TCP connections, but it, it's, it's a complete stream of connections continuously, right?
So that's why this is something that's not always tested because, you know, people, like I said, layer two three, forget layer four seven. So that's when, when you start looking at small packets, small transaction that happens in real life, you start seeing certain levels of problems that getting gets exposed and as early in the lifecycle that it can get exposed, the earlier it can be fixed. Yeah.
So those, those types of like performance journeys, it's really important. Like notice in the difference between these two lines. Um, the blue line here is showing like a lot, you know, you know, a lot less determinism in the, the system at scale.
Um, this is really important to our customers. Our customers want some sort of like, you know, deterministic band in which the performance of the system will remain. So for example, if we can say that, you know, the system will remain under, you know, one millisecond that might be completely fine.
They just don't want, you know, a a delta, uh, happening where you might be one millisecond and then spike into 15 or 30 based on the load that's, um, on the system. Also, the fact that some of the systems are actually multi-tenant, so you don't want a noisy neighbor problem. You don't want somebody that's co-resident on the same machine causing an issue for you.
You can see here again that at these sort of connection per second rate with the work that's going on in the kernel, that in the OVS kernel with TCA hardware offload, you are using a lot of cores because there's a lot of data path rules being programmed in the software data path and in the hardware data path. And this causes like a lot of like OVS is revalidation to, to happen, and you're using a lot of cores to to, to maintain that. Um, so around 28, um, less cores under this profile.
Um, but that's not enough. So we wanna go to 50 K, right? Um, and you can really start to see the performance Deltas appearing at these much higher, um, higher rates.
OVS doca is able to get to the target, um, objective much, much quicker. Um, OVS kernel initially struggles. What, what's happening in this white space between these green, green and blue lines is like, there's a lot of connections probably feeling, yeah, go ahead.
Yeah, yeah. How did you choose four 50 k, uh, as the payload size for this testing? Yeah, Yeah, it's a good question because like, we we're trying to keep the packet size small, so we want to basically stress the interaction between the software path and the hardware path so that the driver has to do a lot of work to figure out when I should hand off to the hardware to program the swi.
Oh, I was just gonna ask, do you have, do you have some um, uh, uh, real one numbers that let's a typical packet Size and or connection per second that you're seeing in an AI training scenario or an ai, um, you know, inferencing scenario to give us sort of, how do we re excuse me, relate this back to the real world? Okay, so in the real world, um, number one, we don't know ahead of time. Mm-hmm.
Um, every workload is different. This front end networking is not used for training. Okay.
Right. So you typically, um, we, we have Infinity band implementation on the backend. So all GP p GPU is highly optimized using Infini Band transport.
Of course there is Rocky and Rocky is RDMA over converged ethernet, um, you know, the Last Okay, You're learning this week. Uh, yeah. So, so on these workloads, what is it that we're really focused on?
It's that movement of data from external sources into the system to feed the, the, you know, that you have a storage cache or some cash in layer built locally on the system, and that's feeding the GPU because they're so data hungry. You need to have that sort of cash there, right? So, Would it be fair to say that all of this is, you're really trying to figure out how to optimize your environment as best as possible for a customer whose workload you have an idea that it's gonna be something, but you're not sure exactly what other than it's gonna be very high performance demanding.
Yeah. I mean, I would love if you came to me as a customer like six months before and told me exactly what you want to do, because then we could work together to build exactly what fits your needs. Mm-hmm.
But as a cloud provider, right, we sell on demand. Mm-hmm. So you can come to us and you can say, well, I'm gonna run this workload, why, you know, can I just run this and not have to change anything?
Right? Um, so being like really fast in terms of like delivering GPU capacity to customers to have all of the, the services we need around that, whether it's block storage, file storage, S3, um, whatever it is, everybody's somewhat different. Nobody does exactly the same thing.
So we try to build the system, and this is part of this testing methodology, is like, number one, we want to understand the absolute performance of the system. We wanna pull it completely to pieces in the lab. We wanna validate it, test it, and as part of the, the testing process, we wanna run those tests for long periods of time.
Why? Because we want to ensure both stability, you know, we want to check this there memory leak somewhere maybe over time. So some tests might run 48 hours, 72, whatever it is that we decide.
But the, the, the sole motivation is to understand what the upper bounds of system performance are so that we can be very transparent with customers and show here's the data of the performance of our system. Now you can come on the back end and you can do some nickel tests, or you can train a reference model and you can look at the performance of the Infinity band fabric, but different customers actually do multiple different degrees of, you know, parallelism in the, in the back end. So, you know, there might be tensor parallel data parallel, pipeline parallel.
Uh, everybody's somewhat different. Sorry for a long-winded answer. No, I Think, I think that's a great answer, really, that set the context.
Okay. So I kind of a follow up question is, um, I mean if you're opening this as a TC connection, why are you only doing one, uh, transaction per connection? Because that's the most stressful on the driver.
So essentially what you wanna do is you want to basically, Okay, so you don't wanna make, you don't wanna hold the stream open. Yeah, yeah. You could, you could hold the stream open, but what's it really telling you about the performance of the system?
These tests, again, are contrived to really exercise the worst possible case. Yes. And that's why, like, you know, side perf is very good for us because it allows us to control those knobs.
We actually do tests where we do multiple different transactions, but the most interesting ones that tend to break the system, and again, I work very closely with the Nvidia, um, doca team on this. Um, you know, they're very interested in what it exposes about the optimizations they need to make in their software pipeline. In what instance would you be with Nvidia and GPUs, would you actually be doing this kind stressing it in this level in real, real, real, real application environment In a, like you, you tend to not want to see these workloads.
You understand that you're testing to this workload, so you're, so clearly you're seeing this somewhere. So what we're trying to see is that as we put some, a system like this into production, that we can basically set, uh, a set of thresholds for the system. So for example, if a customer is inferencing on a system and that was exposed to the public internet, or they left a port open accidentally, and all of a sudden they get hit by a DDoS attack and they start getting smashed with thousands of connections per second, number one that we know how to track, you know, this, this metric.
And number two, that we hit a threshold, we want the dashboard to go red and go, we need to swing in and make sure that there is not something happening with the customer workload Beyond a denial of service. Is there that kind of peak threshold that maybe it's 75% of this number that you're seeing them hit in the environment, the AI training environment? No.
We, we don't, we don't see that okay. That those sort of rates today. But again, it's, if, if you get some customer that imposes that workload at some point in the future, you know, we wanna be ready.
That's Part of our engine. Do you even think we're gonna see that? And if you saw that, what would be that?
I, I guess what I'm saying I I'm looking at here is that AI has, is pushing everything we have mm-hmm. And you're testing at this level of, you know, what NASA and those kind of people, the worst scenarios, right? Mm-hmm.
What could blow up on the whatever. And so then you're saying, okay, so where, where would those threshold, what, what thresholds do you see on the horizon that practically customers really need to be planning for over the next four to five years? Yeah, so, uh, the, there are examples where if you implement this technology as say a gateway service mm-hmm.
Like the type of workload that that gateway service is going to be serving is like a load balancer endpoint. So there could be a massive convergence of traffic on a system that serves, say, a load balance to endpoint. Mm-hmm.
And that could be, hey, we have a really popular model that people love inference in. It might be deep seek, it might be something else, but Yeah. I was gonna ask if you, if you do congestion testing, because that's, that, would I, that's kind of what I would think of as another potential load that you would hit where you would potentially start dropping a packet every now and then mm-hmm.
Mm-hmm. And deal with the retransmissions to figure that out. Do you do that type of testing as well?
Yes. Yes. So it's, it's really understanding that that performance, so again, as we're a multi-tenant environment, so in some cases you have multiple tenants on the same system, we need to be able to number one, set metering thresholds for each customer.
As an example, as part of this testing, we will push each tenant workload to the set parameters that we give. And that ensures that number one, that if a customer really uses opera or eats all the resources that they have available, that they don't bleed over and steal someone else's resources. You know, and it's, it's those types of things.
I think like, uh, quite humbly, I, I learn about new things every day when building this infrastructure. It's, uh, complex. It, uh, you know, like I said before, it's like we don't know ahead of time what the customer's gonna do, but we just wanna really understand where the breaking points in the system are.
Okay. Yeah. This is a good example.
I was just wondering what else you'd tried. Yeah. Jim rinky, CDC.
Yeah. I was just thinking about this because of all the different things you might've deal with everything from like a smart meter output, or I'm thinking about that wind farm that you showed out in the ocean. What happens if a storm comes up in the North Sea, right?
And now suddenly you've got horrendous amounts of data coming in from the turbines because you're getting close to them exceeding, uh, their capacity or, or, you know, maybe, hey, we need to shift them so that they're feathered or whatever it could be. Right. Uh, or even things like, uh, ev EVs suddenly having to charge at much higher rates because it's winter time as we had in Chicago a few years back.
Mm-hmm. And you know, so this is, but I love the fact that you're, you know, looking, okay, what's the worst possible thing and planning for it because having been in that situation before where we didn't plan for it, I would've loved to have some sort of tool like this Yeah. And know that my system could handle it.
Yeah. I mean, it's a good example here is the work that, that we did initially, you know, brought Nvidia and Keysight a bit closer because as working with the Nvidia team, they were like, wow, we are able to really break our system. Let's use your CM test and methodology to actually test internally.
Now they've integrated a lot of this into their automated testing pipeline. So we work closer together, they learn from us, we learn from them. Yeah.
We partner with Keysight and we really can drive, you know, forward value for ultimately customers who invest in a lot of capital into paying for this infrastructure. We need to give them a world class service. Mm-hmm.
We can only be world class operators when we know what the absolute limit of any system is. It Goes to my last slide, don't let your users be your beta tester. So that's what proof is doing different here.
Yeah. So the, the last call out really is like, it took, you know, OVS kernel about five minutes to stabilize at the same connection rates. Um, and you know, that's, this is the outcome that we're really after.
We don't stop here. We go beyond this. We're gonna try and get the 8 million connections and hardware hopefully into the hundreds of thousands of connections per second.
Um, and this is the latency profile here. You can see this green lane is exactly where we want to be. We don't want this non-determinism in the system.
We're able to eliminate that. Um, and yeah, from, from our point of view, you know, we were too small to go and try and build our own traffic generator. You know, having worked with, with, uh, Keysight many times before, this is what these guys do, an emulation company.
Um, yeah. It's is massively beneficial now for us and our customers. Cool.
Cool. Uh, just what's next for us? We'll, we'll get the, the Blackwell platforms on, um, our cloud platform.
We'll keep, um, advancing our telemetry and monitoring as we roll a sim system out into production. And, you know, we're still massively focused on, um, you know, storage optimization and scale. And of course security is like at the front of mind for us.