Netris Presents at Networking Field Day 40
Networking Field Day 40
April 8 – 10, 2026
#NFD40
Transcript
Welcome back everybody. I hope you had a great lunch, snack, whatever you consumed to get you through our last presentation for Thursday, because we have a two-hour block coming your way from Netris. And I was very impressed the first time that Netris presented at one of our events, and they have only grown from there.
They brought this fancy presenter and they brought a fancy billboard too. So it will be a fun time. They have a great presentation lined up, and there's going to be oohs and aahs and a demo.
You're going to want to stay tuned, so make sure you're taking some very good notes on this. We love all of our people out there that are participating during Tech Field Day. Remember that you can leave a comment on our video streams if you're watching us on YouTube or you're watching us on LinkedIn Live, there is an opportunity for you to comment.
And a lot of those comments, especially if they're good questions, we'll relay those into the room to make sure that you're getting the answer that you're looking for. Can you re-share that? Oh, there we go.
com, thank you so much for tuning in. com that includes links to all of the presenters, all of the delegates, all of their social media, all of the coverage that you're seeing and going to see from this event. com for all of the great stuff.
I'm actually going to step out of the way here because Alex has a very good presentation. He's way better at this than me. But we will be back shortly.
Remember, if you need anything from us here in the room, we're @TechFieldDay on all of the socials, and our good friend Corey Dearig, who monitors the socials for us, will definitely help you out or ping me and make me do it. But no matter what, we're here for you. Alex, are you ready to go?
Yep. All right. Take it away Thanks, Tom.
Thanks everyone, delegates, and Netris marketing team for helping to organize this. I'm Alex. I'm the CEO and co-founder at Netris.
Before Netris, I've been in network engineering many years, and I guess between all of us in this room, we probably have thousands years of network engineering background collectively. And what we do at Netris, we work with NeoCloud AI factory operators. We help organizations deploy large GPU clusters anywhere between 1,000 to 40,000, 50,000 GPUs.
Anywhere between few hundreds to many thousands of switches in every cluster. And networking is quite different in AI compared to traditional data center networking, and my today's presentation is about that. So we're Netris.
We're a network automation, abstraction, and multi-tenancy company. We call this NAAM. NAAM is a term we invented, and NAAM is what comes after SD-LANs and intent-based networking.
And that's because when you're building AI infrastructure, networking for AI infrastructure, for hosting lots of GPUs, you have different networking requirements. And you cannot use SD-LANs because everything in software, and you're pushing lots of traffic, and you cannot use traditional fabric managers or intent-based networking technologies, because managing the switch fabric is not enough. And I'll walk you through more details.
And one close example of NAAM that we all have seen, other than what we do at Netris, is basically what large hyperscaler cloud providers, they built and are running under the hood. The technology that we've created is something similar, but it is more designed for AI cloud operators. We're blessed to work with a lot of them, and I'll tell about all that experience.
But before that, before we proceed to AI and our recent success in that space, I wanted to talk just a little bit about our pre-AI customers. So these logos you see on my slide, these are our traditional data center customers before AI, and these are organizations who were hosting many switches. For whatever reasons, they needed to be on-prem, and they needed dynamic network configuration for their infrastructure.
So they were basically building private cloud. So we worked with these amazing customers, and this is where our technology evolved. And when we entered the space of AI, which for our company happened roughly year, year and a half ago, we got tremendous traction.
And it's been roughly a year when we became first NVIDIA-validated network automation vendor, and since then we've been working with lots of AI operators. Over last 12 months, we basically launched more AI clusters than entire networking industry combined, and I'll tell more about that. So some of our customers, these logos on the slide you see, these are all commercial, paying actual customers.
It's not marketing. These are not POCs. POCs is way bigger slide we cannot show yet.
We'll show next year. So these are actual customers. Most of them are running large NVIDIA GPU clusters.
One of them is running the largest AMD GPU cluster. These customers are hosting anywhere between 1,000 to 50,000 GPUs live, not theoretically, live today. And all of them are trying to go and host 100,000 to million GPUs.
And that requires a lot of complex networking. So since last time we presented here, roughly one year from that, so we launched roughly 20 AI clusters, pure AI, so we had customers before AI, but this is, I'm talking pure AI GPU clusters. And that happened to be 12% of NeoCloud market.
NeoCloud market is growing, of course. It's going to be way bigger than it is today. But today that's 12%.
And our team, we had to grow our team. Most of you know that we're headquarter here in Santa Clara. But today we have offices across six countries, because customers are all over the world and the team grew three times.
And the revenue happened, absolutely not important, but revenue happened to grow 800% in just one year. And I'm not gonna talk about commercial side anymore. This is the last commercial slide.
So let's talk tech with respect to AI GPU clusters. So this diagram is basically what everyone who's buying GPUs, this is literally what they are trying to achieve. They'll buy GPUs, they'll buy some networking hardware, some fancy expensive switches.
That's how it works in theory, but in practice, there's a lot of nuances, especially around networking. And this session is for network engineers who are working for organizations where they are in charge of designing networks for their large GPU deployments, choosing solutions for their GPU and AI strategy. And we'll explain what are the nuances and what are some critical considerations we learned from our experience that you may want to pay attention.
So this presentation is about bare necessities when it comes to network automation for your GPU AI cluster. And whether you consider building your own network automation or buy a technology, these are same considerations you need to think about. And among our customers, we have customers who are just starting their GPU cluster journey, who just deployed their first thousand GPUs, but we also have customers who are deploying their cluster number eight, and they are already 50,000 GPUs in their journey.
And we even have customers who are very advanced. They developed their own in-house network automation technology, and they ended up transitioning to our technology. So I'll share more about all the learnings.
So this diagram is a network diagram for a GPU cluster, and if you think this is complex, I have to spoil you because this is smaller than the smallest customer that we work with. And even this tiny cluster, which has only 512 GPUs, it's too small for a Neo cloud or too small for making real money out of GPUs. This tiny cluster alone has 1,200 links, 1,244 network links.
And the big challenge with these GPU clusters is, unlike data centers, networking is not fire and forget kind of situation. Networking is very dynamic. You have to have multi-tenancy, and this means you have to have a way to dynamically reassign your GPU servers from one tenant to another tenant, dynamically reconfigure every network switch.
" Which is not uncommon case. But see, even when Neo Cloud is building a single-tenant cluster that has to be consumed by one tenant, sometimes it's 25 tenants, sometimes it's 500 tenants, but I'm picking this extreme case, one tenant. Seems to be easy enough.
But listen, to service one tenant, a Neo cloud needs to support multi-tenancy because you need to have minimum two tenants in your infrastructure. In reality, you're going to have more, but minimum two is hard requirement. Why?
Because the way it works, you allocate some resources to your tenants, some GPU servers, and you as a Neo cloud, you're responsible for the infrastructure, responsible for power, and just like that you are responsible for networking, the physical networking configuration of the switches. And you are responsible for, let's say a GPU failed, you are responsible to replacing that GPU. And let's say you allocated the access to your tenant, and they are working and they are happy, and then a GPU failed somewhere.
Now it's on you to replace that server. And to replace that server, you as a Neo cloud, you cannot have access to that server. That is your tenant's server.
Think about it. Can AWS access your data? I'm not saying technically.
They are not allowed to access your data, right? Just like that, you as a Neo cloud are not allowed to access your tenant's server. So what happens, you need this other tenant, so you can take that server out from your main customer's tenant, put into your service tenant, and service that server, replace the GPU, reinstall the operating system, make sure it works.
Then oftentimes uninstall the operating system and give it back to your tenant. So that simple requirement makes multi-tenancy hard requirement. And also, oftentimes when you're building a Neo cloud or just an AI factory, you just want to have this ability to share the infrastructure with lots and lots of consumers.
So multi-tenancy is a must. Now, how do you reconfigure 1,200 links dynamically on the fly without breaking things? And this is super simple example.
Real examples are much more complex. So look, AI networking, it's a cake of your own. Maybe everyone heard of this Jensen's concept that AI is a cake of layers.
Now, networking itself has a cake of its own, its own layers. It's not just leaf-spine networking. Leaf-spine network is just one part of it.
Your GPU servers are also connected to east-west or backend network, which is usually rail-optimized network, and it is responsible for GPU to GPU connectivity andAnd because it's rail optimized, we have a lot of switches. And sometimes it's not one network. Sometimes we have two backend networks, sometimes we have four backend networks.
These are called single-plane, dual-plane, and quad-plane networks. Newer generations of GPUs, they require more performance from network, and we built quad-plane networks, so this is lot more networks. In addition to that, rack scale networking or scale-up networking, what NVIDIA calls NVLink 72, that used to be a bus inside the server, but now this thing became a network of its own.
Now you have additional nine switches in each rack. So in each rack, to service a rack of GPU servers, you need few per each rack. You need few front-end switches, you need a ton of backend switches, you need nine NVL switches.
And that's bare minimum. And over time, host networking is becoming also complex because the way workloads grow, there is this growing need to also use DPUs in actual DPU mode, not in NIC mode. And I will talk about this, and I will show this also in the demo at the end.
But essentially, there's this DPU thing sitting inside a server which is like additional switch inside the server, and that's a network of its own. That's additional fabric that needs care. Networking inside the host sometimes becomes challenging in some specific cases, I'll explain.
And also edge is becoming basically an important part of your front end, your north-south network, because your edge also needs to be dynamically reconfigurable and multi-tenant. So let's look at some examples from real-life clusters. " But it's not like that.
I just picked numbers from few customers without exposing their names. There's a customer who's running small, what is considered small, 1,000 GPU cluster. They are a Neo cloud.
They have 50 plus switches in their Neo cloud, and they have lots of tenants. So they do all this reconfiguration, this real-life live reconfiguration basically every day. And their GPUs are NVIDIA.
Their switches are part NVIDIA, part Arista. It's one example. Another example, this customer has 18,000 GPU cluster.
One cluster that has 18,000 GPUs. This is 256 racks of hardware, roughly 2,000 switches. The GPU architecture they are using that requires quad plane, meaning in their east-west network, they need four fabrics, four rail optimized fabrics, all connecting to the same servers.
So on the backend, you can think this way. The GPU's looking towards backend fabric. Each GPU has one physical port and then goes optical splitter and it connects to four different fabrics, and that applies to each and every of these 18,000 GPUs.
And then they have rack scale networking, NVL 72 in their case. So this cluster has eight fabrics. They also use DPUs.
Their tenant has this requirement that they need to use DPUs. So this cluster is using eight fabrics. This is real thing.
It's not future. These are not plans. This is actual-- Yeah?
Could you back up to the optical splitter? Yeah. I'm just curious what that means.
So, when you have multiple fabrics- Mm-hmm ... see, initially, the backend fabric was single rail optimized fabric, so non-blocking, like every GPU can send 400 gigabits to every other GPU. Now, how to deliver more performance?
So industry came up with this new hardware approach where they call quad plane or dual plane or quad plane. So quad plane means you're building four times of the same backend fabric. But on the server side, these servers are getting so dense, you don't have room to put a ton of plugs.
So the way it works, it's still a single plug, but deep inside it is using like a CDW or some kind of optical splitter that gives you multiple cables. So it's physically one plug, one SFP plug, and then you take different cables into different fabrics. Okay.
So it's using like CWDM? Yeah. Along that line.
So when you look from the server perspective, you will see multiple NICs. So if your server has four GPUs, this kind of common form factor for this latest GPUs, four GPUs per server, you'll see 16 interfaces on the server side. Like you type ifconfig, you'll see 16 interfaces looking towards backend network only.
But it would be less speed then, right? Like you're splitting out like 800 gig into four 200s. Yeah.
Okay. All right. Yeah.
Yeah. Okay. That's what I thought.
And there's example of this other cluster, 15,000 GPUs, which requires around 600 switches. It's single plane, so less switches. And it uses mix of Arista switches and some other Broadcom Sonic enabled switches and then DGPs.
So basically what I'm trying to tell here is that all these different technologies are real. We are seeing these technologies in real deployments. And this is maybe also why that we are seeing there are number of new switch manufacturing hardware companies, right?
Three hot startups right now, and they are building new switches. And people will say, "Hey, we have NVIDIA is the leader in Ethernet switch shippings, and then there's Arista as a second, and then Cisco and these others. " This is why, because AI requires a lot more than just pure Ethernet and the NVLink is actively used and the rest of the industry needs equivalent for that scale of networking.
And we're also seeing DPUs actually being used. In the beginning, DPUs were there, but they were running in a NIC mode, so most GPU operators were not using DPUs. But we're seeing that this is changing.
They are enabling DPUs, and they are willing to take advantage of DPUs, and I'll explain what that is. So in terms of hardware vendors, we're seeing a lot of NVIDIA switches, more than 73%, but we're also seeing some Arista switches, and we are seeing other Broadcom Sonic switches. At Netris we're neutral.
We work with any hardware where we see demand, and this is where we see the demand. We're in conversations with every network switch vendor that comes to your mind right now. Once there is demand, we will support them, but this is currently what demand is.
We're also seeing this trend of one AI cluster operator moving from previously they would stick to single hardware vendor, but now they're moving towards multiple. And roughly 26% of our customer base is using mix of more than-- So they have more than one switch vendor in their clusters. So I mentioned this term NAOM, right?
We like it. We invented it. We use it.
We hope that industry will pick it up. So network automation, abstraction, and multi-tenancy. It's not just a bunch of words.
This presentation explains what are some, we call bare necessities when you are building your AI cluster automation. And that actually describes what NAOM is. So it cannot just be your old school fabric manager or your old school SDN.
It needs to have these functionalities I'll explain, and there's more I'll explain following slides. So the same controller needs to be able to manage all these different fabrics, your front-end fabric, your back-end fabric, your out-of-band fabric. The controller needs to understand what is RoCE, RDMA over converged Ethernet, and what is regular Ethernet.
Same controller should also be able to understand what is InfiniBand, because InfiniBand is very common back-end. Still 40% of clusters are running on InfiniBand and the rest are running on Ethernet across our customer base. Host networking, as I explained, also becoming a challenge for GPU cluster operators.
So your same networking system should also handle host networking. In terms of DPUs, in terms of in some cases, there are some edge cases when you need to extend your VxLAN fabric all the way into host, either through DPUs or without DPUs. And also there are cases when you configure RoCE on your back-end fabric and you need to configure all the right things on the host, and it is also a challenge.
And then edge. These isolations, these VxLANs, these VRFs that you configure, they need access to the internet. So you're essentially building a cloud provider, right?
So you need to be able to provide this cloud-like network services. So, I'll explain a little bit how we handle these things, how system is roughly architected under the hood. And let's start with Ethernet part.
So now vendor should absolutely be multi-vendor, right? Same functionality should work across multiple vendors. And although NVIDIA is main hardware vendor, we see the way market is today, most deployments are NVIDIA, but we are seeing other vendors, so these other vendors are also important for us to support.
And the way our technology works, we have an agent that is sitting on top of the operating systemWe have a ZTP server that can provision, install, uninstall, reinstall, upgrade, downgrade your switch operating system, because we're talking thousands of switches. These are critical functions. And we have this agent sitting on top of the operating system, which means that as many switches you have, that many agents you are running.
And if some hardware architecture doesn't allow for running an agent inside the switch, that's okay, we can run outside too. But most importantly, this agent system, it's a distributed system where you have lots of parallel processes, because this is the only way to make your network automation technology scalable. You cannot have this old kind of Ansible approach, where you have all logic generated in one centralized place and kind of trying to make one centralized something tell thousands of nodes what to do.
That doesn't scale, and that's not what we do. The way it works, our controller only has some high-level information, and I'll show you soon, and we have these agents that are distributed across your infrastructure, and each agent has this proprietary logic that knows how to translate your desires that are stored in the controller into actual consistent and validated configurations on every switch in an orchestrated way. Similar to that, we work with DPUs.
NVIDIA has this technology called DPF, which is basically a ZTP server for DPUs. It helps to install operating system, but it doesn't help with the configuration, with the day to everyday configuration changes. And NVIDIA has this framework called DOCA, which is a library that sits inside the DPU, and it basically provides APIs and a framework for something like Netrist to provide configuration for DPUs.
Again, it's a distributed system. We have our container sitting inside the DPU. We form EVPN BGP adjacencies between DPUs and your leaf switches, and that way, DPU becomes an actual member of your switch fabric, and then we can provision VXLANs and VRFs, whatever is necessary across the network.
We also need to integrate with InfiniBand, right? Because your back-end fabric can be Ethernet, but it can be InfiniBand. And InfiniBand as a technology, it's more plug-and-play-ish technology, and there's not a ton of day zero work that we need to do, so we don't have a ZTP for InfiniBand because NVIDIA has this very nice controller called UFM, which takes care of day zero, day one, but that thing, UFM, it doesn't know about your multi-tenancy.
So when you onboard, offboard, resize tenants, NVIDIA UFM doesn't know about that. So someone or something needs to be in synchronization. That's what we do.
Our controller and NVIDIA's UFM controller, they are integrated, so we automatically discover the GUIDs and partition keys out of NVIDIA UFM, and when customer creates their tenants, we automatically push appropriate partition keys and associate right GUID. So we add that multi-tenancy functionality to InfiniBand. Then comes NVLink, which is NVIDIA's term for rack-scale, scale-up fabric.
Now, NVLink used to be a bus inside the server connecting eight GPUs. But today, it made its way out of servers into the rack, connecting 72 GPUs. And to do that, they added nine switches.
Not Ethernet switches, but different kind of switches. And by the way, there are companies who are trying to solve the same problem through Ethernet, which can be a potential path forward. Anyways, NVL is something that is rack-scale networking, nine additional switches, and by default, all the 72 GPUs are talking to each other.
And as a AI cloud provider, as Neo Cloud, if you want to have this ability to split the rack, maybe you want a tenant, you want to have this ability to provide four GPUs to one tenant or eight GPUs, right? How do you do this when 72 GPUs are always connected? So you need to have NVLink partitioning.
That's what we do. We have integration with each rack. There is no rack-to-rack connectivity.
Each rack is isolated nine-switch network of 72 nodes. So we have integration where our controller talks to each rack, learns everything about GPU IDs and existing partitions, and when you create tenants, not only our controller reconfigures your north-south and Ethernet network, but also carves out isolation on your NVLink side of things. That way, we guarantee that your back end, your front end, your rack scale, everything is aligned with your isolation rules.
So to summarize, how NAM is different from SDNs or intent-based networking. NAM covers networking end to end, and by networking, we mean networking for AI use case, for Neo Clouds, for AI factories. NAM is something that helps network engineers.
Network engineers are very important for AI. It's not like-It's not like AI is coming to take network engineers. It's not like that.
AI needs network engineers. Network engineers make AI possible because AI needs a hell lot of networking. So it's a tool for network engineers that helps them from day zero, day one, and day two.
So they design their network through this tool, they simulate their networks, they deploy, they do validation of these massive fabrics. They automate troubleshooting of these massive fabrics on day two, and they perform maintenance across these massive fabrics. And it also provides all the right APIs to others, to consumer side.
Because unlike traditional networking, it's not like network engineering is one camp and compute people are another camp. It doesn't work like that in AI. Everything needs to work together because there's no SDN.
Compute folks cannot use software networking. Compute folks have to use physical networking, and your NAM controller is the only thing that is connected to all these physical networks. It's the only thing that connects your desires, your intent, with actual network engineering constructs.
So- Alex, real quick. When you say it covers end-to-end, that includes, and you probably covered this in a slide that I'm forgetting, but that's front and back-end networks as well, and automation. That includes north-south or front-end, same term.
East-west or back-end, as many back-end networks as maybe one, two, or four, whatever, maybe in the future, eight. Who knows, these AI- Right ... people.
NVL, edge networking, so your NATing, your load balancing, because that thing too needs to be aware of your multi-tenancy. Right? You can have overlapping IPs in your tenants.
Your DPO networking, because it's like a switch in every server, and you have a lot more servers than switches. So if you have 2,000 switches in a network like that, you have roughly 4,600 servers, that many DPOs. And all that needs to be orchestrated.
And even NIC networking, sometimes there are some edge cases that networking on the host itself needs some additional help. So consider this very simple workflow, common kind of generalized workflow. Somewhere in any AI cluster, there is this PaaS, IaaS layer.
It can be a commercial product coming from someone, it can be homegrown something. But there is some kind of system that is responsible for managing compute, servers, containers, virtualization. It may have customer-facing portal, it may have connection to billing system.
Very important thing. It helps Neo clouds make money. But this thing, unlike traditional networking, remember VMware had its own SDN.
But that model, after acquisition of Nicira, I love Martin Casado, but it's just that that technology is not for AI. Mm. Because these systems cannot use software networking because of the enormous amount of traffic.
So they have to collaborate with something, your NAM controller, that is in charge of your physical networking. And there has to be this API-level collaboration or integration. So consider this.
These logos at the top, these are compute orchestration or platform companies, products. Some of them are open source. Many of them are commercial products.
These are our partners. And we are grateful to NVIDIA for connecting us to these amazing partners and leading us the way to build very powerful partnership with all of them, starting with NVIDIA Nico, which is this open source product that recently became open source, and it's kind of gaining popularity. Or vCluster, Rafay, Spectro Cloud, Red Hat, Mirantis, or CloudStack, which is another open source product.
These are different technologies that Neo cloud builders are using to orchestrate their compute, install, uninstall, operating system, Kubernetes. Very important, but these things depend on networking, and these things have little to no network orchestration skills. And you know how marketing happens to technology, right?
Sometimes there is alignment, and oftentimes there is misalignment. And especially the bigger is the market, the more marketing misalignment you might see. " So this presentation is kind of guide to network engineers to help them evaluate their Neo cloud orchestration stack.
It's very important that you do networking right, and it's important to distinguish what network configuration can compute orchestrators do and what... kind of network orchestration you need from your NOM platform. So your networking, because it's so complex, it needs to be really, really specialized and I'm glad that our partners and us, we are in alignment on this.
And it's been a lot of work for partners and for us to build all sorts of API integrations. So basically we both can go to customer and say, "Hey, these companies, this product, and Netris as an NOM work together. " And now we got to part two of this.
So, I'm Alex, I'm the CEO and co-founder of Netris. I just finished part one of our network automation, abstraction, and multi-tenancy presentation at Networking Field Day 40. It's a presentation where we share from our experience of deploying and helping many large-scale GPU cluster operators to deploy and automate their networks.
And part two of this presentation is about life cycle of the networking. Now, in enterprise data center networking, there was some life cycle, but here there's a lot more of a life cycle. It's very nuanced in a sense that there are two things.
One thing is these networks are very complex, and obviously deploying these networks is more difficult. And there's also this other factor. See, the way this AI business works, you're never designing one cluster.
If your company is sustainable in AI business, your company is going to be deploying more and more GPUs over time. Newer generations of GPUs are being announced, and then your company is buying those, and when you're architecting a system, it's not for one cluster. You're architecting something for long-term strategy.
That's why you're going to have a lot of life cycle. And other nuances are these valuable GPUs, when they are shipped to the data center, you cannot play with them for a month or so till you figure out and you go live. No one will let you do this because they are so valuable that the management of the company you work for, they want these GPUs to go live immediately.
The moment the power is enabled, they want to go live and start making money. So this means that your preparation work, your figuring out, your learning should happen early on before they arrive. And sometimes even your configuration troubleshooting and validation should happen before.
Now you're probably thinking, "Alex, what are you talking about? The switches are not in my data center. How do I validate configurations?
That's crazy. Doesn't make sense," right? Unfortunately, this is how it has to be done, and this is what we've been doing with our customers.
So we do a lot of simulations. We use simulations, very powerful approach. But before simulations, we do modeling, and modeling is source of, it's like source of truth, but sort of next generation of that because it's slightly more than source of truth.
And please don't think switch configurations. There's no configurations in modeling. Modeling is information about your topology, what switches you're going to have, what servers you're going to have, how all these cables are going to be connected, some IP addresses.
But why I'm saying this is beyond source of truth, because your services, the services that you're going to create, the VXLANs, VRFs, VLANs, and cloud constructs that you're going to create, all these things will be in the modeling. So what happens on day zero, once you order the hardware and the company's waiting for hardware to be shipped, and then racked and stacked and powered on, network engineering team is working on the preparation. Oftentimes, it's a collaboration between customer, us, and system integrator or Nvidia or whoever your GPU vendor is.
Oftentimes it's a collaboration. And what we do during that process, we model the network. We know what is ordered, and how are you going to wire these networks is known in advance.
We model that, and we put that into Netris controller running somewhere, and then Netris controller can create a simulation. So think of this. Controller knows what is your master topology, what is it that you are trying to build.
But we don't have hardware yet to bring up that network and see whether these BGP sessions come up or not. Alex, hi. Scott Robohn.
Is it okay to think of this like ContainerLab or other IP-focused virtual-only environments? It's slightly different. Eh?
So we use two technologies for this. Nvidia has technology called DSX Air, which is a place where you can, through API, you can simulate infrastructures. But we also have built our own.
It's called Cloud SIM. It's not a product, it's used internally. " So we had to build a technology for us, at least to make our product work in the first place.
I notice you haven't called it a digital twin. Well, it's commonly called digital twin. Okay.
Twin-ish. I'm not trying to lead you anywhere, but okay. Yeah.
Yeah. Sometimes industry calls this a digital twin. Okay.
So the idea is that you have model in the controller. Once you have physical hardware, controller can program your physical hardware and bring it up. " You put their credentials and your simulated environment comes up.
And then maybe network architects and engineers want to look through it, discover some errors maybe, or sometimes it's like, "Oh my gosh, we forgot about our upstreams. " And then with that pre-validated model, you are starting to go live. " And your hardware is there.
Now you need to bring up this, maybe hundreds or maybe thousands of switches. They are powered on, but there's no operating system. So as hardware is being correct in stack, someone needs to copy the MAC addresses.
Until you open the box, you cannot see the MAC address. When you open the box, it's written on the switch. It's how it works today.
In the future, we hope we're working on this, that when factory ships the switches, we will know the MAC addresses for each switch and its precise location, but it's not there. But it will happen. Today, someone needs to put that MAC address manually, and MAC address is important for the ZTP server to identify the physical box with logical box on the master topology.
Alex, quick question for you. Jason Gennert. As far as topologies, do you have best practices, kind of validated design topologies that are preloaded into your product, or do you have to build out your topologies by hand?
How does that work? Yes. We have that.
You have both? Yeah, we have both. So, there are reference architectures coming from vendors, right?
Like Nvidia has a number of reference architectures attached to each kind of GPU family. We have some of those, and when we implement actual customer deployment, sometimes the real deployments are slightly off compared to that prescribed reference architecture, and that's fine. They share with us what is it exactly, and we tweak the topology, because Netris can work with any topology.
But what if I haven't ordered it yet, and I don't know what my final topology is going to be? This is when, like what you just said, we have these pre-built topologies, which are basically 100% aligned with Nvidia's array. Which is like, hey, you don't know what exactly you're going to do?
Hey, take this example array, which is Nvidia array, build it, launch it, bring up a digital twin, and learn. So this is used for trainings, for demos, for testing things, for two vendors to integrate together. So like, yes, and that's very actively used.
Okay, cool. Another question for you. Do you integrate with a source of truth to pull in variable data for those topologies?
Like a NetBox or something along those lines? Yes. There's multiple ways.
There's no magic button between Netris and NetBox. Okay. But oftentimes, this data is coming from various places.
Sometime we have to work with an Excel file with P2P links, and we have a way to load in Excel file of that. With NetBox, some customers use NetBox, and what they do, they would write few lines of code between two APIs, and then kind of that code will load something from NetBox and insert into Netris. But inside Netris, you need to have your Netris-specific source of truth, your IPAM, your topology.
It has to be in Netris. And some customers absolutely combine that with NetBox, but we do not have that nice, beautiful button. Okay.
Yeah. Thank you. But still, it's easy.
Alex, I have a question. Phil Gervasi. Do you also generate simulated traffic, GPU to GPU traffic, so that way you can test failure scenarios, kill a link, take a spine out, change buffer keys, things like that?
Yeah, great question. Yes. It's very limited today.
Okay. It's lots of pings, like that. But I think there's...
I don't know if I can share their name, but there's a vendor who specializes in that, basically testing. Well-known vendor. And they are working on integration with this kind of simulation technology, so to simulate more realistic traffic.
So you can do an outright validation that way- Mm-hmm ... of performance measuring. Yeah.
Okay. Yeah. And of course, I think probably primary for a lot of architects is isolating hotspots and failure scenarios- Mm-hmm ...
more than anything else. So that's interesting. I'm going to talk about validation in my next slide.
So validation, to your point, is big, important part. And what you said, some of that exists today, some of that will come later. So how we do today, so when this large network goes up, so what you see here, these different colors, so number one is Netris master topology in the beginning.
Switches are offline. And you start with that, and switches are starting to get operating system through ZTP, and then Netris agent loads, and it starts automatically configuring switches, and not only configuring, but also monitoring and validating. So, again, there is no such thing as switch configuration in Netris controller.
Network engineers do not write configs in our world. They just provide master topologies, and they provide policies, but not configs. If you replace a Cumulus switch with Arista switch, you're not going to translate Cumulus config into Arista config.
Algorithm will take care of that because there was no config to begin with. It was always generated. Now, these hundreds or thousands of switches get auto-generated configurations, and because algorithm knows that I just configured BGP on this link, and I'm expecting this link to be connected to that link, because algorithm knows what to expect, algorithm is able to do lots of testing.
So for example, there's going to be a lot of miswiring, and algorithm will say, "Hey, this cable was supposed to be connected to port five, but it's connected to port six. " And there's number of things that, for example, when NVIDIA deploys these clusters, they have a protocol of things that they test, and we have implementation of all these things. " And that's kind of we work with NVIDIA and the customers.
And then comes troubleshooting. With respect to troubleshooting, there are kind of two approaches, troubleshooting from the perspective of each node. Remember we have agent on each node, and this agent knows what controller knows, so this agent knows who else this switch is connected to, and who are other loop back IPs in this fabric.
So at least it can execute some pings. So, these are techniques to troubleshoot things. You cannot misconfigure this network because it is automatically configured.
But what can happen, sometimes control planes can break somewhere, or some hardware can be stuck in some weird situation, how you discover things like that. And when you scale your network at a larger scale, you come across things like that more often. Sometimes there's one switch, it looks good, power supplies are there, fans are spinning, ports are up, BGP is up, everything.
But something is off. So there's kind of number of tests that we are doing to help network engineers identify kind of weird situations like that. And that part keeps growing.
With each release, we keep adding more and more smart tests. Okay. And now part three is about consumption, how we consume the network.
So, I'm Alex Oroan. I'm the CEO and co-founder at Netris. I'm presenting at Networking Field Day 40, and my presentation is about network automation, abstraction, and multi-tenancy.
We call this NAAM. It's for AI factory operators for NeoClouds, and this session is to help network engineers evaluate solutions and designs and architectures when they are building NeoClouds and AI factories. And part three of this presentation is about consumption model.
So these networks that we all are building, they exist for a reason. Someone needs to consume these networks. And that someone can be the cloud operator, or that someone can be a compute orchestration product that's talking to network controller over API.
And lookCloud been around for a while and over time, 10, 15 years ago, cloud introduced some abstraction model, VPCs, subnets, VPC peering, elastic IPs, elastic load balancers, and time proved that it's a very useful model. Now this AI cloud operators, people on consumption side, not network engineers. Network engineers need to do their job, yes.
But part of network engineer's job is to enable the consumption. And consumers don't want to consume switch ports or VLANs or VxLANs. They don't want that, especially given that average GPO server has between 10 to 40 connections.
Really difficult for a consumer to manage stuff like that. So we're using a model that is very similar to cloud networking abstraction, VPCs and endpoint and things like that. So I'm using this diagramming method here where from one side we have these VPCs as a cloud construct.
When someone is using the cloud, they don't see all the switches around it, but they do exist. And as a network engineer, you are responsible for that switches. While your consumer is responsible for consumption, you are still responsible for that switches.
So the very, very basic concept here is the VPC, right? " So they want to have this ability to create very, very simple request. Now, what you're seeing here, it's a screenshot from our web interface, and obviously everything can be accessed also through REST API, which is how it's usually done, or through Terraform.
We understand Terraform language. But this is just visual example to keep all of us on the same page. Now, this is how you tell the system, "Hey, I want these two hosts, these two GPU servers to be isolated together.
I'm going to give this to my first tenant. " Now what needs to happen underneath, the system knows what your topology is like, what your Ethernet is like, what your backend Ethernet is like, RoCE and front-end Ethernet, you configure differently. The way you configure VxLANs in some place use layer 3 VxLANs, in some place you use layer 2 VxLANs.
Maybe you're using InfiniBand, but all that is not consumer's business. They don't care. And your network automation needs to take care of that.
So what happens? All these algorithms know what exactly to configure in front-end fabric, backend fabric, and your rack scale NVL fabrics. And then in some time, this isolation is enforced and now these two servers that belong to alpha can talk to each other, but like this other server that belongs to bravo cannot talk to each other.
So VPCs are in isolation. Obviously in Ethernet, this achieved by using VRFs and usually L2 VxLANs, L3 VxLANs are also supported. Using partition keys and GUIDs if it's InfiniBand, and using GPU partitioning, if it's NVL.
I'm actually going to show how NVL works in a demo soon. So VPC is this very, very foundational construct, and then there is this other construct we call VNet. AWS calls this segment.
Essentially VNet-- Sorry, AWS calls this subnet. And VNet is basically Netris term for universal segment. It can be layer two segment, can be layer three, can be VxLAN, can be VLAN.
It can have a default gateway or doesn't have to have a default gateway. So it's a network engineering construct. So back in the day when we all were working in data center networking, there was no VNet.
VNet was happening between my compute, my Linux administrator colleague, and me. He would come to me and say, "Hey, Alex, we're deploying this new five servers for storage or whatever. " That was the VNet.
Today, it's a construct in a self-service portal, which is important construct when you are building a cloud because these things, you still need to service that colleague, but now you need to service them programmatically and across different fabrics, not just Ethernet, but also InfiniBand and these other fabrics. So, that's the idea of VNet. And the VNet can be created individually one by one, or when you are doing this thing, so this thing is basically an object with two servers.
You can tell Netris, you can create templates inside Netris, and you can say, when my user is creating this kind ofsub-cluster, I want Netris to create these five different VNETs automatically behind the scenes. So this is mechanism that kind of connects network engineering world with the AI cloud operator world, and that's in the hands of network engineers. Network engineers define this translation.
So VNET and VPC are these very important constructs, and when you configure VPC, VPC becomes a VRF from Ethernet perspective, and when you configure a VNET, you basically turn certain switches into VTEPs. Now, obviously, Netris' algorithm knows which servers are connected to which switches, and we do not go and replicate all these VRFs and VTEPs. No, we only configure VTEPs and VRFs where applicable.
We take care of your TCAM and we don't waste your TCAM. It's bad. Wasting TCAM is bad.
These GPUs are all big, but TCAM is still limited. Same technology, just more throughput, but same technology. And now what if running a VTEP on a switch is not enough?
What if you want to split a server? Because listen, when VTEP is on your switch, when you enforce this tenancy and VxLANs on the switch port, which is very important use case, your smallest unit of isolation is one server, right? Your tenant has access to server but not to your switch.
So switch is this safe place where you can enforce their isolation. But what if this cloud that you are building, this new cloud you're designing, needs to provide also granular multi-tenancy, splitting a host, virtualization. And you don't want to do this virtualization through using VLANs and sub-interfaces.
You can do that. We do support that. But that is not as performant as GPU operators want.
And if you're designing for lots of inferencing traffic, there is this thing called KV cache. It is very bandwidth-intensive use of north-south interface, a lot of traffic exchange between GPUs and storage. And you don't want to use VLAN sub-interfaces for this.
You want to use DPU-accelerated isolation, so that VLAN or VxLAN sub-interfacing has to happen on DPU level. Now, so we came to sort of intricacies of host networking. What are some networking things that needs to happen on the GPU host itself?
So one very simple thing, and these things are kind of optional. We have some tools here. Some clusters need these tools, some clusters don't.
But my job here is to explain what is for what, and if you're designing cluster for AI, you'll know which to use or you'll know what exists out there and maybe you'll reach out and we'll help you further. Anyways, host networking, this is your server, the GPU server. It has two interfaces, eth0, eth1, looking into north-south front end.
And this is no DPU yet. DPU is running in NIC mode. Hardware is there, but DPU pretends it's a regular NIC.
So two NICs looking into north-south and eight NICs looking into east-west or your back end. Basically, one NIC aligned with each GPU. So they are physically connected and traffic is not going through the CPU, but traffic entering this NIC is directly going into the GPU.
Nvidia calls this technology GPU Direct. It's like RDMA-level technology where NIC talks directly, can write directly into memory of the GPU. This is important for RoCE.
So your north-south fabric, your east-west fabric, back end fabric, is running RoCE, which requires certain configuration on the switches, which is in Netris is one checkbox. You enable that and algorithm knows, "Oh, okay, it's RoCE fabric. " And now your switches act the RoCE way.
But you also need to configure IP addresses and certain routes, certain way on host side. And sometimes it's eight interfaces, sometimes it's 16 interfaces. Sometimes you want to write a script to do this.
Sometimes you want to use vendor's plugin to do this. We have a plugin, very optional. What it does, the nuance here is your host cannot have access to your network controller.
This is your tenant's machine. Your tenant has root access on the server. So this plugin needs to learn some information that only your network controller knows, but without getting access to network controller.
So what we do, we have this optional functionality where you can expose some insensitive information over LLDP, so your switch can signal the server over LLDP, sending custom TLVs that this cable, this rail should use this IP and this static route. So this is basically IP auto-configuration and routeRouting protocol over LLDP that custom created by us. Some customers use it, some customers just use some custom scripts to configure that IPs.
Yeah, we call that NHN, Netris Host Networking plugin for east-west. And then comes this other thing we call EOH, EVPN on host. Now, sometimes you have servers that are not DPU-enabled, but they need to become part of your EVPN fabric.
And you want to expand, you want to extend a VXLAN all the way into host. Why that might be the case? So imagine you're building a larger cluster, you want to support thousands of tenants, and you cannot use VLAN sub-interfaces because there's no enough VLANs.
You want to go beyond limits of VLAN. So what you can use is VXLANs. But how do I make my server, my host, to kind of understand VLANs?
How to make it part of control plane. So we have this plugin, you install on the server, and it will create bridge interfaces on the server, and these bridge interfaces will be end points in Netris. And you can say this virtual bridge endpoint and that switch port, physical switch port, should be together in the same VXLAN.
The typical use case for this is when you have some servers for control plane of Kubernetes or control plane of something that you manage for your customer, your tenant, but it's not the GPU host itself. Because GPU host has a DPU. If you want to do that granular multi-tenancy, if you want to do that extension of VXLAN into server on a GPU host, you just enable DPU, and you do what I just explained in a hardware-accelerated way.
So how we help here, we have auto-configuration of HBN. Nvidia calls this HBN, host-based networking. The idea here is that the DPU and leaf switch, they form EVPN BGP adjacencies and start exchanging EVPN routes.
And that way, the host stops seeing your regular NICs. These two NICs that were previously visible to the host, they disappear. Instead, DPU shows up some additional kind of simulated NICs, but they are hardware-accelerated.
You can push 400 gigabits of traffic through these so-called virtual NICs. And you can have many of those. You decide how many, like eight, 16, whatever.
And then you can connect this, they are called VFs, virtual functions. You can connect these VFs to your virtual machines, to your containers, and now you have different containers or different VMs connected to a hardware-accelerated sub-interface, which is first-class citizen in Netris controller. Then you can say this VF and that switch port, they should talk to each other, and that will be 100% hardware-accelerated at wire speed.
Now, there's a method of doing this DPU overlay thing, okay way, and there's a method of doing this really, really right. Now, sometimes there's this misconception of like, hey, these DPUs are like switches. Let's run BGP adjacencies between DPUs but not within the switch fabric, and let's run an overlay between DPUs.
Using the switch fabric as pure transport. Well, technically that works, but it leads to some problems such as, in that case, only DPU-enabled endpoints can be your VTEPs. You can create a VXLAN between any DPU-enabled servers, but what if you have non-DPU-enabled servers on the cluster?
What if you have a cable coming from another data center and your tenant needs to read a lot of data, which is very common use case. What if you need to connect your edge gateway to provide NATing or load balancing functionalities? So this approach of disconnected DPU overlay, it's a problematic approach.
We do not recommend doing this because it leads to workarounds after workarounds and after workarounds, and that does not scale. So the way we recommend doing this and what is our default way of doing this is what we call complete BGP EVPN integration. Meaning that your switches, they all talk to each other EVPN BGP as well as they talk EVPN BGP to your DPUs.
Meaning, DPUs and every switch in north-south fabric, they exchange EVPN routes, and this means that you can create hybrid VTEPs. So basically, a DPU can become a VTEP for one service, and switch can become a VTEP for another service. Make sense so far?
This part? So that way you can... Yeah?
So the VTEP on the server is doing my encap and decap for the tunnel? On the DPU, yeah. On the DPU.
Okay. So we're not doing it on the leaf, I mean. Doing both.
Doing both. Both. DPU on the server has its control plane shared with the leafSo any route that DPU learns, your switches learn that route too, that prefix.
So that allows you to create a VXLAN that has combination of switch ports and DPU virtual functions. So it can combine a virtual machine and a bare metal machine together. Okay.
Like in this example, see? There's a server in VPC Bravo that is running in a NIC mode. There's no DPU.
But it can become in the same VPC together with a DPU-enabled host. I have more examples coming soon that kind of gives more use cases for this. So, in this section, we are talking about how we provide, how we deliver this consumption model to our cloud operators, cloud consumers.
And the common question is how do we as a Neo Cloud, let's say you're building Neo Cloud or AI factory, how you organize shared services? And a common example is storage. We have some storage system.
We have our different tenants, tenant Alpha and tenant Bravo want to have access to storage, but they should not be able to access each other, should stay isolated. So storage is usually a group of servers. You put them into an isolated tenant.
So storage machines need to talk to each other, synchronize data, whatever, and how to make other VPCs consume that storage. So we have this construct we call VPC peering. Our job here is to help network engineers deliver that cloud consumption model so their AI cluster can be consumed as a cloud from network perspective.
And the construct that helps here is VPC peering. It's a simple construct where you say VPC A should have access to VPC B. It's kind of easy for consumer to set this up, and from, like in this diagram, this bottom part kind of shows what is logically configured, and top part is showing what is actually being auto-configured on the switches.
So in this case, Alpha has VPC peering with storage. Bravo has VPC peering with storage. And on the switches, our algorithm will automatically configure VRF leaking.
So, VRF4, which is VRF for storage, will receive routes towards Bravo and Alpha, but Alpha will not be able to access Bravo, or Bravo will not be able to access Alpha because if traffic leaving Alpha entering that VXLAN of Alpha's trying to get to Bravo's IP addresses will not be routed because that VRF has no idea how to route. So ASIC will never route that traffic. So it's safe, it's easy to implement, easy to consume, easy for you as a network engineer to implement when you have tool that does this, and it does the job.
And this also works if you have kind of mix of DPU-enabled hosts and non-DPU-enabled hosts because we always combine control planes of DPUs with leaf switches, and so like switch or DPU, we kind of treat them the same. There's this example of direct connect. So, what happens if your tenant has some data located in some remote site and they do a lot of data exchange?
This AI is a very data-intensive business, so when you serve these AI tenants, you'll realize that they are exchanging a lot of-- So like, your tenant willing to send and receive two terabits of data in and out of your data center. That's like a normal ask. And how do we organize this?
And let's say they have this remote site where they have a ton of data, and there's a cable coming from remote data center, and this is not one time thing. This is not like, "Hey, we'll make it work," and whatever. It's something you want to enable your cloud operators to be able to do this easily, programmatically.
So to organize that, we have a construct. It's a BGP construct. It is used also for defining your uplinks when it comes to connecting with your border router.
But in this example, we're using BGP connection for forming a BGP connection between a physical switch port and the VPC. So this is similar to what AWS calls AWS Direct Connect. When you have a VPC inside AWS and you have a cable coming from AWS to your data center, you can configure a BGP on your router between your router and your VPC, right?
Now in this example, you are the AWS. You are building AWS. So this is how.
This construct allows doing that. It's the same construct you use for configuring your upstream connections between your leaf spine fabric and your border routers. With the only difference that in that case, you would configure this BGP session between your border router and your infrastructure VPC, your own, your network engineering VPC.
But in this case, you are configuring this for the VPC of your customer. " Okay, one more thing. So we talked about how we configure lots of things across switches, high performance data, terabits of traffic between storage, GPUs, and how we isolate all that on a hardware level.
And then there's still one missing question. How do we provide connectivity between VPCs and internet? So AWS, again, to use this analogy, AWS has these functionalities they call Elastic IP, which is a construct that allows to map one public IP with one private IP.
And Elastic Load Balancer, which is basically mapping one public IP with a group of private IPs. And if you as a network engineer, as an architect, you are building AWS, again, you need to be able to provide this functionality to your tenants. And to help with that, we have built this technology we call Netra Softgate, and it evolved since last time we presented here.
So now Netra Softgate is horizontally scalable, and it's multi-tenant. What functions it provides, look, it provides three basic functions, NATing, layer four load balancing, and DHCP. " The challenge with this is that when you try to deliver these very basic functions, in a multi-tenant and horizontally scalable way, that is very difficult.
Most devices that are designed to provide NATing and load balancing or DHCP, they are designed for data center use case. So you can have big, powerful devices, two of them, working in HA cluster, but most of them are not horizontally scalable. They are not designed for cloud provider use case.
And they are not multi-tenant in a sense. Yes, you can configure VRFs there, but how Softgate is different from multi-tenancy perspective, it fully integrates with your front-end fabric. It becomes part of the fabric.
Switches think that Softgate is another thing versus Softgate in reality is a server. You dedicate some servers for this function, and if we look under the hood of AWS, they do something very similar. They have their own proprietary accelerated software gateway for these functions.
And we realized that if that approach works at AWS scale, maybe anyone who's building AI clusters, in getting closer to AWS scale, they will need this technology. And we created this technology, which is XDP accelerated software. So you dedicate several servers.
They connect to your leaf switches. They form EVPN BGP adjacencies. That way, this cluster of Softgates can tap into every VRF, into every VPC, obviously according to your instructions.
And it can scale horizontally. So you start with two or... Sorry, you need four.
Four is the minimum. You start with four, and then as you monitor utilization of CPUs and RAM, and you add more Softgates, and it will automatically rebalance the traffic. Now, delivering these functions, NAT, layer four, LB, and DHCP in a multi-tenant and horizontally scalable way is really, really, really difficult.
This is why Softgate exists. And look how it works. So from a network engineer perspective, you just deploy several Softgates, and these are your kind of data planes.
And then consumer starts creating load balancers, which can be TCP, can be UDP. It is VPC-aware. They pick some front-end IPs from a pool that you define.
They can do some health checks. Very importantly, we have our proprietary implementation of Maglev algorithm. This is very important when you're building scalable systems because think about it.
These Softgates, they don't share stateYou're not allowed to share a state. When you're building cloud scale something, you cannot share a state. It doesn't scale, and it becomes very challenging.
So we implement Maglev algorithm that helps us with layer four load balancers, like different softgates, unaware of each other, are load balancing same traffic destined to different servers, and they all end up with the same consistent load balancing. There's additional things that we do to make sure that NATing works correctly. Again, a lot of math and a lot of creative algorithms involved that we implemented on low level using XDP and C programming to deliver the right performance.
This is example of how you configure NATing. So this is a VPC-aware NAT. You can select your VPC, you pick your public IP address, you pick your private IP, and then traffic coming from the internet passes through your fabric.
So this red line is load balancer traffic. It's coming from the internet. Softgate is like a router on the stick.
It's connected to a few leaf switches. It's not connected to your border router directly. It's connected to border router indirectly.
So the traffic physically passes through the fabric. It comes into a softgate. Softgate decapsulates this traffic, and softgate knows, oh, okay, this traffic belongs to VPC Alpha.
And if Alpha and Bravo have overlapping IPs, that is fine. Softgate knows how to handle that. And then it uses Maglev algorithm to decide whether this packet goes to server one or server two, and then checks health checks, and then encapsulates with the right VxLAN and sends further into fabric.
Similarly to that, NAT flows work, they arrive at softgate and destination IP or source IP is changed, and then NATing happens. Quick question for you, Aleh. I'm sorry if you covered this.
Where does the softgate live? You dedicate several servers for this function and connect to some leaf switches. Okay.
Usually, you would connect to a border leaf, but doesn't have to be there. Any switch is fine. Okay.
And what is it? Is it a container, VM, any of the above? No.
It's the whole server. Oh. It's not containerized- Bare metal server.
Bare metal server. Got it. So multi-tenancy is not implemented through running different VMs.
That's not performant. The data plane that we developed, it's low level, like XDP, literally in low level C++ code. Got it.
That data plane is able to deal with overlapping IPs, and it understands VPCs. This is what makes it performant. Using a bunch of open source components and containers, that will work for a lab, but that will not be performant.
We've tried that. You want cloud scale performance, you need to deal with low level development. All right.
Thank you. And then with providing all this access between VPCs, between internet, between direct connect, with all these things, maybe you're asking, okay, how do we control traffic? Maybe when I let my tenants to access my storage, maybe I don't want them to be able to SSH to the storage.
Maybe I want them to access certain ports, right? So, yes, obviously, there is a construct for providing VPC-aware ACLs, which is on consumption side. It's a simple construct where you say, this is my VPC, this is my source IP, destination IP, these ports I want to permit or deny.
And what algorithm does, see, we're talking about hundreds or thousands of switches and lots of DPUs, and where to insert these rules. If algorithm inserts these rules everywhere, just in case, we're going to run out of TCAM very quickly. So we're not doing that.
The algorithm is smart enough to continuously, on the fly, analyze your routing tables in each VPC, and dynamically inserts all these rules in the right places. So if you have ACL that is, let's say, is trying to permit or deny traffic coming from the internet through Softgate, behind elastic IP, for example, that enforcement will happen on the place that is closest to the entry point on the Softgate itself. But when you are trying to access controlled traffic that's going between two switch-terminated endpoints, let's say between storage and the servers in the VPC Alpha, then you want this enforcement to happen on switches in the hardware and not degrade your performance.
So in this case, we automatically apply this on switches using TCAM. That's the only way. But apply on the DPU.
That way we save TCAM usage. We reduce TCAM usage that way. So where possible, we would like push away these ACLs to minimize TCAM usage.
Where not possible, we'll use TCAM. So, I walked you through what we call bare necessities for building network automation for large-scale AI clusters. We call this NAAM, Network Automation Abstraction and Multi-Tenancy.
We created a checklist of all these things, of bare necessities. So there's a link. You can go through that link, put your email and information, and then we will send a checklist, and this checklist will help you kind of go through what are the bare necessities for building AI cluster network automation.
And if you ever have questions or have interest in the product, please reach out to our sales team. They will connect you with solutions architects. We'll provide access to Cloud SIM so you can try the technology.
They will help you to figure out whether this technology is right for you or not, and will help with throughout your AI infrastructure networking journey. Okay. And in part four of our networking field day 40 presentation, I'll show some demos of some of these things that I just covered.
We would like to see. Yes. Okay, let's do that.
I'll grab a chair. Okay. So look, this is Netris controller.
This is the day zero situation. Hardware is ordered, it's in the process of delivery, and we're waiting on hardware, so we're modeling something in the controller. This is empty controller and, see inventory, topology, everything is empty.
And the first step is to model, to put some initial information into the controller. And, so to do that, I'm going to use this Terraform module because Netris understands Terraform. And we can insert this information through different ways.
We can ingest a CSV file. But, when using Terraform, this is kind of a predefined topology based on one of the common use cases. So this one is for two-tier Ethernet-based switch fabric.
So in this case, we're going to generate infrastructure for 64 GPUs. So it will derive the actual number of switches and cables and everything from this number. So let's put 64, and here we define some IP addresses for loopbacks, for management, and it will be picking IPs from these ranges.
And that's for east-west, that's for backend network. And this part will define different parameters for front-end network. So we'll, so basically, we need to define how many leaf switches we need, how many spine switches we need.
We define the things, and then we apply. So this window here is SSH terminal into the same controller, into basically this same controller. So this Terraforming process is happening on the same machine.
This module is calling Netris APIs to create different objects. Everything in Netris controller is object-oriented, so we're creating objects. So we have created this topology, and we're going to simulate this topology using Netris Cloud SIM.
This is all Netris constructs. You're not generating device configuration here yet. Not yet.
Okay. Not yet. Yes.
Very good point. I just started the simulation. It takes five minutes, so while it's coming up, I'll walk you through- Okay ...
what I just did. But you're right. I just created only Netris constructs, no configuration.
So controller at this point doesn't know... And in general, actually, controller never knows about configurations. It's not controller's job.
Yeah. That would depend. No, it actually doesn't know about configurations.
I believe that's how you've constructed your system. Yes. Yeah.
So what did we create? So we created some inventory entries for each server, which is basically the name of the server, and other parameters are just left to default values. There's this custom field, which is kind of a data storage.
If you're using InfiniBand, we would auto-detect GUIDs and some information from InfiniBand fabric, and we will store that information here. But for current purpose, this is not important. And then we've created some switches, because we do more stuff with switches.
We define more things with respect to switches. So we have a switch name, operating system name, so system knows what operating system to provision. AS number- Hey, Alex, quick question for you.
I noticed in this screen, you've got IP address, and I've seen in other screens where you had IP family. Mm-hmm. How does IPv6 support look?
There's certain parts of it that support V6 or across the board? What does that look like? It's fully supported on services side.
Okay. Oh, so you mean like gateways, like NAT and- Yeah ... like the layer three services.
All that is dual stack. Okay. But on this side, like on the management side, it's V4?
Yeah. Okay. So these things are with respect to, like, this is my loopback IP of the switch, my management IP of the switch.
So these are IPv4s. Yep. Links between switches, they are either unnumbered or they can have IPv4.
But we're working on IPv6 support for switch-to-switch links. Why? Because some future clusters that are coming, they need more IPs than IPv4 can support.
They will have more links than you can have private IPv4 addresses. Thank you. So we created that, and the simulation is loading.
We created entries in IPAM, so subnet for loopback, subnet for management IPs. And we have this topology. So topology kind of is just information about how it should be cabled.
And we have this concept we call inventory profiles. So- Can I ask you a question? Yes.
Will this build cable maps? So this way you can kind of hand the cable map off to your install folks? Yes.
Export cut sheets. Here, we have export cut sheets. We also have export labels.
Perfect. So we can print and and stick. So this top part here is the back-end fabric, east-west fabric.
This bottom part here is the north-south fabric. This middle part, these are GPU servers. In this architecture, we have eight GPU servers, so this NICs one to eight, each GPU connected to east-west fabric.
And NICs nine and 10, those are production NICs. They're typically bonded from the server side. And 11 is IPMI management interface of the server.
We have these inventory profiles, one profile connected to east-west switches, one profile associated with north-south switches. And this is where you define things like what is my DNS server or who can SSH to the switch. But also you define things like these switches should be optimized for QS and RoCE.
And each hardware vendor has their own how you configure RoCE kind of recommendations, and we know what is it for different vendors, and combination of operating system, and this checkbox tells our algorithm what exactly to configure. So from user perspective, it's just one checkbox. What's going to happen underneath, a lot of configuration will be different.
Now, we can see that this topology part is changing colors, and if we go to our health dashboard, see what's happening. Controller doesn't know that this is a simulation. Controller thinks that these switches are coming alive and they are getting operating system, and then controller pushes what we call Netris agent on the switch.
And then agent, using out-of-band management, kind of calls back home to controller. Controller is usually hosted on-prem in the data center. But the way communication works, switch will call back home to controller.
And when that happens, we receive heartbeats, we show heartbeats, and we can see that 22 nodes, which are basically 18 switches and four south gates, already established contact with the controller. And on this side, we can see that a number of critical alarms are showing up. What's happening here, this agent is waking up and there's this brand-new operating system with no configuration.
Agent is like, "Oh my gosh, no configuration. What are we going to do? Who I am?
"I have these neighbors. I'm supposed to have these neighbors. Okay, let me try and bring some BGP configurations up.
Let me see if I can do that. And then both automatic generation of configuration happens and automatic monitoring happens. So not only agent tries to bring up that BGP session, but it starts monitoring lots of things like- Quick question for you, Alex.
You're talking about configuration generation. Is that performed by the agent itself and then applied to the switch, or does the mothership, the Netris controller, create the configuration? It's done by the agent.
Okay. Controller- Follow-up question. Yes.
Oh, sorry. I'll let you finish the answer there. Controller has no idea about these configurations.
Controller can receive information of like, hey, this port is not coming up, but controller has no idea what was or what is supposed to be configured on that port. Okay. And then I have a follow-up question.
It's probably a silly question, but with your product in place, do your customers still back up their configurations? Do they still need to back them up somewhere- That's a- ... when they're generated on the fly?
That's a great question. No. There's no backups in our world.
No backups of switch configuration. You do need to back up your controller, and controller will create daily backup files, and you just need to copy it. But there's literally no backups of switch-by-switch configs.
Switch dies, you replace the switch, and the agent is again, "Oh, my gosh, empty switch. " Okay. I know of some enterprise networks that also don't back up their network configs, but that's a different story .
Yeah . Okay. So now we can see that now most alarms are gone.
Wiring is consistent with the topology. Surprisingly, all cables are connected correctly- Connect a cable award ... in the simulation.
In physical data center, our practice shows that roughly 5% of cables will be miswired on average, and it's a big number at scale. And then you might spend few weeks of fixing this. But having a list of what is miswired, it's really, really helpful.
So now that the fabric is mostly up and running, no alarms, let's connect to some of the servers. Okay. So the way simulation works, you can use your controller as a jump host, and from inside the controller, you can SSH to any switch or any server.
And in real life, at this point in time, servers would not have operating system, right? Because it is job of another system to install the operating system. But for the sake of training and demo, we bring up these machines with some Ubuntu operating system, so we can execute pings, and these machines have a ton of interfaces.
So we use this script called cluster ping. It takes host ID and executes parallel pings across each rail. So one command and 10 pings.
Useful. Now, Host 0 is pinging itself, and obviously all interfaces respond. But if Host 0 tries to ping Host 1, we're getting timeouts.
So on this other window, I will connect. So left and right windows are same jump host. On left side, I'm connected to one of my servers.
On right side, I'm connecting to one of the switches, one of the leaf switches. And here I'll do some show commands. So this is list of VRFs so far configured.
And this is a list of VLANs, VXLANs, and bridges so far connected. And I'll go back to the controller to create some tenants, to onboard some tenants and see how configurations change. So as a network engineer, we want to provide consumers a way to easily create these clusters.
So this is that construct. We call it server cluster. So in this construct, you put some very simple parameters, just a name, admin, like a local admin for Netris, which data center, whether to use existing VPC or create new VPC, and then a template.
We select this template, and I'll show this template soon. And then here from consumer perspective, we're selecting just few servers, right? Four servers in this case.
Add, and this is being provisioned. And let me create one more. Maybe we call this Bravo, and all the same parameters.
Again, create new VPC, same template, but the only difference that servers from zero to three are occupied by Alpha, and I need to select other servers. So the intent here is to simple use case. Two bare metal subclusters is what a consumer is trying to create, right?
Simple descriptionsBut it's network engineer's job to turn this into VRFs, VXLANs, and all that stuff with the help of NetRis. So network engineer defined these templates. Remember, I used the template.
As a consumer, I use this template. And this one is example template, and you can, different customers do different templates. But the idea of this template is network engineer is telling the system that, "Hey, for this kind of subclusters, I want you to create one VNet, one segment for backend networking for East-West, and that one I want you to create L3VPN kind of template.
" Now, Ethernet 1 to 8, that's with respect to the model that we have in the controller. That has nothing to do with actual interface names on the server. If you go and if config, there will be different names.
And then the template also tells that, "Hey, I want another VNet in my North-South, and that one I want to be an L2VPN type, and I want to have this as a default gateway on that one. " Now, in this case, two customers have overlapping IPs, and this is kind of choice of this engineer. And then we can see that in VPC section, we can see that two VPCs inside NetRis have just been created.
See, the time been, like two minutes ago. And if we go back to this switch, remember we had this show command that showed that existing VRFs. And if I rerun that command again, we will see that two new VRFs have been just added to the switch.
These two. And these numbers are aligned with VPC numbers, so it's just for convenience. And then if we go and check this VPC section, the VNet section, we will see that three VNets were created for Alpha and three VNets were created for Bravo.
So basically, that template is what connects network engineer's world with cloud consumer's world. And with the help of that cloud consumer is dealing with the simplicity and network engineer is dealing with network engineering constructs, VXLANs, VLANS, whatever network engineer needs in this case. And if, again, we go back to our switch, we initially had this list of VLANS, VXLANs, and bridges configured.
If I rerun that same command, we can see that some additional VLANS, VXLANs, bond interfaces were created, so this versus this. Right? This compared to this.
And we can even see that on some interfaces, even bonding has been configured. Why? Because we detected that server has LACP configured, and it's like, oh my gosh, the server is trying to do EVPN multi-homing.
Let me configure from the switch side. And if we go back to the server, which was failing to ping its neighbors, so Host 0 couldn't ping Host 1. If I rerun that, we can see that Host 1 is responding on each rail, production interface, management interface.
Just like that, Host 2 is responding, Host 3 is responding, but Host 4 belongs in different tenant, so I should not be able to access it. So now that these four machines can talk to each other at the backend, front-end, lot of capacity, lot of throughput, cool. How do we access this thing from the outside, from the Internet?
So I need to copy this private IP, and I'll go back to my controller, and I'll configure an elastic IP. Okay. I select DNAT.
We support source NAT and different things. But in this case, we need DNAT, right? Destination natting.
Here we select the right VPC. We select the source. We can do on protocol basis, TCP port, but just a simple case.
And from here, we pick the public IP, maybe this one. And it's a pool that network engineer designated for this function in IPAM. You mark the pool for this function.
And we click Add, and I need to copy this IP address, and so hopefully this works. So here, see, this one, this window is not using any jump host. This is basically Wi-Fi in this room.
So hopefully there is routing. Oh, looks like there is. And now there is.
See? So configuration happened. We have this ability of connecting simulation to the world.
And just to prove the point, and who knows? Let's create this very simple application on my server, and let me try to telnet to that TCP port through Wi-Fi. Yeah.
So data exchange happens and it actually works. Okay. Now we have four more minutes, and I want to use this time to quickly show some of the DPU stuff that we have recently built.
So, where is it? Yes. So for DPU demo, I'm using a physical lab, not simulation.
We cannot simulate DPUs yet. Try that. So in this setup, we have these three servers.
Server one has a DPU, server two has a DPU, server three doesn't have. It's regular host. And- Are these Bluefields?
Bluefields. Yeah. Bluefield, yeah.
And here, this machine, I'm connected to this machine. This is server three. And on this machine, I have this...
Again, this is continuous ping, which is pinging its own gateway, and it's pinging two private IPs that are on some VMs behind that Bluefield. And it's currently failing because I haven't configured much. So I'm going back.
Let me quickly show, or actually, let me create a VNet for that. Okay. We need to select the right VPC.
This will be VPC for demo. You can tell it's a physical lab, right? A lot of things configured, our engineers are using this lab every day.
We select Admin, and then we add an IP. So this is the IP, this is the default gateway IP address for this subnet. And then here, I need to add server three, its eth1, as untagged.
Then I want to add server one. And for server one, see, because it's DPU-enabled, I'm seeing this pf0, vf, these interfaces. So I need to include, I believe I need to include pf0, vf7 for server one.
And for server two, I believe I need to include pf1, vf7. Hopefully, this is right. Yes.
Click Add, and it's provisioning. It will take some time, and it's still failing because it's provisioning. Takes one minute.
While we're waiting, let me show a couple things. There's one minute. It should work in one minute.
In this case, with a DPU, when I edit the server object inside Netris, see? It's slightly different. So compared to regular server that doesn't have a DPU, see?
There's DPUs, says zero. And this other server, in this case, is saying DPU is one. DPU is like a switch, like another network engineering thing.
So we need to manage this thing. So we need to have a loopback IP, management IP, and AS number. There's a way to generate these things automatically in a fast way, yes.
And let me quickly show you the diagram of what it is. So, this is server one. It has a DPU.
Server two has a DPU, and there's server three, no DPU. And I have a VM inside server one that has this IP that ends with 101, and it's connected to vf7, pf0, vf7, because there's 16 of these PFs, VFs. And on server two, VM with IP 102, and on physical machine, IP 103.
We're out of time. And when we go back, we can see that provisioning is done and- Fricking amazing. Bare metal can ping its gateway, and bare metal can ping hosts behind the DPU.
And you can even tell, see the numbers are, there's round trip time is slightly lower with the default gateway than behind, because it's two hosts away, two hops away. That's it. Three seconds over.
30 seconds over. I'm just impressed you did it live. Yeah, all of us, yeah.
Yeah, that was pretty impressive. I hope you all got all of that because the quiz after this is that you have to do that. You were taking notes, right?
Were you taking notes at home? Oh, too bad. But I think if you were very interested in trying out Netris, I'm sure that you'd be very well served to give them a call, send them an email, mention them on social media.
I think somebody might be able to help you out. We're going to go ahead and call it a day here. We've got a little bit of Q&A with the Netris team after this, but then I think we're going to go beat something up.
Not a person, it's a room. It's a room full of Catalyst 6500 switches. Mm-hmm.
Probably not. But we're going to do that. m.
with day three of Networking Field Day 40. We'll have a couple more presentations headed your way, including one of our delegates presenting. So make sure that you are tuned in and ready for that.
But for now, for Thursday, thank you very much for tuning in. Currently, I have to go unplug a whole bunch of DPUs over at the Netris office. I'll see you guys later.