How it works. Multi-Tenancy & Network Automation for AI Infrastructure Operators with Netris
Netris, as presented by CEO Alex Soroyan, offers cloud-provider-grade network automation and multi-tenancy software tailored for AI Infrastructure operators. The core of their solution lies in the Netris Controller, which acts as the centralized source of truth for network engineers. It allows for the modeling and simulating network infrastructure using tools like Terraform and CloudSim, while also providing APIs that can integrate into cloud provider platforms, facilitating the creation of VPCs and managing network functions. A key component of their offering is SoftGate, a gateway on Linux servers that provides functions such as elastic load balancing and NAT. It offers a streamlined, integrated solution compared to separate, third-party products.
The presentation details Netris’ approach to day-zero and day-one operations, highlighting the use of Terraform for infrastructure-as-code methodologies and how the controller facilitates the deployment and management of various switch vendors. The system supports granular multi-tenancy through VXLANs and is designed to integrate with shared storage solutions. Netris facilitates access and isolation by allowing access to the network from the storage and the tenants, intending to integrate directly with storage vendors via their API. This setup allows for a cloud-like experience for AI infrastructure operators, streamlining the onboarding of tenants and the allocation of resources.
Netris differentiates itself by being multi-vendor and providing cloud networking constructs not typically found in traditional network automation platforms. The presentation emphasized the efficiency and integration provided by SoftGate, which eliminates the complexity of connecting firewalls and load balancers while supporting InfiniBand through integration with NVIDIA UFM. Alex expressed confidence in Netris’ position, particularly given the growing demand for cloud-provider-like capabilities in the AI infrastructure space.
Presented by Alex Soroyan,CEO and co-founder, Netris. Recorded live in Santa Clara, California, on April 24, 2025, as part of AI Infrastructure Field Day. Watch the entire presentation at https://techfieldday.com/appearance/netris-presents-at-ai-infrastructure-field-day-2/, https://techfieldday.com/event/aiifd2/, or https://www.netris.io/demo for more information.
Transcript
I'm Alex Soan, CEO, and co-founder of Net. Uh, we're a software company focusing on network automation, obstruction, and multi-tenancy for AI and cloud operators. Uh, and this section is focused on how, how we do this, how we automate and, uh, deliver multi-tenancy for, uh, cloud providers for AI and cloud providers.
So, just, just to remind you, just just to set context, uh, we, we, we, we had this notion of three pillars, which is basically an ability for, uh, infrastructure operator to have the notion of VPC, like in the cloud, to have the notion of cloud networking functions. Basically, functions that switches do not do, but are required in the network. And fabric manager functions, functions for managing, operating the fabric for day-to-day operations.
Now, how it works is, uh, this, it, it starts with a, with network controller. Controller is, uh, a centralized place where, um, you know, this is centralized place, first of all, for a network engineers, you know, on day zero, on day one, network engineers need to set this thing up, right? So, uh, and, uh, back in the day, network engineers were using a lot of, you know, Visio and like Excel sheets to, to kind of, you know, work with some data and do some preparation, which is perfectly fine.
Here we provide some tools which, uh, which help network engineers to do, to do some of this preparation right inside the controller. So you're not only drawing a diagram, but your diagram is becoming kind of set of objects talking to each other, basically helping network engineers to, to, to define the source of truth for, for the network. Very helpful with day zero, day one, uh, operations or network engineers.
Then they plug the hardware. Uh, they, they, they, they bind physical hardware with the, you know, logical, uh, objects in the controller. We understand that, uh, we, we, we install the operating system, then we install a lightweight agent on the switches.
That way agent talks to controller. And there's number of algorithms in this agent that this algorithms are, you know, translating this high level intentions into actual validated configurations. And that way network works.
Now on day two, we also provide APIs to cloud providers, which are not network engineer, that are not necessarily for network engineers, but they are more for compute and storage for other guys. So, as a cloud provider, you are connecting all this compute, storage and, uh, networking together. You need APIs to, to make your different platforms talk to each other, and especially when you are building your user facing portal, you need, uh, you, you need APIs to call on, on the backend.
So, for example, when, when a tenant comes in, you wanna create tenant in the system, and we'll see these examples in in the coming slides in more details. Question, I have two question, uh, one about the zero one, do you have a kind of infrastructure code capabilities to, uh, deliver when the switch are connected? Yeah, that's a, that's a, that's a great question.
So, uh, we, we do, uh, our controller, and you will see this a little bit in the demo, but in short, controller has three interfaces. There are equal web, uh, rest, API and Terraform. Oh.
So if you want, if you as administrator of that AI cluster, if you want to treat your network infrastructure through infrastructure as a code methodologies, you can talk to network controller in Terraform. Okay. And another question is, for a platform engineer's perspective, uh, about monitoring telemetry, do you support any open telemetry or stuff like that?
Um, yeah, good, good question too. So, uh, we collect telemetry from the switches and send telemetry to our controller using our proprietary tech methods. Okay.
But that way, uh, see, you, you, you have all this data in our controller and some customers, they like to have like all this data in Grafana or somewhere centralized. This data is collected through proprietary method into controller, but in the controller, it's available to you. So you can take all this data from the controller, uh, and you can view in your Grafana or, or whatever you're, you're using Integration with Grafana and, um, gotcha.
It's not, It's not our integration, it's community integration, community integr. Yeah. Community is doing this and yeah, yeah, this is all possible.
Gotcha. Thank you. Do you have Ansible integration as well, or do you have it planned?
Uh, no, no, no. We, we, we don't have Ansible integration. Well, we have rest API, which can be consumed in that format.
And where are the, how do you get access to the APIs? Can you just go to your, um, landing page and get to it from there? Uh, so API documentation is built into the product.
Uh, during the demo, I will show Okay, a little bit. Yeah. But it's, it's, if, if it's a new customer and they don't, and they don't have access to API, we, we can, we, there, there are ways to, to help customers to read the, to, to familiarize with API like upfront when they're doing evaluation of the product.
Okay. No, nothing to hide from that perspective. Okay.
And, uh, component number two is, uh, what we call soft gate. Now, switch management and isolation, all that is very important, obviously, but there are some functionalities. The switches cannot do, like elastic load balancers, like net like, uh, uh, network address translation, uh, like elastic ips and some others.
Now, these functions are critically important. And, uh, one way to, to have to have these functions is to involve, uh, other products like dedicated firewalls or dedicated routers, which is fine. That's okay.
And that is compatible with net with, there is no conflict, but we offer this additional alternative option, uh, where we have this component called soft Gate. It's a gateway that's running on Linux servers. It is optimized, it is horizontally scalable, it is aware of tenants.
So you can have millions of tenants, millions of VPCs. And this thing is designed to, to scale like that. And the great thing is that when you use external router firewall, you, you are gonna have the challenge of, you know, connecting, making APIs of two different platforms, talk to each other, APIs of your switch fabric to talk to APIs of your firewall.
That's, that's, that's a tough, uh, tough task for a cloud provider in case of Southgate. All that is done. It is a hundred percent integrated in naris.
It works seamlessly. And, um, I will show this in the demo now. Uh, some of, uh, you know, just an example of, uh, what, what kind of like, what typical network looks like, and, uh, kind of to walk you through the process from going from like day zero to like having running cluster.
Uh, typical AI cluster has two networks, front end and backend network. Uh, our controller talks to this different switches through out-of-band management, which is typically not network managed. It's kind of separate network for safety that we leverage to make these agents running on every switch to talk to our controller.
Same management network is also used for, uh, our soft gate to, to talk to our controller. And, um, we also optionally manage the host networking. Not every user needs this some, but some need this.
And in that format, we, we also run an agent inside the host, and that agent communicates with the network, uh, actually using LLDP and some, some smart algorithms over LLDP to learn, uh, different things about such, such as like IP addresses, routes, and, uh, program, some ips and routes for especially that is especially useful for backend network in AI fabrics. Um, and that, that agent also configures, uh, super nick, uh, that also configures DPU and in the future releases. We, we will also support granular multi-tenancy through, uh, you know, terminating VX loans, RFS, and other constructs directly on, on the DPO inside the server.
Uh, now network controller is typically used by network engineers as well as other platforms. So if cloud provider has this end user facing user interface, oftentimes we are kind of backend for, for network functions of that. Or some customers are using, uh, different paths, platforms.
Uh, again, we have integrations with different paths, partner platforms where they again use, uh, our APIs to basically provision networking to create VPCs, uh, you know, networking functions, et cetera. Uh, if backend network is based on InfiniBand, in that case, uh, you know, we don't touch InfiniBand, which is directly InfiniBand world, has this amazing fabric manager called Nvidia, UFM. So instead, we, we, instead of like loading agent on the switch, we tie our agent to, uh, UFM fabric manager.
This way, you know, ethernet customers and InfiniBand customers, they get exactly the same experience, same APIs, absolutely zero difference. Now, uh, on day zero, day one, uh, what happens, uh, first step in, in the data center, in, in data center, first comes network. You, you need network to be there to be able to install compute and storage, et cetera.
And, uh, how do we bring that net network up? So network controller comes first because it is source of truth. And then typically either implementation team, uh, health kind of models, what network is going to be in network controller implementation team can be customer, customers team can be our professional services can be, uh, you know, system integrators, professional services or vendors, professional services.
We, we work with all of them. Now, we, we provide tools that help to, you know, through Terraform and using, uh, DevOps methodologies, help them to describe their infrastructure in net risk controller. So through Terraform you can say, Hey, I'm building lift spine fabric, or I'm building AI infrastructure based on Nvidia Spectrum X architecture.
Now we know what is Spectrum X architecture and NVIDIA's RA for that. And in that case, we would say, Hey, how many nodes doing need? How many, how many uplinks doing need?
We will ask very basic questions, and we would help you to populate the, the controller. Then you can make some simple changes. Once you have, you know, told controller, what's your building?
You can simulate that. So, you know, um, there's NVIDIA Air, which is simulation platform, which we are compatible. We can simulate, we can create a digital twin in Nvidia Air, but we also have our own platform.
We call it Cloud sim, which is actively used by our developers, our customers, evaluating customers, our partners who are, uh, you know, integrating with us. It is incredibly easy way, uh, to, to like experiment and build and, uh, evaluate infrastructures once it's done, evaluated hardware is there, switches are wired, and we, we, we detect switches, we identify them by MAC addresses, and we install operating system. Our agent run configuration, they're up.
And your network is kind of in kind of a, like a ready, ready state. You can start onboarding your tenants. Now, let's say you get two customers, two tenants, let's say Coke and Pepsi.
Um, you, so they knock your door, Hey, I need five servers. I need 25 servers. Whatever.
You, you, you describe that very same thing in net risk controller servers, 1, 2, 3, 5, go to 10, one servers, 20, 25, 23, whatever, go to 10, two, enter. Um, and because net is source of truth, it knows all your topology, it knows how your servers are wired connected, how is your network is dealing net knows, is this NVIDIA network? Is this Arista network?
Is this like, is this InfiniBand? What, what are the best practices? So this algorithms automatically go and generate appropriate configurations in every switch.
Um, you know, if it's a bare metal, uh, the isolation would happen on the switch ports, basically on the switch layer. If it's, uh, you know, after that, you know, provisioning is done, VRF VLANs V excellence are created, and your, your servers can talk to each other. Now, there's always this question of granular, uh, multi-tenancy, like how to do this per gpu, per container perm levels.
Uh, this is done through, you know, extending VX excellence to the host. So basically making, either making the, the entire server participant of the VXLAN fabric or making the DPO participant of the VX non fabric. So meaning the same protocols that switches talk to each other, they also talk to switches, treat your servers as if servers were, were switches to kind of like that, using EVPN and, and B-G-P-E-V-P, um, signaling methods for theon.
Um, then, um, uh, very, very common question. Uh, you have these two tenants, Coke and, uh, Pepsi, for example. Uh, and, uh, how do we provide shared storage?
Now, uh, different storage vendors, they have different methods to identify customers and to isolate. Uh, but most, you know, typically they support either of two methods. One is, one is to have different VLANs for, for different, uh, tenants.
And one is through what is called VRF, uh, leaking, or we call it VPC hearing. Now how AI infrastructure cloud provider can, can implement this. You know, it is, it is easy for a storage provider to say, oh, just, just configure v RF leaking.
And, uh, it'll work. Of course, that's that storage, uh, provider's requirement. But it is, it is challenging for, uh, AI infrastructure operator to configure, uh, that peering, uh, dynamically on the fly without making, uh, you know, when you are implementing configuration for one tenant, you cannot, you know, you cannot break operation of other tenants or using workloads is a, it's a, it's a production.
So it's, it's cloud experience, right? Your customers don't expect things to stop working just because you are onboarding another customer. So, uh, again, we, uh, because we, we have this notion of VPCs.
We, we have this notion of VPC peering, just like in typical cloud, we're from cloud provider's perspective, it's a simple object that says here, this VPC with that VPC, in this case, you wanna say peer, my tenant coke with my tenant shared storage and peer, my tenant Pepsi with the, with my VPC, uh, shared storage, uh, that way Pepsi and Coke cannot talk to each other. They can talk to storage, but they cannot talk to each other because storage would be, would have its own VPC, its own VRF. It is not transiting any traffic and the configuration on the network, and this is AI network, hundreds of switches configuration on the network is kind of like that.
You see that, you know, green lines, it's, it's basically hundreds of green lines like that on every switch. And it's not, you cannot copy and paste if every, every switch has its own configuration, but it's the algorithm that knows how to properly organize, uh, the VPC peering, VRF peering. And it's kind of from cloud provider's perspective, it's, it's not their problem anymore.
It is our algorithm's problem, which is obviously rigorously tested in, you know, virtualized and physical environments, and it's also being run by lots of lots of customers. Right. Good question.
Question. Thanks. Uh, Brian Martin, signal 65.
Um, when you're looking at VPC storage and storage VRF functions, you have a qualified set of storage vendors you work with, um, that you interact with, uh, so you can configure them, or how are you managing that interface? Uh, yeah, there's a, uh, so we we're intending to work with everyone. Uh, there are some storage vendors with whom we work, um, closely.
Um, when time comes, the announcements will, will follow, but, um, we, we work closely. And the idea there is, uh, for, for them to integrate with our API and them by seeing, you know, different parameters in our API, they can kind of automatically understand, uh, who, who to permit access. Now, let's say particular AI customer is using storage vendor that we don't have integration.
Uh, we, we are seeing all, all the storage vendors in different deals and, uh, we, we don't have kind of a p integrations with all of them. We, with some of them it's in progress, but even if there is no integration, uh, this, this solution, uh, is, is is already sufficient because what we do, we ba we, we provide the access from networking perspective, right? From storage vendor's perspective, it is, it is one network.
The network of storage is still, uh, living inside the same switch that we manage, right? So we, we basically manage access to that network. We say tenant K can access storage and storage vendor itself.
They know that tenant Koch has this ips, and if traffic come comes from tenant Koch, or if traffic comes from this ips, I should, uh, serve because I trust that traffic is coming from really coming from that tenant. Because NET is enforcing this in, in, in the VRF leaking LA layer. And in the case of the storage, they know that that IP address is acceptable because that was configured a priori eventually through your API.
Yeah. Okay. Yeah, they, they, they trust that, uh, VRF leaking is configured appropriately.
Uh, in, in the future, we will do some integrations with some storage providers where they will not even need to know the IP of the tenant. They, they will see it. But I assume you're just using agents that you've put together to deal with the storage APIs?
Uh, we, with, with storage in, in today's, uh, situation, uh, agent is not required. Uh, there's no need for integration with Do you Have a direct integration with the storage vendors you work with? Now Today, we, we just control network access and network is still managed by net see storage, uh, storage vendor is basically that, right?
Storage, storage servers, yeah. Are connected to north south fabric. And North south fabric is entirely managed by net.
Now, the way ai, uh, uh, infrastructure operator configures the system, we, we, they give us the notion that this switch ports are not connected to tenants, but are connected to storage. So, and then because we know that this is storage, this is not random server, we, we were able to, when they say, uh, you know, provide VPC peering means we actually yeah. In this slide, when, when, when provider says Koch tenant Koch should access storage, we understand, okay, we will go and configure all the appropriate rules in the switches.
Mm-hmm. So traffic, which is exiting server, uh, from Koch, trying to travel to storage, will not be blocked. Okay.
But, uh, do you have any provision for handling multi-tenant storage? Uh, storage should be provisioned. Uh, se separately storage should know that there are, there's Coke and Pepsi tenants storage should expect this traffic.
Yeah. But it, it would be potentially over the same network connection that you would access it. Uh, that's a, that's a, that's a good point.
So today, uh, AI infrastructure provider needs to tell Naris about tenancy and storage provider, you know, okay. Two separate calls possibly in the future, and we work with some storage providers on this. There will be integration where AI infrastructure operator will tell us about tenancy and storage.
Could, could, could pick up that information from us. Okay. Question.
It seems like, it seems like you're trying to build that abstraction that would make us, if I had to go build all of the declarative terraform myself to do this at scale, right? Mm-hmm. I'm talking multiple TF files and I have to then ingest all of the vendors.
Is that, are you still building a declarative approach out to all of the networking infrastructure? So you're, you have templates, you have code that can, that can be defined as it goes to each vendor. What's kind, what's kind of the layer between net controller and the switch?
Is that something that can be modified or looked at? Uh, that can be you, you mean by customer? Like, like how, how, how much can customer extend?
Yeah. By customer or by use case? Yeah, all of the above.
So we, we, today, we provide, uh, most essential obstructions, uh, in, in terms of vendor support. Uh, we support these four vendors, Nvidia arisa dell edge org. Um, for us, this is individual development for every platform, right?
Every platform has different commands, different best practices. Something is possible, something is not possible, and you need workarounds. So our algorithm needs to take care of all that, that minor details.
Now, we, we, we develop our roadmap based on, you know, basically demand and customer needs. It is very much focused on AI and cloud infrastructure use cases. Most our customers are either running AI or they are cloud provider.
So for these two use cases, we, we have really good coverage. Now, if someone wants to do like something else, um, we, we, we should talk, uh, you know, sometimes, sometimes we would say, Hey, customer, what you are trying to do, we have, you know, we have the building blocks, maybe you can leverage this, this, this APIs. Here's example code.
You can do a little bit of your own additional development, and you can extend to your use case. Okay? There was storage.
And, um, another very common, uh, question is, okay, we figured out storage, uh, sharing. And the other question is shared network services. So how, how do I provide internet access to, to my tenants?
How do I, if, if I have like a data scientist that's trying to access this cluster from the outside, how do I give them elastic ip? And, um, of course it's possible to do with like dedicated firewalls, routers, but the level of integration you need to, to do you as a ai, uh, cloud provider, you, you need to invest so much efforts into this automation between firewalls interface and your, you know, fabric managers. API, uh, we have this alternative option we call Soft Gate.
You know, through Soft Gate, we are able to provide this kind of same, same simple abstraction where you can say, Hey, net, let's create a net rule. Or Hey, net, let's create an elastic load balancer. And, um, how it works is, uh, following, see in, in a switch fabric, in a modern switch fabric switches hooked to each other through E-V-P-N-B-G-P, and that way they, they run the excellent tunnels.
And that makes, uh, the switch fabrics very flexible and scalable. Uh, now the soft gate nos, uh, are basically Linux machines running trics, uh, with some acceleration, with some, you know, smart algorithms that we've created to do some things right, and scalable. Now, soft gate nos, they talk to the fabric using same BGP and DVPN and VXLAN methods.
So to switch fabric, salt gate is like another switch. But in reality, salt Gate is not a switch. It's a server.
It can do different things, things that switches cannot do. And that way, soft Gate, if, if Net Wants Soft Gate can access any VX one, any VPC, any tenant basically. So when customer says, let's provide, uh, you know, load balancing service for Soft Gate, first of all, uh, first of all, how, how Soft Gate gets connected to the internet itself, uh, you know, typically you have your border switches and they have, you know, links connected to upstream router or upstream internet provider.
Now Soft Gate forms A BGP sessions with this upstream providers, which is kind of very easy for network engineers to, to figure out how to organize this. That way Soft Gate has access to the internet, soft Gate can terminate some public ips, and when a cloud provider or their customer is requesting for Elastic IP or load balancer, soft Gate can pick an IP from its pool, and, um, that IP is visible on the internet and the traffic coming on that IP can be serviced by Soft Gate and load balanced for further, uh, towards one of the tenants or, or, uh, you know, network address translation. So, so that's the principle how Soft Gate works.
And I will show this a little bit on, in, in the demo. Uh, so oftentimes, uh, you know, when we, uh, when we show the product to engineers, they say, Hey, uh, common questions like, how do we, we're evaluating this for our use case. Um, what, what are the important things?
What are the important angles we should keep in mind when, uh, we're doing this evaluation? How, basically how you compare to typical hardware vendor products coming from typical hardware vendors. Uh, there was a question, are restock cloud vision, right?
Very common question. Uh, or, or like other offerings, uh, right, other independent, uh, network software vendors. Uh, so, so big one, some of the important differentiators, uh, are we, we're multi-vendor.
We, we don't manufacture hardware. We, we don't care what hardware use. We want you to be successful.
Um, we provide this cloud networking construct, uh, which is, you know, typically, uh, traditional network automation platforms. They were designed more for, uh, enterprise data center use case, and they don't necessarily optimize for this kind of AI cloud builder kind of use case. Um, we, we have this soft gate, which is, uh, which is kind of u unique in terms of we haven't seen, uh, you know, similar, uh, product elsewhere.
Uh, typical approach is to have a firewall to have a separate load balancer, which is okay, we support that use case, uh, absolutely no problem. But for a lot of customers, and a lot of customers are saying, can I use, can, can we use Net without Soft Gate? We're saying, sure, you'll be surprised to know that we probably have only like one or two customers that are not using Soft Gate.
They all use it. They all love it because they're like, this is cool. Like, well, it works fast and it's all integrated.
Uh, we support both, uh, ethernet and InfiniBand. Yes, it's a simple plugin through Nvidia UFM integration. It's not much, and we charge very small amount for, for that, uh, small integration, but it, it makes, uh, customers lives so much easier.
And we are verified, uh, Nvidia Spectrum x uh, solution. Question, uh, Brian Martin, again, um, other ISVs, when you look at the market and what you're offering, who would you say are your top handful of competitors in this space? Uh, I think, uh, that's an, that's an interesting question.
Uh, not, uh, not, uh, typical hardware vendors. Uh, in, in tier, in, in a regular CPO cloud. I would say typical hardware vendors.
Like, like Arisa Cloud vision, like Cisco, a CI in ai. Uh, not them, uh, but the main competitor in my opinion is, you know, doing it yourself, like building kind of artisanal automation in house. We see them as a main competitor.
Okay. So you feel like you're the leader in this space, um, bringing a solution to companies who've been trying to do with themselves? Uh, we we're, uh, we, we, we, we were lucky to be building something, uh, starting from 2018.
Um, okay, 2018 people were asking me, Alex, what is it that that net is building? I was saying software for cloud providers. And people were like, cloud providers, like Amazon.
I was, no, no, no, others and people who are these others, right? In 2018, like three years ago, there were some, and people were saying, how is it going? I was saying, it's going, it's not bad.
We, we are getting some customers, we're still alive. We, we can survive this way. But today is a different situation.
Today every AI infrastructure operator is basically cloud provider. So I think we've got lucky, uh, to, to be in, in the right time, in the right place. I know a lot of, a lot of, um, other companies are, are building solutions in the space.
Uh, there are competitive solutions, of course. Uh, some of them are decent. I don't know, you know, everything about all of them, but obviously the mar market will catch up, right?
Uh, it's, it's a huge market, but it's, it's a big enough it, there's, you know, room for everyone. Are we a leader? We're, I think we have the opportunity, uh, to, to be, to become one.
We're hoping to be one.