Design, deploy, and monitor networks for AI with Aviz
Thomas Scheibe, Chief Product Officer, offers solutions for designing, deploying, and monitoring networks for AI workloads. Their focus is on addressing the specialized networking needs of AI, including multiple networks, differentiated Quality of Service (QoS), and the integration of compute into the end-to-end network topology. They aim to provide automation and orchestration for faster deployment, service activation, and infrastructure expansion. Their product, ONCE, supports Sonic and Cumulus network operating systems, focusing on streamlining network management through design, modeling, deployment, and monitoring capabilities.
The Aviz presentation highlighted the evolution of networking in AI, emphasizing the shift from a single data center network to multiple networks, particularly the separation between front-end (user access) and back-end (GPU communication) networks. Aviz recognizes the importance of lossless behavior, different methods to address AI application requirements, and the integration of network settings on both the switches and the network interface cards (NICs). The company partners with hardware providers and uses reference architectures like NVIDIA Spectrum-X to automate network configuration. This allows enterprises to define networks and configure network separation.
Aviz offers comprehensive support for Sonic deployments in enterprise data centers and at the edge. They are automating deployment workflows for the NVIDIA Spectrum-X reference architecture, with the ability to configure multi-tenancy and extend the fabric. Aviz simplifies network management in AI, allowing users to deploy and manage their networks quickly and efficiently. They offer a comprehensive suite of solutions to design, deploy, and monitor networks for AI, focusing on automation and orchestration.
Presented by Thomas Scheibe, Chief Product Officer, Aviz Networks. Recorded live in Santa Clara, California, on April 25, 2025, as part of AI Infrastructure Field Day. Watch the entire presentation at https://techfieldday.com/appearance/aviz-networks-presents-at-ai-infrastructure-field-day-2/ or https://techfieldday.com/event/aiifd2/ for more information.
Transcript
Quick introduction. I'm Thomas Shier, uh, with vis, uh, we have Ravi here and then we have Krum here. He's gonna observe us, but presenting is gonna be Ravi and I, so the two of us will get you through one hour.
Um, what I wanna walk through, and I know this is all about Amazon questions, so please ask questions. I do have a few slides, you know, set up. Uh, but, but please chime in.
Um, I will do a quick overview, uh, for the folks that don't know what AR is doing, and yes, we're not the rental company ar these networks. I will tell you, I will tell you what we're called ar these networks. Uh, and then you see what we're going through here.
Um, let me, let me do the quick one of these networks. Uh, we we're a pure software networking company. Um, we started off as a, as a company that provides support for our enterprise customer.
One deploy Sonic, the open source networking operating system that got born in the cloud or was born in the cloud. That's where we started. But what we're gonna talk about today is actually the stack, the software stack we have and how do we enable customers to actually not just deploy Sonic, but actually use it and what other operating system we're gonna support in this space.
Uh, the little question behind why, why are we, um, couple answers to this. We wanted a name that sat with an A for obvious reasons, as you probably all realize. Number two, if you look at the three circles, it's like the idea behind this is like broken rings.
You actually can fly free and Avis actually means free Bird and Portuguese. And you're gonna see this in our little mascot at the end, but that's the end. Mm-hmm.
But that's the story behind the race. Um, and I, you know, and I say this because all my first question when I joined the company is like, what is up with this name? Um, so yeah, what we're all about, we're, we're a pure networking software company, which I think we're probably one of the few pure players, most networking companies, and I know this, I work for one of the big ones for 20 years.
It's a combination of hardware plus network operating system plus tooling stack. Uh, the reason why we do think this is needed is 'cause we do see, and it comes out of where the cloud providers came from, providing flexibility choice. And that then will result typically in a better RIE.
Uh, we're not a hardware company, but we are partnering very closely with all the hardware providers. 'cause obviously to deploy a network operating system, you need a hardware, you need switch. Um, the portfolio, what we have is three things, probably a little bit more, but if I have to just put it together, is one is what we call once, uh, and it stands for Open Networking Enterprise Suite.
Uh, it's really think about an automation suite. And now we'll go into detail that supports Sonic and we're adding Cumulus, actually we have added Cumulus. And the reason since we're at the AI field there, you might know why, because we do think this is where the pack's gonna go, uh, in terms of what customers will deploy.
And some of you probably look at me saying, well, I'm not quite there yet, but that's what we think where the future's gonna be. Uh, the second area, what we do have, we're not gonna cover this today, but I put it on the slide so you guys see it. Uh, we have a whole product around network, observability, software product.
Again, all what we do is software stack that runs either on switches, on service, uh, and then the last one, uh, is what we call network co-pilot. It's actually trademark that we have. And I was also very surprised to hear that none of the other large networking a winners actually sold about trademark there.
So this is our trademark networking co-pilot. Uh, and as you might imagine, this is really around how do I help a network operator get to ays faster? And we're not here to build over tools.
It's really, as you really think about it, it's a, it's an intellectual language interface. I don't wanna say a chat interface to all your tools that you're using today. That's actually in my mind, really exciting given what I tried to build in the past.
I think this is actually changing and will help how you can operate your infrastructure without changing. And then what time, obviously you're gonna go down, uh, and looking building infrastructure. So you saw a little tagline there, AI for networks and networks for ai.
Uh, what we're focusing on today is really the, uh, networking for ai. And this is really around what's changing in terms of networks are being built. And I'm probably preaching the choir since this is the AI infrastructure for you take, uh, I put it in two trends that really happened.
I said Sonic was born in the cloud. It was built, uh, in driven by Microsoft started somewhere in 2016. Uh, I vividly remember, uh, having my first sonic distribution on a separate company hardware.
Um, and so what we are really seeing this is moving to being deployed outside of the large cloud providers. Uh, this is going on for the last five to six years where this is seriously being deployed. So that's one threat.
It moves into large enterprise data centers. Uh, and it will move into enterprise edge, particularly where we're seeing this in retail environment. And there are good reasons for that.
Uh, and the other thing that will really drive this is AI inference needs moving obviously from training and large clouds to enterprises wanna have their own deployments. And these are particular around inference. So that's one of the trends.
The other one is really, and this is networking. Again, I don't think I'm telling you anything new. What used to be one network in the data center is now two or three.
Um, and I'm not here to debate whether what technology it is in my world, it's all ethernet. And I set that very consistently over the last five years. But it's more than one network.
Uh, that's the interesting piece. And you see it, it's really the separation that we see between what some people call it, the front and the back end. Some vendors call it the east, west, or north south.
Um, in my previous word, I I call it the scale up, scale out and the front end. Uh, but it's basically different networks. One is connecting GPUs outside of a server.
The other one is what you normally use to connect users into your, your server cluster. And then storage may become for some customer dedicated network just because storage in the AI environment is a little bit more following what GPU needs from a network than what a typically user network needs. And so, excuse me, can we Yeah, please ask questions.
Very Interested in, um, Sonic, and you don't have to go to Deep, but I'm looking at a blog post I did in 2016. I was at Dell at the time, uh, and we had our own OS 10 built on it. And I used to know the details, but I forgot.
Um, so how has that evolved and is that part of Open Daylight or is that, where is that, is that in the Lynx Foundation? Lenux Foundation? Yeah.
It And where Does it sit under, Where it sit under in the Linux Federation? I actually do not know the answer. I should and I don't, but I will get back to you afterwards.
Are You part, are you working with the Linux? Yes. We're Okay.
We, we representing, uh, as a company, we in there. Okay. Uh, the, actually the, the, yeah, lemme come back where it sits.
Okay. Um, but yeah, it moved there. It used to be driven on, and then Microsoft moved it into the Linux Foundation to make it a community project.
Uh, and you see, if you look at the chair, uh, group, there's Microsoft is on there. Um, most of the vendors, including Dallas, still on there, uh, Cisco's on there. There's, you know, all the large windows on there.
Um, And so a little bit on Sonic. So how many of you actually have touched look at Sonic at all? One, two.
Okay. So a little bit behind this Sonic really was the idea of, and this is becoming important 'cause we're gonna have a demo univers. So the idea behind Sonic was, is I should be able to standardize on the operating system independent of who is the hardware vendor for the device and independent on what selection vendor sits inside.
So if you look at what Sonic basically did, it says they defined an abstraction layer for silicon called sci, uh, SAI, right? That is owned by the silicon winner. That, uh, provides a common interface to the operating system.
And then Sonic is an open source community project. Uh, very similar vitality analogy. Linux, right?
And then what you see is you have different vendors building around that, their distribution. Uh, and then obviously you have a community version that you also can build a build, certify, build an image, uh, if it really comes down, what you have to deploy, and you will see this, you basically take Sonic, which is the control plan, the operating system. You do need a piece of software that covers the rest of the switch or otherwise would switch a router, right?
Fence, power supplies, LED, Blinky, blinky. You need to mess that happen. Uh, and then obviously you need the, the ASIC obstruction interface, the SDK, right?
And if you merge this together, that becomes here switch image like you would buy today from any hardware when you buy an operating system. So that's starting in a nutshell. And how has it over the last, like I said, the last time I really Yeah.
Followed it was about You were very early. That's literally when it started. So when did, um, and we had our own distribution and uh, Debbie and Base.
Yeah. How, so what has happened in the last 10 years? How has it changed?
Or is it just been minor money? No, significantly changed. As I said, it started off really, and if you think about it, right, what you basically have is you build an operating system and you're standing around what features should I add, right?
The initial feature set was all driven around what is relevant in a cloud provider data center, right? Think about Azure, think about your standard top of rec, or now we call it a leaf spine topologies. What features do I need?
Right? It went from there to, hey, this is great. Now we've proven, now runs in very, very large scale, right?
If you imagine, look at the large providers that have it, most of the cloud providers have a version of Sonic or as derivative of this. We're talking a hundred thousands of devices. It's probably the largest in install base of any os.
Uh, which is one of the nice things about it because you actually get more SOAP time than you would get in any other os. Uh, but what you saw back to your question is the feature sets that get added, and this is a community effort. There's a public roadmap.
Mm-hmm. We go from data center cloud to more features that are relevant. Enterprise data center, I think about VX and EVPN capabilities, and then in the last two years we saw interest around edge features.
Think about retail branches, campus, POE, uh, one or two to A, B, C, D, uh, one or three acts. You, you have these authentication capability, they are coming in, right? And so it's really driven primarily around what features get added by where users wanna deploy it.
Um, expectation. I have community and I'm sure some people have their own proprietary Correct. Through APIs that they Yes, correct.
So yeah, very, very similar. The analogy to what is going on in the next world on compute, we're like 20 years behind, but the same thing is gonna happen. Yeah.
Okay. Thank you. Good point.
So I'm pretty sure this one, most of you have seen one way or the other. Uh, what I try to depict here, if you, if you look at what used to be a front end network, which is on the bottom, right? You had your server and then you have your to or leaves, and then you have spines, right?
And then maybe a large, large deployments, you might have three tiers. What we really see here, which is very interesting, is now we have a backend, which is just connecting the GPUs between the AI service, right? There's another network inside the server or the rec, as you all know.
Uh, but really this is what happened. And so what's really changing here is a couple of things, right? Lossless behavior, uh, there are different ways to achieve it, right?
The, what the industry had for a long time was around, uh, uh, RDMA and and PFC that's around forever. But clearly there's a lot, a lot of innovation going on. How to achieve lossless behavior or fast, fast remitted, right?
There are different ways how to actually address what the application cares about, which in this case is GPU complication, right? One thing is to make a loss in network. One thing is do it on the nick, uh, or the combination of both, right?
There's congestion control algorithms that get improved. There's selective retransmission multipass that you implement, which is not standard ethernet, uh, as it used to be. Uh, and then obviously in the, in, in the past, mostly when we're about network, we didn't really care about what's going on in Nick besides dual homing for redundancy pass.
But now if you look at what, what really needs to happen when you talk about configuration of network infrastructure, you have to do the settings post on the nick on, on the switches to actually make them work together. Otherwise you would not get the desired behavior. Um, the other interesting one for all of are networking forever, and I'm, uh, aging myself a little bit over 25 years.
Uh, GPU networks, they just 10 to 20 x network. Uh, and so that's another interesting piece, right? Because even if you wanna build the same network on e, that you will need a dedicated network just for the fact that you need so much more bandwidth on the backside, right?
Uh, a server used to have two nicks for redundancy. It typically ISO has at least eight additional GPU nicks, right? You have each GPU has its own nick and that has four x bent with what you typically have on the front end.
So that on alone just adds, uh, change how we design. Uh, the other thing which is interesting from a design perspective, we go through this in the demo, is GPU backend networks tend to be very, very well defined. They're not like enterprise networks where, not here, not here.
Maybe I put my management switch on this, or maybe I connected directly into the this point. None of this is happening on a, on a backend network. Highly constrained topology, which actually makes it very, very nicely to, to this.
So what is the same, and I say this right, because not everything is new. Multi-tenancy is a requirement that external pass, and that just translates going forward on what you have to do on a backend as well. Uh, leaf spines we talk about forever.
Same thing here. And mostly well-defined. The backends, the backend.
When I say backend, I I really mean the GPU network. Very well-defined networks. The front ends tend to be a little bit more of a mix my experience, but you probably have your own and your own networks.
So on the prior chart yeah. Um, you listed, um, on the, the right hand side storage indicating that there is something new and changing there. Yeah.
But nothing, yeah. User GPU storage. I thought that was a weird kind of combination.
User GPU and storage. So what's the, what's going on with the storage side Here? Yeah.
If you look on the bottom, it's a good call out. It's, it's either traffic and storage. Yeah.
Um, and so there's a little bit of a shift going on from what we see. Um, it used to be, when used to be in this kind of funny saying this, because we were really talking about the last two years when we built these in network words, it was back in G-P-U-G-P-U connectivity outside of a server. And then in the front, when I say fronted that same server, how do you get data actually into the AI server to load what you actually run your inference on?
And then obviously user access to just get, you know, manage the, the, the server as well as probe. Um, what we see is there is these, these inference cluster get big enough that you actually want to put storage in a, uh, in a pool instead of having it directly attached to service. And you want to have access over the network.
And now you get to some of the similar requirements you have on the backends around latency. Mm-hmm. Uh, you see deployments of, uh, NVME over fabric mm-hmm.
Happening. And so you can run it over the same nick, when you think about the server and you think, Hey, I have my storage traffic going over the same nick, but you might end up having a dedicated storage nick. Uh, just because now I can control better the quality of transmission very similar to what I do on the, on the, on the GPU nick and then just have what they call a user normal standard access.
Right? And so we see both of these modes. Um, if you look at some of the reference designs and recommendation from uh, GPU vendors out there, they have both options.
Uh, my expectation is we will see these AI servers having literally three types of nicks. Uh, one for standard the traffic, one for storage traffic, and then one for GPU. I think the GPU next and the star next are most likely the same.
Uh, they're just dedicated where they, where they hook in terms of what, what's That way? Yeah, I was gonna say, you don't necessarily want user traffic on the same traffic on the same network as your storage traffic. You, you can, but preferably not.
Well, given the amount of money you spent on that server, you probably would just have another storage Network. Well, also, depending on what type of storage you have it, yeah. It might not necessarily be secure.
This is another one, right? Where do you do security, right? Do you do encryption?
Which some of the, the newer nicks that have EPU capabilities, you can actually do encryption on fly. But if you don't, you want more like a physical separation, that's another good reason why You want Yeah. I mean if, if you wanna fry a few eggs, you can do IPSec on your storage wrap.
No, no, I'm not. I'm not. Yes.
So, so Thomas is, is the, is the architecture still being figured out then? No. Or is it, is it very much reference design and we're just following Those different options and we get to this, we get to this.
Okay. Gotcha. Yeah, what you see on and my what, what I have seen and I think you see is in general, given how sensitive the requirements around latency congestion, you want to separate as much as you want GPU traffic, start traffic from everything else.
Mm-hmm. Uh, and they're very well-defined recommendations what to turn on, on the nick as well as on the switches. And you see this across awareness.
I see you see this across the traditional networks, awareness, Juniper, Cisco, Arista, you see it from Nvidia. And we will go through this. Right?
So before we dive too deep into the architecture, can you tell us more about who is, is a venture backed, is it a huge company? How did it get formed? Did people leave a, a small company or a Juniper and, and create this or Yeah, so the quick pitch, we're a startup company.
Uh, we are around since 2021. Uh, we have around eight funding as of last year. We're roughly 80 to 85 people.
Pure software. Most of these engineers, uh, our DNA is an sonic network operating system. Uh, if you look at the founders and a bunch of the engineers, we build operating systems.
We know how to support operating systems. And so the business process is, you saw the three products I had initially. Customers that want to apply Sonic.
If you're a Microsoft, you do your own support. If you're an enterprise, you're not gonna have a support team. You look to someone like it used today, if you buy a switch from Arista, Cisco dealer support the us.
If you buy Sonic, who do you go to? That's us. So you saw a gap as far as nobody was out there doing it for enterprises.
And you said, right, Because the traditional model is you buy the OS was the hardware. And when you want Sonic, you really want to buy Sonic, definitely from the hardware. And so there was absolutely a gap, which everybody wants to mimic what the cloud providers do, but they do need a support organization that does it and they don't.
And enterprise, if some of you that work in enterprises, you don't wanna hire 20 people just to support the os. So that's the gap and we stepped in. Okay.
Cool. And then you said you're venture backed. Do you have Angel?
Are they people who are in the networking area or is it like a Sequoia or, yeah, you look, if you look at our investors, we have a couple VCs, but we also have investors, which are hardware companies. So if you look at Cisco's an investor and us, uh, Celeste is investor, Edgeco Acton as an investor, uh, Qualcomm venture as an investor. So we have a mix of VC funding as well as what I would call strategic investors that want to take advantage of what we do with Sonic and apply it to their hardware.
I mean, that would be great to put up front 'cause it establishes credibility. It's not just, um, getting money from doctors and dentists. You're, you're actually getting nothing Wrong with doctors and dentists, but Yeah, exactly.
Because Then you can do what you want. That way they Leave you. No, listen, we, we are, we are well established.
So we're deployed in maybe, you know, I was more going somewhere, but we, we have deployments in fortune hundred enterprise companies. We have customer deploying as in 5 cents switches and more. We see phasing in, as I said, that's going on.
This the sonic, any enterprise is not new. Mm-hmm. Uh, this is going on for the last five years is really what, and it's just accelerating because of what's going on with ai.
It's just saying, Hey, I need to re rethink how build networks. So now I'd actually look at this, right? Even more.
Uh, but yes, we're, we're not, we have large deployments. We have a daily customer support organization doing nothing but supporting customers. And you have a, a csuite that knows all about this.
This looks like you came from Juniper and Cisco and all these different places. Yeah. I spent, I dunno, some of you might know, I spent 20 years at, at Cisco managing data center networking there.
I used to manage Nexus product line. So, and then, yeah, I spent one year into managing dpu. So Cool.
That's why this is, it comes together as I I I hope you Yep. Got a little bit. Yeah.
So now let me actually go in because I know, I wanna make sure we want to get the demos. Yeah. Um, so what we have is once, which is basically our automation software tool set.
And so when we started this, we built this to support customers around monitoring. 'cause we do actually, when an enterprise says rolls out Sonic, we give an SLA 24 7. If you have a step one within 30 minutes, we are with you.
And we'll figure out what's the problem really what a standard enterprise expects. So we built this to what we call once for monitoring. We get requests, uh, and you will see it.
Can you extend this and also add design capabilities? Like do you have a blueprint and now you can automatically design what is the configuration for a fabric? Can you model it?
Uh, and I'm not, and and so the two models you will see, one is reusing. If it goes to, we're a big fan of container labs, which is open. Uh, and if it's Nvidia, Nvidia has Nvidia Air and we can use that.
They actually push out of our orchestration tool con uh, uh, uh, configurations into, into Nvidia air model air. And then once you, good pun intended, you push and deploy to the switches and we walk you through. So that's what we do.
If you look at this, it like a slide, um, on the left side, really, we're partnering with either Venice, Nvidia, others, they have reference architecture, value design. We take them, uh, we can start building models in this tool, uh, validate 'em and then push 'em and manage the switches directly. Uh, and then really as I said, it's design, modeling, deploy and monitor.
Um, with two main strengths in the, in this tool, where we started was with this everything for Sonic, we would support Community Sonic. We support propriety sonic distribution as well. Anything that is Sonic we can support with this tool.
Um, and so I put on the right eye chart, but what I want try to say different deployments where we see Sonic today in which we're actually supporting the, the typical one is an enterprise data center deployment, standard leaves, spine topologies. Uh, you might have multiple fabrics. You might have a DMZ, anything you would expect from an enterprise data center.
We support this today. Uh, we can visualize the topologies, we can monitor, we can configure, we can push, we will demo this. Uh, and then the other one, what we're seeing is in what I call edge branch, think about your, in a retail location, you maybe have 10 to 15 switches.
Network is not changing because it's a blueprint for every one of that location. You can model this template, you can figure this out and then you're gonna push, push, push and monitor the deployments. So that's what we do with Sonic.
The second one, uh, we're partnering with Nvidia. We announce this, that GTC, and you probably have seen, actually I think this, you had somebody here. Um, it's around their spectrum X reference Archite.
If you look at what this is, and I have it on the, uh, my right, you're right too. Uh, if you look at what their reference architecture is for ethernet spectrum X is really a combination of a spectrum switch and a Bluefield, DPU snic and a blue field Nick, it's on the front end. Two different flavors of Bluefield Nicks, but basically D***s, um, very well defined.
And this was the question earlier, very well defined how this network needs to look and what the configurations are. What Andrad came and said, Hey, we wanna partner is can you automate the deployment? Can you automate the workflow for a customer based on very few inputs?
Like how many GPUs do you want? And what's the starting subnet for your nick Based on that you can define everything else, right? And you basically can do this.
Then the next step is what you want to do. You want to add tenancy. And that's why I say that earlier, right?
That didn't change, right? If you buy 256 GP PS it's most likely want to give 20 to one in your organization, 50 to somebody else, right? Can I do this?
Can I now do the network separation? Which is if you're a networking guy, straightforward, right? It's some flavor of v accident segmentation you're gonna do.
But you wanna automate this, right? You want to be able to, Hey, I wanna move this GPU from this tenant to this tenant. I should be able to do this.
And then the other one that we're gonna do is if you say, Hey, I'm running out on my capacity and need to buy another 256 GPUs, can I now extend that fabric without impacting the initial deployment? So this is what we're gonna do is with Nvidia spectrum, uh, x reference architecture. That's based on the current one, which one, three, which, uh, assumes an HEX was h uh, GPS in there.
But you might imagine, uh, more coming. And then the, the certain, this is more of teaser, we're not gonna show this today, but as I said, we have a second product, which is a network co-pilot, which is really more of a language interface chat interface. You can use this in conjunction with the standard orchestration tool to now ask question in English or Japanese or Italian, whatever your language of choice is.
And we have a customer that uses it in Japan. Um, and basically ask questions like, Hey, tell me what's going on in my fabric. How many of my switches are up?
How many are down? Are they all on the same west version or not? Right?
These kind of things. And you know, yak, is this something you couldn't do today on the ui? Yes you can, but you have to know where this on the ui and it might be not formed in the way you want.
If you can ask a question, you get the answer very quick. And so that's the integration we do with co-pilot.