23. Etnernet isn’t Ready to Replace InfiniBand – Tech Field Day Podcast
AI networking is making huge strides toward standardization but Ethernet isn’t ready to displace the leading incumbent InfiniBand yet. In this episode of the Tech Field Day Podcast, Tom Hollingsworth is joined by Scott Robohn and Ray Lucchesi to discuss the state of Ethernet today and how it is continuing to improve. The guests discuss topics such as the dominance of InfiniBand, why basic Ethernet isn’t suited to latency-sensitive workloads, and how the future will improve the technology.
Transcript
InfiniBand is king when it comes to ai, but a new challenger has appeared on the horizon. Ethernet will the young, scrappy upstart be able to unseat the unparalleled king of all things data communication. In this episode of the Tech Field Day podcast, ethernet isn't ready to displace ethernet yet.
Welcome to the Tech Field Day podcast, where we bring together a group of IT technical experts to discuss a single idea about key concepts in the enterprise IT industry. This podcast features a variety of perspectives from members of the Tech Field Day delegate community, and is often recorded in association with one of our events. Tech Field Day is a part of the Futurum Group, and this podcast is also published on our sister company Site Techstrong tv.
On this episode, we will be discussing ai, but before we jump into the premise, I would like to have our guests introduce themselves, starting with Ray. Yes. Uh, my name's Ray, Ray Lui.
com. com. I've been in the storage industry lots and lots of years.
Hey everyone. Scott Roon. I am your friendly neighborhood internet plumber.
Apologies to Stan Lee. Um, I've been very fortunate to kind of fall into a networking career, uh, starting back in the early nineties, uh, been through a variety of operations, uh, and technical sales roles, um, and now an independent consultant involved with the founding of the Network Automation Forum. Um, and a new initiative called Total Network Operations.
Alright, and my name is Tom Hollingsworth. I'm an event lead here at Tech Field Day, which is a part of the Futurum Group. Let's jump into the premise for this episode.
If you're talking about AI infrastructure, there's a lot of things to dissect, but one of the most important ones is InfiniBand. It is the technology that is running the backend networks that all of our inferencing and model building happens on, but InfiniBand is old. It's the way things used to be.
The future is ethernet, just like it always has been. There's a way to make ethernet do exactly what we wanted to do, or is there, in this episode of the Tech Field Day podcast, we talk about ethernet isn't ready to replace InfiniBand today. Now, before we jump into that topic, I think we kinda let need to let everybody know what is it about InfiniBand that makes it so useful for these backend communications?
I mean, Nvidia basically built their AI clusters because they have access to InfiniBand. Yeah, yeah. I I think it's lossless and, uh, non-blocking technologies, which makes, makes a big difference there.
The, the, the protocol stack also bypasses the kernel, right? So that low latency, by not having to be processed by the kernel. Understanding RDMA kind of helped everything click for me when I was starting to learn about this.
You know, every GPU has memory co-resident, you know, on whatever it's mounted on right there in a, in a particular box. And our DMA wants all that memory pool associated with all the other GPUs to look like a single direct memory access pool. And to make that happen and to keep the GPUs cranking, whether it's LLM model training or what have you, you need that low loss, low latency, um, fabric interconnect to make it all seem like one big cooperative machine.
Yeah, we were at, uh, Juniper here, uh, uh, for Cloud Field Day 20, uh, probably a month or so ago. And they, we were making a big pitch to try to, uh, use their FIN event, fin event ethernet technologies to support, uh, the backend GPU network. Uh, I mean, according to Juniper and, and many people there, there are three networks that are associated with AI activities.
The front end network typically all ethernet. There's a backend storage network, which can be ethernet or FIN band. And this is a backend GPU to GPU network, which historically has been infinite band and, and, uh, Juniper was making a pitch to say that we can possibly get there.
There are some problems, but, you know, with, with various, uh, capabilities that they talked about, uh, they have, uh, they believe they're close. And just so that everybody is clear, that having talked to people who are developing these new versions of ethernet, um, Jay Metz of the Ultra Ethernet Consortium specifically, there are three very distinct networks that, that we talk about. So the backend and FEA band network that communicates amongst the GPUs, that's a type three network or mode three network.
The network that's used to send the data to the clusters to do all that work is a type two or node two network. The way that we interface the network to do that through laptops and other client devices is a type one or Mode one network. We're never gonna touch the mode one network that is ethernet and wireless, and that problem is solved.
Mode two is where we're doing a lot of development of ultra ethernet today because we need to be able to feed the monster as quickly as possible. And initially when pe, when people were just trying to describe whether they were going to be displacing ethernet or InfiniBand for that, there was not a whole lot of talk about the communications between the GPUs on the backend until sufficient ethernet technology was developed in order to be able to overcome some of the latency issues and some of the loss issues. 01% of my packets, I'm totally fine and I don't care.
01% of my packets, I'm freaking out. Yeah, yeah. Juniper mentioned a couple of techniques.
They have, uh, identified dynamic load balancing flow let control, and there was one or two other ones that, that they believe will lead to, uh, less loss, but not, not completely lossless and, uh, you know, ECN, which I guess is, uh, ethernet, uh, con Explicit congestion notification. Yeah. Yeah.
So you're, you're finding out that which paths are, are congested and that, and using the other paths as, as at the switch level effectively, or even at the, well, probably at the switch level, maybe at the port level. I don't know. The, I think the, one of the huge encouraging efforts here is the ultra ethernet consortium, right?
That's coming together with multiple vendors trying to figure out, okay, how do we extend and modify ethernet to be appropriate for that GPU interconnect use case? You know, I, I go back to 25 years ago when I never thought sonnet and TDM infrastructure would, you know, and the rich fault management fault identification capability that's built into sonnet, um, could be replaced by ethernet. And, uh, boy was I wrong, um, took time.
But all the OAM, um, all the other fault detection mechanisms that we need to use, you know, ethernet, gigabit ethernet, 10 gig ethernet, um, for a wide area networking protocol. I mean, it's obviously, you know, totally taken off today. And while I think InfiniBand has a lot of momentum right now, install base, um, single supplier for all your GPU and interconnect needs, don't bet against ethernet in the long run.
Right? It's, it's been proven to be adaptable for however many decades now. And I think, we'll, we'll continue to make progress here too, through specific vendor innovation and what gets standardized in the UEC.
Uh, and the other thing that, that Juniper mentioned, they had, uh, they had recently announced an 800 gig per second, uh, gigabit per second interface, but they were talking the backend GPU network. It's typically four oh gig bits per second. And it's, uh, you know, obviously from an fin band perspective, lossless non-blocking and, and low latency.
Yeah, I, well, part of, part of the thing is I've talked to a lot of people who've said, don't bet against ethernet, but I guess my question is why? I mean, we have InfiniBand, it works really well. The, the upper limit for node scaling is somewhere in the neighborhood of nearly 50,000 nodes.
And I don't necessarily know that we're hitting that headroom right now. What is the allure of ethernet? And before you come at me with, oh, well, it's cheaper, I'll give you, I'll give you a, a non cheaper answer, and I'll say, it's an excellent question.
Um, who's the dominant GPU supplier today? Nvidia, Where can I buy InfiniBand Nvidia, Where else can I buy InfiniBand Nvidia, Right? So there are many enterprises, you know, that are interested in building their own large language mod, LLM training infrastructure that want competition, that don't wanna be, you know, stuck with a single integrated stack, um, from a single vendor, you know, 'cause that gives you certain leverage over your customers.
But consider the fact we're talking about thousands of GPUs spread across thousands of, of servers that are all trying to, you know, do direct memory transfers from one to the other at the fastest possible rate. Can infinite keep up? Uh, I don't think so.
And that, that's part of the problem that we're running into, is that we're, we're looking at two distinctly different versions of ethernet. Um, and, uh, founder of tech field, a and, uh, you know, future and group, uh, coworker Steven Foskett loves to quote something that Bob Metcalf, the founder of ethernet, told him one time. Um, I don't know what the future of networking is gonna be, what's gonna look like, but we're gonna call it ethernet traditional 10 megabit, half duplex ethernet looks nothing like the 800 gigabit ethernet that we're installing in these systems to be able to run them at speed.
And yet it's all ethernet under the hood, except that it's not. You know, Scott, you kind of alluded to the fact that a lot of the, uh, administrative infrastructure that we would need to see here is, you know, wasn't present for the longest time. Things like OAM and and, uh, error correction to detection.
I mean, when you look at the way that this network would be built on the backend, it's not ethernet. It's a fabric that's running all kinds of things like, you know, uh, really weird ECMP packet spraying, um, very explicit congestion notification mechanisms because ethernet by its very nature doesn't like to work the way that InfiniBand does. It's not a reliable transport, and yet we're trying to do everything we can to make it do that, and like holding on by our fingernails to claim that it's ethernet.
Is this really what we should be doing? So I'll go back to the single source issue. I think competition, um, for a technology brings a lot of value.
And the InfiniBand ecosystem doesn't bring that today, the ethernet ecosystem does. And, you know, we can, we can dive into the weeds on, you know, what that transformation of ethernet has been from, you know, pre 10 meg half duplex to something very different today. Um, that's okay.
There's a lot of the ecosystem and chip design, um, business that's common enough where building all that gear, even though it's that yet another modification on ethernet that's that's known and it's, it's well understood. And if I can get it from 4, 5, 6, or more major vendors, that does incredible things, you know, from a technology perspective. I think the other thing about this is that a lot of that sophistication ends up being in the switch rather than being in the port.
So, uh, ethernet ports are relatively not that intelligent compared to things like, uh, infinite bat and others and stuff like that. So the ports are cheaper and the, the ports being cheaper allows, uh, you know, you to configure a 400 gigabit per second, uh, network, you know, 30, 40% cheaper than you could with an INFIN band, something like that. I, I think there's some value in that whole intelligence where, where that happens.
Because that's one of the reasons why a technology like RDMA works is because the card is the important part. The, the switch, if you want to call it that in InfiniBand, is effectively just kind of a dumb optical interconnect. Uh, that's one of the reasons why it doesn't scale very well, is because once you hit that, that upper limit of, you know, I think it's 48,000 nodes, um, there's literally no more space to interconnect those things.
And that's when you have to go to effectively like a, a, well, I guess it would be a two stage or possibly even a three stage CLO fabric at that point. Yeah. All ethernet nerds out there are probably freaking out.
'cause you know what this looks like, but I mean, one of the, the things that we keep talking about is this cost effectness, does this decision, is it impacted by the fact that we're still trying to figure out if it's better to run this technology on-prem versus in the cloud? And would a cloud provider have a different way of looking at this? Certainly whoever's hosting the GPUs, that's their decision, right?
Whoever's building the infrastructure for that. Um, I think, you know, if I'm, if I'm not gonna build my own AI networking infrastructure, I'm just gonna use cloud services, I kind of don't care, right? As long as I get the performance.
And you're certainly gonna have more enterprises using cloud-based AI services to build models. Um, I think just by the numbers, I don't know if I would, I would say that by revenue perspective, but you'll have more enterprise customers engaged in doing stuff that way versus building their own AI data centers. I I, I think you have to make a distinction here between LLMs and, and everything else, quite frankly.
'cause LLMs are, are order of magnitude, maybe two more, uh, scale, uh, scalable or, or, you know, have, have the scale configuration that that normal, yeah, the whole AI world at the LLM training level is a different, it's like, it's a different world than anybody else, whether it's done on-prem or, or in, in, in the cloud. I think doing a LLM model in the cloud would cost you four or five of your appendages if you have them. So maybe the, the better question there then, Ray, is if we have a purpose-built communications infrastructure that adapts well to AI workloads, whether they be LLM or, or other kinds of things, why are we needlessly complicating it just so we can say it runs on ethernet?
The other thought I had to some extent was that the predictability of the workload matters. I think ethernet can handle a more predictable workload, whereas, uh, InfiniBand can handle a more unpredictable workloads. And, you know, if you're a cloud provider, uh, you're probably gonna have to deal with this unpredictability more so than, than than anything.
So I think you might go with an infinite band in that case. But, um, I think, you know, the, the expectation and even Juniper and his discussions of cloud Field A 20, we're talking about relatively small models, you know, in doing, you know, fine tuning or RAG or something of that nature on those models and, and, and doing some training of dlms and things of that nature. But nothing, nothing to an LLM level, uh, because I mean, the scale is, is two orders of magnitude more.
I mean, you know, I can put together a an A GPU system with an expensive configuration, but it's, it's not gonna have that much of a problem, you know, moving data between the GPUs and the same server to some extent, doing that across a thousand GPUs or, or 8,000 GPUs across 1000 servers. Another question entirely. Well, certainly one size never fits all right?
Um, so the emergence of different solutions for, for different scale, you know, we've seen this movie before. Um, Tom, I forget the way you asked the last question you posed, but, uh, I, my response would be, you know, the, the ethernet switch manufacturers wouldn't be going after this if they didn't think they could handle, um, some of the, the, the, the more stringent performance requirements. Of course, it's a growth market, right?
And they wanna go after areas where they, they think they're gonna be able to, you know, continue to grow market share and keep their, uh, investors happy. But, you know, I think any of us who've dealt with being on the other side of seeing requirements come in and what it means to work with product management and engineering to implement something, those, those decisions never get made lightly. So I think it's safe to say they see promise probably in different, um, levels of achievement, but, uh, you know, they're going after it intentionally and think that they can, It's all innovative dilemma to some extent.
You go at the low end of this, of this stuff and try to build a market from that perspective and then move up. Uh, so I, I think if, if ethernet can handle a relatively small training activities for, you know, DRM or, or, or minor LLM models and things of that nature, it makes sense for them to go after it with the hope that ultimately they can go after the big prize, which is the LLM training cluster, Right? And that's what everybody wants, ultimately at the end of the day, is they want to be able to get into a piece of this huge pie that Nvidia has been, uh, baking, you know, somewhere in the neighborhood of $3 trillion.
But we talk about this a lot on the, on the chip side of things. Nvidia has a huge headstart. They went all in on this sure.
Years ago, and they've slowly been building out InfiniBand, and we know that they're working on their own alternatives. Uh, spectrum X is one, uh, that heavily leverages Bluefield dpu and, and some of the ethernet properties that they've gotten from their MRL acquisition, which ironically enough is where they got the InfiniBand from, right? And, you know, they're, they're also working on using NVLink as kind of the, the proprietary networking protocol interconnect, um, because it keeps things in the family.
But my question is, if we know that Nvidia is developing their own version of ethernet in-house to replace Infin Band, or at least be an alternative to it, is there a hope that a standards based ethernet version will ever be developed to kind of break into that market? Because if it's the situation that we've had for years and in a lot of other technology areas where there is a clear market dominating leader who consumes all of their own champagne, dog food type stuff, what hope is there for everybody else? They're just gonna be basically getting the table scraps.
I I think UEC is a, is a path forward in this respect if, if UEC can, can conquer the personalization of, of, of networks for, for AI, a large language model training centers, and, and allow it to approximate the low latency, approximate the lossless approximate the non-blocking characteristics of Infinite band, and it's a standards based solution, and any of these guys go out and adopt it. Uh, that's the question. Can it do it ai, you have to talk to Dr.
Met his team to see what the story is there, right? And, and the existence of, you know, the efforts on, you know, ultra ethernet and the pressure that can put on development of the in fin band ecosystem, that's a benefit too. If, if, if all it does is make Infin Band better, you know, from a cost per bit perspective, that's a positive outcome too.
I, I think the challenge with UEC to some extent is it, it is a, is it's a, it's not, it's, it's not a, um, not an easy thing to plug into this world, right? I mean, it costs, you have to switches and you've got, you know, maybe even ports that have to be more intelligent, et cetera, et cetera. So, I mean, there's, there's things that have to go on in order to adopt it.
Can it be adopted? Sure. Um, and at what price does it compare to InfiniBand at that level?
Yeah, I don't, I don't have a crystal ball. I don't know if either of you two, be careful. It might be a Palantir.
Um, you know, I I, I can't tell you, you know, three years from now what the slider Bar is gonna look like for, you know, uh, improvement, uh, and e ethernet, um, grabbing market share here, but it seems hopeful. Um, you know, a year from now we'll probably have a better idea and seeing how well the ultra ethernet consortium's actually doing. And that's a challenge that we have right now, to be honest with you.
Like you kind of alluded to, they don't have anything like, and, and I'm not speaking out of school here. I've talked to a lot of those people, and they're still kind of in that forming stage where, sure, we've gotta get things together and we have to figure out what our direction is because we'll use packet spraying as a perfectly good example. Um, Cisco is a champion of packet spraying because they have silicon that can implement it.
Broadcom is also a champion of packet spraying, and they have a, uh, group of silicon chips that can do it. But those two companies, even though they're using the same technology by name, they're not using the same implementation of it. And I mean, this is the thing that we've, we've dealt with over the years in networking.
Um, you know, Scott, all I have to do is say things like ISL trunking Sure. Or Cisco P oe, and you know exactly what I'm talking about, right? When you have competing standards, there's only one way to, to solve this problem.
And of course, it is the tried and true tradition of picking one company and deciding to do the opposite of whatever it is they said, because you just don't like them for some reason, right? So can, can the UEC finally say with certainty that we will get to a point where we will have a standard for people to adopt, or will everyone just kind of agree to go their separate ways because, well, I didn't get what I wanted There. There's always the risk of that, you know, taking your ball and going home.
Um, but you know, this is where I'll, I'll get super optimistic on you and say like, I trust the Silicon Valley ecosystem to produce useful solutions more than I trust any one company. Um, and, you know, for every negative POE or ISL example that you can cite, Tom, we can pick other things too where, you know, you start with a, you know, a standard, uh, an RFC for example, somebody has a better implementation or an option that seems to work the best, and that becomes a defacto standard or the most commonly implemented way of doing a thing. Um, you know, look, we can see things like what started out as VRRP became HS rrp mm-hmm.
Right? Became, became interoperable. So there are, there are some things to be hopeful, um, about from, You gotta look at this.
This is trillion dollar market here. I mean, these guys have got dollar signs in their eyes if they can just agree on something, go after it. I, I don't see, I, I don't see this, uh, bifurcation being a, being a being a problem, but, you know, I'm not a networking guy.
But really, are we all networking guys at this point? Because all we're doing is building a slightly different version of ethernet in service of another application. I mean, Scott, you and I have been through this a number of times.
You know, networking is dead. It's S-D-N-S-D-N is dead. It's this other thing.
And, and this, this cycle repeats itself. And my biggest problem with it, if you wanna call it that, is what happens after that? What happens once we've built a purpose built network that runs on a custom ethernet version and AI hype dies down, it fizzles out.
Is there a way to salvage what we've taken out of ethernet and leverage it for something else? Or is it so purpose built that this can never be expanded and kind of has to be jettisoned like the ATM lane networks of old I I, I think the UEC has got, um, a couple of different personalities it's going after. It's not just ai, it's HPC and it's, it's a couple other, you know, big pharma and things of that nature.
So, so I think it's, it's got the potential to handle multiple different styles of, of usage. And if, if AI turns out to be a bust, which I personally doubt, but I I could be wrong, uh, I think these other solutions will still be there and, and, and still have need for networking personalization. And Finman is big in HBC has always been big in HBC, and, and the reason is lossless non-blocking low latency.
If, if they can address the HBC requirements and the AI requirements. I think UEC has a potential here, but I'm not a UEC spokesman. So, but let me, let me play your argument back to you, Tom.
Um, so you, you basically just made the same case against InfiniBand InfiniBand, you know, it goes after a small subset of use cases, but there's a market big enough that sustains it, right? Um, this, you know, development for, you know, advanced enhanced ethernet via UEC or other mechanisms doesn't mean that the rest of ethernet goes away. It's not a mutual exclusion.
So, you know what, maybe, maybe there's wasted development effort, um, in it if it pans out. Um, you know, we'll see. We, we've obviously seen that before too, and you just triggered me and you run in my trauma with the lane, um, less bus LECS and proxy ls.
Thanks a lot for that, Tom. Um, you know, some technologies work are useful for a limited period of time, and, and Lane was an extremely limited of time. Yes.
But let me, let me kind of expand on that. You, you talked about the fact that in finna band is kind of, you know, it it has this market and it's, it's very focused. And part of the reason that I'm gonna argue that InfiniBand is still around is because of ai.
If InfiniBand had been a niche solution for HPC, I don't think anybody really would've cared about it. It would've been, you know, SNA or it would've been any other thing that we've developed over the years that frame Relay, I mean, it's deployed widely in a very specific case, but AI networking kind of brought it to the forefront and effectively made people sit up and say, I can displace this. But if ultra ethernet doesn't pay off, a lot of companies are gonna be frustrated that they invested a lot of money and didn't get anything out of it.
If InfiniBand recedes into the distance, it still has HPC to fall back on. Sure. I mean, I'm, I'm, I'm, now that I've triggered you, I'm gonna trigger Ray next.
Um, we fought this battle with Fiber Channel. Sure. I remember everyone was a champion for fiber channel over ethernet, and does anybody know what happened to Brocade?
Yeah, they're still around. They're owned, they're still making Fiber Channel still exists. It's still, it hasn't gone away.
A fiber channel over ethernet. Yeah, not as much. But, but the point is, is that those specialized technologies that we developed over the years are still doing their specialized jobs.
They're not trillion dollar industries by any stretch of the imagination, but do they need to be? Yeah, I, so HPC is a big enough market in and of sale in and of itself to continue in fin band for a long time to come. If anything, a lot of your customers, a lot of our customers nowadays are looking more and more like HPC environments.
They're doing a lot more scientific, a lot more data analysis, a lot more big data types of stuff that's, that's going on in these en environments that are looking just like HPC did or does. I think it'd be a really interesting set of discussions on the floor at like supercomputing 2024, like November and Atlanta, right? Ask people who are, you know, that's the heart of the HPC community, you know, what are they thinking about?
Do they see, what are they thinking about testing, um, ethernet solutions for what they're doing? Or are they wholly committed to saying in Finman works? And I'm just gonna stick with it.
There, There are, there are a couple supercomputing systems out there that, that are implementing ethernet only environments. Not a lot, but there are certainly some, You're getting that pulse like on an, on an annual basis, right? Is the needle moving at all?
Well, I was gonna say, we also have another event coming up similar to Supercomputing, where you can talk all about the AI data infrastructure that goes on behind the scenes. And that, of course, is AI data infrastructure Field Day, which will be happening October 2nd and third. com and learn more about that, including seeing a lineup of the delegates who are gonna be there as well as the, uh, presenting companies.
And I'm sure that they're gonna have a lot to say on the subject just like we have had here. Yeah. Yeah.
I mean, to some extent they're more focused on the, the data aspects of it rather than the, the networking aspects of it. But yeah, for sure. As you can see, the premise is solid.
Ethernet's not ready yet, but as, as our guests have talked about, there are advances along the way that are getting ethernet ready to go. I don't know that ethernet is going to displace what InfiniBand does today, but I also don't know what the future of AI looks like. And it could be that when we develop enough technology to accelerate those workloads, maybe the developers start writing for the tech that we have on the horizon instead of the tech that they're saddled with from days gone by.
And in that case, I think that ethernet is the clear winner. If only because ethernet is the chameleon of the networking industry. It always takes on the characteristics of whatever it needs to do.
It is very flexible. There is no rigidity to it at all, and ultimately it wins. I want to thank everyone for tuning into this episode of the Tech Field Day podcast.
If you enjoyed this discussion, please subscribe to our YouTube channel or subscribe to the Tech Field Day podcast in your favorite podcast application so you don't miss an episode. Consider giving us a rating and a review because that is a great way for people to learn a little bit more about what we do here. com/podcast or watch us on Techstrong tv.
Thanks for listening, and we will see you next week.