49. Not All AI Infrastructure is the Same
Learn more by visiting the Networking Field Day Event Page on the Tech Field Day website.
Enterprises require vastly different infrastructure for AI. When building your next network, you need to understand what is required in order to achieve specific outcomes. In this episode of the Tech Field Day podcast, Tom Hollingsworth is joined by Scott Robohn, Brad Gregory, and Ron Westfall to discuss the various different types of AI infrastructure. They talk about inferencing and models as well as how to effectively utilize what you currently have. They also discuss what to look for when buying new equipment and how best to put it to use in order to maximize return on investment.
Transcript
The world is a wash in ai, and you are going to need to figure out how to use it. But before you can do that, you need to figure out what your infrastructure looks like. But I've got news for you.
It's not gonna be a cookie cutter, because as we find out in this episode of the Tech Field Day podcast, not all AI infrastructure is the same. Welcome to the Tech Field Day podcast, where we bring together a group of IT technical experts to discuss a single idea about key concepts in the enterprise IT industry. This podcast features a variety of perspectives from members of the Tech Field Day delegate community, and is often recorded in association with one of our events.
The Tech Field Day group is a part of the Future Room group, and this podcast is also published on our sister company Site Techstrong tv. On this episode, as we head into networking field day, we're gonna be discussing AI and infrastructure. I'd like to take a moment for our guest to introduce themselves before we jump into the premise starting with Brad.
Thanks, Tom. I'm Brad Gregory, um, founder and owner of MCN Strategies, where I try to demystify and simplify, uh, hybrid multi-cloud networking for the enterprise. Thanks, Tom Scott Roon, uh, co-founder of the Network Automation Forum, and host of the Total Network Operations podcast.
Thank you Tom Ron Westfall, research director here at the Futurum Group, heading up our AI networking coverage, including mobile ecosystem and host of the 5G Factor podcast. All right, and of course, I am Tom Hollingsworth. I am a practice lead for networking, security and mobility here at Tech Field Day, which is a part of the Futurum group.
Let's jump into the premise for today's episode. No doubt, you've heard a lot of talk about AI out there, and you're probably wondering to yourself, well, what exactly is AI and how can I make it work for my company? And that's usually where the real research starts to happen, because you're either going to have to figure out whether you're going to be training models or consuming the inferencing that they provide.
You're gonna have to figure out what infrastructure you need to purchase, and then you fall down in the rabbit hole of trying to figure out exactly which infrastructure matches which use case. And some folks get really confused and just kind of give up there. We're gonna make it easy for you today because I have a little secret.
Not all AI infrastructure is the same. So let's kind of ground it right there because we were having a great conversation about all of the things that people need to understand about why AI applications can be vastly different, even though they're all under the umbrella of those two little letters. So I'm gonna go open it up to our guest here to kind of help explain what is it that makes AI infrastructure so difficult to nail down when we're in the planning process.
Well, I'll take the leap here and, uh, and get us, get us started. You know, we, we talk a lot about, um, infrastructure for AI and network infrastructure for ai. You really have to split the use cases between inference, um, and training, model training.
And everybody loves to talk about model training because it's the cool thing that's technically very challenging and requires this, you know, very complicated synchronized swimming between hundreds or thousands or tens of thousands of GPUs, you know, participating in a weeks long or months long exercise where any error can halt the whole training process and give us real trouble. And it drives, uh, requirements like the use of infin band for extremely low latency, high reliability cut communications and the training process. But there's much more to it than training.
And I'll let Brad take that as the next step. Yeah, thanks. Uh, Scott, uh, I started my career as a, as a low balancer person, right?
Um, where we really had to dig into the application, you know, at the time they were very basic load balancers, layer for load balancers. Uh, basically was the service up or down, and it was up, it sent it, if it was down, it didn't, and then they evolved. They get a lot more, um, sophisticated, and then they become application delivery controllers, right?
Um, so when, when I think about inferencing, to me it's just, um, the same application delivery construct that we've had for years and years and years is just the, the next, uh, iteration of the complexity of what it has to do, right? So, um, you know, I think at one time as we went through the name, it, it was a low balancer, and before it became, uh, an application delivery controller, I remember people calling context switches, right? So it would look at the context of what was coming through and then switch based on that.
Um, so I've seen some stuff out there where a lot of companies, uh, in the load balance of business, they are, uh, looking at the context of what's coming through with the prompt, right? So if, if somebody asks a question about writing code and says, Hey, you, the, the best model to give you the information about the, you know, what you, uh, sent to me about writing, uh, code should be low balance to this model or sent to this model. Uh, if you're taking a trip to Italy, you need to go use this model.
So the, the, the inferencing part of it, there's a lot more than than just a load balancing. You know, there's security, there's everything that, that we associate now with application delivery. But, uh, but that's essentially the inferencing part of it.
It's, it's kind of the front end to get, you know, to get to the, to the train model. So, um, and, and, you know, I, I really think most everybody is gonna be, uh, dealing with inferencing, but not everybody's gonna deal with training because, you know, a lot of the average enterprises, they'll, they'll buy pre-trained models or, uh, they, they won't have the, the, the capacity to build, you know, these massive networks and these expensive networks to tie a bunch of GPUs together on the backside. So, so really, you know, putting those two things together that they're mutually exclusive, but obviously very complimentary.
And because you have to, um, deliver the whole AI application And to follow on, uh, what Scott and, uh, Brad shared, I had the good fortune of being at the Mobile World Congress 25 event, also known as AI World Congress. And you can get, that was a dominant theme. And to answer the question, you know, why is AI networking infrastructure still kind of like this open-ended aspect to the overall AI training and inferencing, uh, dimension?
And that is, you know, we're still fundamentally at the front end of the AI journey. You know, you talk to any operator out there, or enterprise or organization still, uh, yes, uh, the experimentation is still going on, but there is finally some tangible traction in areas where they are seeing improvement in areas, uh, such as productivity gains, improving, uh, coding efficiency and so forth. But, uh, linking back to, okay, uh, AI networking, what is it that's driving the decision making here is that the enterprises and the carriers are still trying to right size their models.
That is, you know, have data sovereignty prevail, making sure that their secrets will never, ever be leaked out, god forbid, on the public cloud or, you know, through, uh, you know, security attacks and so forth. And so there is this, uh, caution, but I think what is important is that when it comes to the AI networking, yes, certainly scaling is critical when it comes to the data centers where the GPU clusters are, you know, churning away, doing, you know, the heavy duty, um, training and so forth. And thus, you know, 800 gig is now pretty much, uh, expected on the roadmap.
6, uh, terabit is around the corner in terms of, you know, the, uh, networking switches and so forth that need to interact with the compute capabilities out there to meet these demands. But linked to that is automation, and that's something that I think we'll see some breakthroughs, not just in 25, but over the near future. As you know, training becomes vital in terms of how, you know, an organization can leverage these AI models within their own specific environments.
And also it's about the energy efficiency. Yes, the data centers are generating headlines. We need modular nuclear reactors to support all this AI training, but networking is not as critical in terms of the energy efficiency, uh, objectives there.
However, it needs to be viewed as a friend to, you know, what's going on in the data centers, how can networking contribute to the overall, you know, uh, ability for an organization not to have their sustainability goals blown out, uh, the water, because, you know, they're investing heavy heavily in AI training. So I'll stop there because I think those are like, uh, two of the theme main things that came out of that event. I think it's interesting that you bring that up, Ron, because you mentioned the fact that, uh, automation is gonna be king, but you also kind of talked a little bit about how, um, it's, people are still trying to kind of right size things to figure out what's going on.
And it's, it's a, a key topic that I think people need to understand as well, because it ties AI back to automation, back to SDN, if you will. And it's this idea that we have a shiny new toy, we just don't know what to do with it yet. So we're throwing it at the wall to see what sticks.
And you know, when you look at the first generation of AI applications that were leveraging things like LLMs, you know, they, they kind of honed in real quickly on being able to summarize documents quickly or being able to, you know, uh, auto generate, uh, information for ticket updates and stuff like that. Which reminds me an awful lot of the first days of automation when we were trying to figure out, well, what kinds of tasks can I automate? How much logic do I need to include in those kinds of discussions?
And then I think all the way back to companies like coho data, which took SDN and used to solve NFS Mount problems in storage, uh, which were novel concepts, but were they solving the right problems? And I think where I'm going with this is one of the issues that we're having that's leading to data center sprawl or uncertainty about AI infrastructure is we don't really know which problems we're trying to solve yet. We're just buying all of these new haswell clusters and deploying them somewhere and saying, alright, make AI magic happen for me.
I want to fix all of the things. Well, what do you want to fix? I don't know.
Do something awesome. Well, I'll, I'll throw in the, you know, maybe the semi, semi controversial statement that model training at scale would not be possible without automation built in from day one. And if you think about how our DMA remote direct memory access works and the coordination of thousands or tens of thousands of GPUs, manual operator intervention does squat for you there, right?
The, the, the software orientation. And, uh, may, maybe I'll even say, you know, it's a good thing that this wasn't designed by network engineers that really wanted to keep up on their CCI and all the iOS commands they've memorized, um, the, the, the heritage of AI model training coming from the software world and knowing that, you know, I've gotta make sure this all operates automatically without human operator intervention. We, we wouldn't be making the progress that we're making if that weren't built into the baked into the cake from day one.
No, I, I was gonna, I, I don't mean to put you on the, on the spot here, Scott, but, uh, you know, orchestration, totally get it right. It, and, and automation's gotta be there. But observability, I think is gonna be really key too because, um, you know, the, the way we monitor and manage networks where you pull in every 30 seconds, right?
I don't think that's gonna fly in a, in a, in a, uh, training infrastructure, uh, or, or a training deployment, right? Because the, it's just so critical, you know, the tail end latency and, and, and things that, um, you know, the ultra ethernet, the ultra ethernet consortium is trying to come up with new standards to, you know, do, do we get out of the Infin band world and all that and really try to make it an ethernet environment. So don't mean to put you on the spot, but how does, uh, the, the observability aspect need to change, or what will need to change so we can effectively monitor and manage networks, uh, that, that have different constructs than we're used to?
Sure. Well, I, I, again, um, I don't, I, I don't come here to praise and find a band, um, I'm not trying to bury it either, but, uh, you know, a lot of that is already baked into in Infinite Band and high performance computing for decades. Um, again, because of the tight application application coordination requirements, I don't think observability is viewed as a separate thing, right?
That we, we need observability more in a higher error environment, right? Um, and that is by design less of an issue in these very expensive clusters that we're asking to do, you know, again, weeks long and months long jobs. Now, to your point, um, for ethernet to get there, you're right, we've gotta add all sorts of observability error reporting, reduce latency and similar mechanisms in for, for it to get to the point where it can be good enough or just as good as what, um, InfiniBand can do today.
And I'll just take you back in history to, uh, you know, sauna infrastructure, right? Who knew, you know, that ethernet would replace this robust bell system heritage, you know, section line path. I've got reporting mechanisms for anywhere I could lose a bit.
Um, and that this sloppy, um, ethernet thing would be adapted for wide area links. You know, you asked me this 25 years ago, I, I, I remember seeing it as part of the Juniper Tech. Um, there's no way Ethernet's gonna take over, but you know what, we added all the right OAM capability to do exactly the kinds of things you're saying, um, to make it an effective replacement for sauna.
And, you know, previously TDM WAN links, that gives me hope that it's gonna be just as adaptable for this in the large scale ai, um, training cluster environments. And one thing, Scott, to add on, uh, what your observations are, InfiniBand, you're right, that's gonna be around for a while. Uh, it's going to increasingly, I believe, coexist with ultra ethernet.
I think, you know, there are plenty of organizations out there. Uh, yes. You know, InfiniBand has proven itself, it's reliable, but we wanna, you know, migrate toward a non-proprietary type of approach.
And so, yes, uh, we've seen this before. It will be years in the making as this, uh, journey, uh, evolves. And I think it's tying to something, uh, that's even, uh, more broad in terms of ecosystem impact.
Uh, Tom invokes LLMs, and we saw Deutsche Telecom invoking perplexity as the LLM AI agent that will be in their AI smartphones that they'll be selling later in the year. So I think we're gonna see more, you know, these endorsements of, you know, which LLM uh, do you wanna bring, uh, to your device? But I think it's also pointing to the fact that, okay, LLM is like, so six months again, now we have a agentic AI and multimodal AI to add onto all of these, you know, scaling challenges and ai, workload distribution challenges.
And I think the bottom line is that the world will be hybrid ai. Uh, we'll definitely have training in the data centers, but inferencing will be increasingly distributed at the edge, including the far edge that includes smart devices, you know, iPads, what whatever is going to be best suited for that, uh, optimization of the ai, uh, implementation, uh, that is, uh, required. And so I think that's something that will, I think fuel, you know, the idea that, uh, inferencing is going to be something that's gonna just grow, you know, in the double digits on a pager basis, uh, for quite some time.
Uh, also, you know, correlating strongly to what's going on with AI training. Ron, don't you think a lot of the, the, um, all the focus right now has been on training, right? Because you got a training before you can get it out to inferencing, but, um, I, um, what are you seeing out there?
Are, are you seeing, uh, people really understand the inferencing, the, the need that's coming, how it's gonna be distributed, how you have to secure it, all these endpoints out there? Or, or is everybody still living in training land right now? Excellent question, Brad.
And if mobile, while Congress is any indication, I'm gonna say inferencing is definitely becoming more critical and integral to the decision making and, you know, understanding and awareness of how do we optimize AI for our organization. You certainly talk to somebody like Qualcomm and they're already, you know, you know, inferencing is the way to go, but it's more than the device, uh, manufacturers or the silicon providers, you know, supporting, uh, the device manufacturers. You talk to, IBM, you talk to Dell, you talk to Lenovo, you talk to HPE, they're talking pocket to cloud, as you know, uh, essential to how, you know, AI is going to work for an enterprise.
That is, how can it be monetized? How can it, we get beyond, you know, simply using it for, uh, you know, say some productivity gains, but making it, you know, just part of the overall DNA of the organization. So my vote is, yes, inferencing is definitely now essential in terms of when you talk to anybody who's not exclusively or heavily focused on AI training in terms of how do they want to make AI work best for them, and how to optimize those AI workloads.
So I will see your comment, Ron, and I'll raise you, you know, look at the, the reported stats on, you know, the, the tools that have come out and how long it's taken them to get 5 million subscribers, 10 million subscribers. We're all users of inferencing. 9% of the people listening are watching this.
Um, we're, you know, we've all inferencing has been taking off that we get attracted to the shiny object of building this huge industrial strength AI networking infrastructure, which is wicked cool, right? No, no question. But the quiet unsung hero, so to speak, has been the adoption of, you know, the models that are out there.
And I'm gonna very selfishly pivot this a little bit, and I talk about this as the million model problem. We have so many new models and such an increasing velocity of more, more models that I can use if I think about just network operations scenarios. 'cause that's where I live and what I, what I focus on, how do I know what models are the best models to use for real NetOps use cases?
I'd love any comments anybody has on that. Well, I, I think, uh, uh, one part of the conversation is, okay, how can we use, you know, AI agents or, you know, chat GPT to help with, you know, overall, uh, skills, uh, within the organization. That is how can a, an operator or an enterprise, uh, better optimize, you know, the AI capabilities for NetOps or for SecOps and so forth.
And so I think it's, uh, there's multiple dimensions here that actually, uh, get touched on with each implementation of ai. Uh, that there's certainly the training side, which we've covered, but there's the, the inferencing side, but also it's like, okay, how will it impact, uh, the, the workforce? And I think what we're gonna see is an evolution where the, uh, people who are using, uh, AI to improve NetOps, both in terms of, you say broader auto automation organization-wide automation that, you know, linked to traditional machine learning capabilities, but also, uh, quite simply using those agentic AI agents to figure out, okay, how do I make sure I'm, I'm using, uh, the, the best approach to make net ops work better for the organization.
And so I think that's one practical, uh, aspect here as to, you know, how, okay, with all these models out there, how can it practically support, uh, net ops improvements? I think that is certainly one, uh, way that, uh, that is becoming, uh, more attractive in terms of, you know, what the decision making is driving out there Out there. Can, can AI learn enough where we don't have to have, uh, as many network agents all over the place, you know, do, do we need to, to capture and collect as much data at all these data points?
If AI can infer a lot based on what it knows over time, that was always one of the, the frustrating things for me. I, you know, and in the, I'll say the old way, probably is still the current way of doing things. You had to collect so much data and, and you were just awash in data.
It was so hard to correlate. It was so hard to pinpoint down to the problem, ironically, right? I mean, I've got all this data and it's so hard for me to wade through it.
Hopefully AI can, you know, figure out what's going on without needing GOs and GOs and GOs and goss of data to be ingested into some tool. Well, Brad, let me give you hope, right? So, so the emerging use cases that I see for very practical NetOps, um, applications, number one, ingest of log messages, you know, whether it's just plain old cis log or streaming telemetry.
Um, and in whatever volume, you know, you, you may have available to you, um, you can take a set of messages and say, tell, describe the behavior to me. Like I don't understand how to parse all these log messages. Tell me what network events really happen.
Boil it down to me. Gimme the key 3, 4, 5 events that happen based in, you know, this, you know, 50 pages or 200 pages of, of log blast at me, right? That's a really popular application.
I also see people throwing packet captures into chat tools. You know, okay, what's, what's anomalous about this behavior, right? You know, is there A-A-T-C-P connection that's being broken?
And I have, you know, I have an open port that's inactive, things like that. Um, those are, you know, these, and, um, a set of another, you know, couple dozen emerging use cases are out there. And the question I'm trying to get to is I've got a million models to choose from.
You know, should I be using Claude or chatt PT or Grok or, or what have you? How do I know which models are gonna be the best ones to use? Um, and I think, I think that's a maturity level that I love to see us drive toward.
I think there are lots of people out there doing great experimentation and being really open about it. Um, and, and watching them hack their way through the jungle of AI is inspirational and incredible. But we also need to get to, okay, how do I teach an organization?
Here are seven things you can do in your NetOps team using these models. You probably need to refresh that in a year, right? Uh, and, and then less than a year that, you know, those models might change, but those are the kinds of things that I wanna see the NetOps community get to.
You shot a flare up for the integral role of data management, because one thing hasn't changed. It is data in and data out. If you're putting in garbage, it's not going to make a difference how elegant or advanced your LLM is or your agentic AI implementation is.
And so I think this is going to shine a spotlight on, okay, how can we make this, uh, more effective for the workforce? How can improve the workforce experience, but also the outcomes for the organization? It's at the bottom line is we need this to improve business outcomes.
And so I think this is a very valuable, you know, okay, what is it that NetOps is going to do? I, I like the anomaly detection. That was definitely hot at the show in terms of, okay, security needs to be wrapped around this.
And actually, uh, before I forget the top decision making criteria, that is when, uh, folks are selecting, I want a new, uh, solution, AI expertise is at the top of the chart. So there's just no way around this. It's an AI world, a hybrid AI world.
And so yes, you can have great, you know, wifi access point technology, observability capabilities and so forth, that has to tie in to an end-to-end AI proposition, or you have to convince them that you know what you're talking about when it comes to the AI dimension. Over to you, Brad, Um, pivot a little bit. So back to inferencing, what is a good starting point?
You know, there's a, from an infrastructure standpoint, an architecture standpoint, because, um, there needs to be a baseline, I think, where we, where we take the, the, the current best practices, the current ways we design things, use that as a baseline to say, this is what I have to have now. And, and you know, I think back, uh, when we went to software defined networking, right? There was a baseline of networking that happened way before software defined networking came along and said, okay, to be credible, you have to do these things at this point, right?
Um, so when, when I think of inference, again, I go back to application delivery security. You know, Ron, you brought up, um, you know, what happens if people start trying to leak data out of, um, of ai, right? When, when you, when you prompt it, when you send a prompt over to it, well, that's data loss.
You know, we've, we've dealt with data loss for how long, right? We've got tools around that. So, um, I I, I would say, you know, for the average enterprise on the inferencing side, use application delivery as a starting point right now and say, I need to, I need to secure ai.
Interesting. The same way as I secure an e-commerce app, you know, at least start there, right? I need to deliver it the same way, right?
I need to be able to context switch on certain things, you know, uh, uh, cost is a big thing with this, right? So certain prompts come through, use a model that's not as expensive when you hit it. This is really important.
I'm writing code. You may have to hit a more expensive model, you know, when it comes through. So, um, I I, I really, you know, I am gonna belabor this point a little bit that I just see AI as being a another application, uh, that we need to serve up to users and all of the things that we've learned over the last 20, 25, 30 years of securing an application, high availability, um, you know, the application delivery aspect to users can be applied here, right?
That should be the starting point. No, no, no, Brad, it's the new special snowflake, and we, we need to reinvent everything around it to secure it. Yeah, no, and I'm of course, just poking, right?
Totally, totally agree with you. Um, we don't need to reinvent the wheel here, and like, as there's increasing awareness of, uh, you know, again, pseudo controversial, you know, maybe, maybe, you know, chat, um, AI tools are really the next search engine, right? Um, that's certainly a very broad use case, but you know, again, a lot of the same security principles can apply, but, um, we need, we need the right, uh, corporate policies to say, okay, here's, here's what you can do with chat GPT, here's what you shouldn't do with it.
And oh, by the way, in the, the NetOps examples I pulled out, you know, you put a packet capture or, or, uh, log stuff in, uh, for analysis, you are probably putting, um, important information from your network IP addresses used, maybe other distinguishing information that you need to think about before you actually pump that into a public tool. And that triggers roadmap. Uh, what is it that is being prioritized in terms of how can we make this work, uh, inferencing?
It's here to stay and it's important because of security. It's, you know, preventing, you know, uh, more transport of data to the public cloud, or at least, you know, minimizing it. Also, it's about network APIs and, and encrypting them and securing them along with more focus on prompt poisoning, how to again, minimize it and avoid it.
And so this is all being, uh, talked about and thought about, but all this is a part of that all critical AI journey for all the, that organizations out there. Well, as you can tell, not all infrastructure is the same. And you've probably been thinking about infrastructure as just the mere hardware that goes into building these AI systems, but as you can tell, there's more to it than that.
You have to have a solid use case. You have to understand things like security and data hygiene. You have to understand applications of ALK manner of other things.
And if you haven't already been thinking about that, now you know what you need to do. And these are the kinds of conversations that you really need to have before you jump in with both feet to use ai, because I'm sure that every enterprise is littered with the burned out hulks of some paradigm shifting technology that was a brand new gadget that had no use case, and sits in a corner now quietly collecting dust until a time when it can be carted off to the recycler. And if you don't want that to be your shiny new Haswell water cooled AI spectrum, X ethernet, ultra ethernet whizzbang thing, then you need to make sure you have the checklist filled out before you ever hit purchase on that po.
Thank you very much for listening to this episode of the Tech Field Day podcast. If you enjoyed this discussion, please make sure that you subscribe on our YouTube channel or in your favorite podcast application so you don't miss an episode. Please consider giving us a rating and a nice review, uh, so that everybody knows what we're all about here at the Tech Field Day podcast.
Uh, this was brought to you by Tech Field Day, which is the home for IT experts across the enterprise, which is a part of the Futurum Group. com/podcast or just check us out on Techstrong TV where we have an archive of all of our episodes. Wanna thank you all very much for tuning in, and we'll see you next week.