15. On-Premises Networks Need to Work Like Cloud Networks – Tech Field Day Podcast
On-premises networks are still very common for specialize applications and need to adopt cloud network operational models. In this episode, Tom Hollingsworth is joined by experts Ron Westfall, Chris Grundemann, and Jeremy Schulman as they discuss how to better implement these preferred methods. They also debate how each model has different requirements and may face headwinds in an enterprise.
Transcript
We don't live in an on-premises world any longer. Things live in the cloud. We're excited to move all of our applications and workloads there.
But should we be moving things from the cloud back, specifically the operational model of the way we use our networks? In this episode of the Tech Field Day podcast, we postulate that on-premises networks need to work like cloud networks. Welcome to the Tech Field Day podcast, where we bring together a group of IT technical experts to discuss a single idea about key concepts in the industry.
This podcast features a variety of perspectives from members of the Tech Field Day delegate community, and is often associated with one of our events such as Networking Field. Day 35 Tech Field Day is a part of the future and group, and this podcast is also published by our sister company Techstrong on this episode recorded before Networking Field Day. I'd like to take a moment to introduce our delegates slash guests before we jump into the premise for this episode, starting with Ron.
Thank you, Tom. Yes, Ron Westfall, research director at the Futurum Group, and I head up our networking coverage and definitely looking forward to our conversation. Hi, my name is Jeremy Schulman.
I go by Network Auto Maniac on Twitter slash x, and I lead a organization at Major League Baseball that's transforming the network infrastructure team on the space or the caveman is Spaceman journey. Hi, yeah, Chris Grundman here. com.
I'm a network engineer by trade and now do lots of things in automation. Alright, well thank you all for joining us today. Let's jump into the premise for this podcast before we move forward.
We all know what an on-premises network looks like. We've been operating them for decades and we all know what the cloud is because we've been dealing with that for at least a decade at this point. And we know that, uh, there are cloud networks and we know that they're a little bit different than on-premises networks, but should they be in this episode of the Tech Field Day podcast, we would like to explore the premise that on-premises networks should work like cloud networks.
I know I just tossed a big one out there and I can already hear people type in like they're ready to to disagree with me, but I wanna open the floor up to some of the experts here. What is it about cloud networks that differ from a traditional on-premises network? Well, There's certainly a couple of things.
Um, and just like most things cloud, what we're essentially talking about is, is outsourcing, right? And so just like cloud compute is somebody else's computer, in large part cloud networking is somebody else's network. And what's interesting, I think about like modern enterprise networking is now we're stitching together these various networks that are run by someone else under the hood, but we still need control of them.
And so I think, you know, if you boil down far enough, there is an on-premises underneath that cloud somewhere. And so I think that leads back to our premise, which is that if, uh, the cloud providers can run an on-premise network in this way, why can't the rest of us? Yes, and I think, uh, we see in increasing choice here, uh, for example, there's colocation implementations to help ease enterprise comfort with using a a cloud type capability that's not strictly on their physical on-premise implementation.
And also, uh, I think we're seeing the cloud providers themselves, themselves saying, Hey, we can deploy our servers and other capabilities at the customer's premise to basically combine, you know, the best of on-premise with cloud. And so I think we're gonna see more of this. I think, uh, one another way of understanding it is what's called the hybrid cloud model, where we're seeing, uh, organizations warming up to we wanna keep certain capabilities, certain applications strictly on premise because of legal requirements, uh, comfort zone and so forth.
But also we wanna use the cloud where it makes the most sense, you know, for, you know, customer facing, uh, interactions and so forth. And so really I think what we're looking at is a blended world that's really the destiny that's a hybrid cloud world. And so, uh, on premises will increasingly emulate many of the best cloud practices.
And also I think the cloud providers and also cloud implementations will look to make it more comfortable so that it does parallel or emulate some of the best on-premises practices that have been around for decades. As you pointed out, Tom, What, what I'm, what I'm finding quite interesting is when people think about cloud networking and they operate cloud networks, that that experience of operating cloud networks is still very different than the way we operate on-premise networks. And how do we bridge that gap?
Because if we are wanting to treat our on-premises networks in, in a cloud oriented operational model, I think there's a lot of work to be done there to, to figure that out. I I think that there is an opportunity here, but I don't think we're there yet. So I actually, I wanna make a a a statement here because this is something I was thinking about as I, I watched an old video that I made talking about the difference between operating a cloud network and an on-premises network.
And it comes down to more than just mechanics. Like my favorite thing is is go ahead and log into your Amazon dashboard and click the, the button for this VM that says vMotion. And you'll find that there is no such button because there is no concept of layer two in the cloud.
Everything operates at a layer three level. But I wouldn't necessarily operate my data center in that way because of the way that certain applications have certain kinds of interactions or things like that. But at the same time, I would never try to port applications that need things like vMotion or stretch layer 2D CI in the cloud is part of our reaction to the way that these networks operate.
The fact that we're being forced to shift away from traditional development and deployment models like virtualization, hypervisors to a more cloud centric model such as Kubernetes and things like that. To a degree, I think that's, uh, right Tom, and I think you're definitely in the ballpark of what's going on here. To some degree, it's this limitation of choice is actually part of what allows cloud networks to operate the way that they can, right?
And so that example of, you know, we're running on layer three also, maybe we're only buying a certain type of hardware, we're only deploying a certain network operating system. We're only allowing certain features. This limitation of options allows the network to be standardized at a level that allows it to be, you know, more of that kind of pets versus cattle model, right?
Where, um, turning up a new switch isn't a big deal because we're not trying to figure out what all the weird stuff that we're doing in our network needs to be translated into this switch. We know that the network operates this way, period, right? And so I think, uh, in order to get a on-premise or kind of a traditional physical network that you may be operating for your organization more like a cloud network, I think standardization is a huge part of that.
Um, now standardization may look slightly different for each organization. Uh, there's certain things you can get away with telling people no on and certain things you're gonna have to say yes to, but I think standardization is a huge part of this. And then once you've standardized and you've got this kind of network policy standard of like, here's what we provide, the network is now a service to the rest of the business and here's how you consume it, then it gets really easy to do some automation and some different things that the cloud providers do to provide that interface back to developers, back to apps, back to executives that they can consume that physical network like a cloud.
I think I, I think that there's a kind of a bit of a challenge there though. I mean, one, for one thing, I think when, when you operate, you know, uh, a multi-cloud environment, there is no standard, right? I mean, all, all the cloud providers, you know, have their own ways of managing their own cloud infrastructure, much like, you know, our ne our on-premise network vendors have their own ways of managing the bits and the bobs, right?
And I, I always find that the argument of, well, if, if we only had a standard that we could all adhere to, wouldn't it be great? And, and a lot of times that just doesn't quite happen the same way. So what I'm looking at is, is what are the, the, the fundamental principles that operating a cloud network can be applied to operating an on-premise network?
You know, can we constrain the user experience much the same way the cloud providers constrain, you know, the user experience in a way that allows for some kind of basis to do a self-service model or a resource allocation model or something that allows them to have the same operational experience that they tend to enjoy. But there's no free lunch with that. I mean, you know, when the network changes, you know, do we version control our networks?
Like, you know, if you look at how, um, enterprises might deploy networks in the cloud, they have perhaps Terraform files. Those terraform files are version controlled, so they can literally version control their infrastructure, you know, on a, on a, you know, on a day basis, on a release basis. How can we achieve the same kind of experience with on-premise equipment?
I would, I would postulate we can, I mean it's, I I believe it's possible to do this. It's not, but we're not doing this yet. You know, and, and does that kind of along the vector of our, of our, of our discussion, like how can we do a thing with the on-premise network if we had this capability, would we be closer to that operational model?
I think that's a fair point about standardization. I don't think the market's gonna wait around for, you know, a magic set of standards that can make everything akay. I think we've learned from history, yes, standards are important, uh, they can level set, you know, certain implementations and so forth.
But I think the market is really driving, for example, the multi-cloud capabilities that are gonna be fundamental because any large organization out there, let alone business is going to use, you know, more than one cloud provider. It's just, you know, diversity 1 0 1 in terms of having choice, in terms of having flexibility with the cloud journey. And each organization is gonna have its own unique cloud journey in terms of how do we optimize the workloads?
Uh, are we gonna, you know, keep this workload on premises because of, again, mandated, uh, legal reasons or, you know, can we actually transfer it to the cloud to gain the benefits of cloud economics or, you know, cloud flexibility because now there's a sovereign cloud that's being offered. And what I mean by the marketplace is, you know, helping drive this. We're seeing cloud providers like Oracle Cloud and Google Cloud getting together to provide interconnects that logically enable the customer to use, say an Oracle database natively through a Google Cloud interface.
They don't have to use, you know, two separate silos to be able to make that work. And we're also, you know, seeing the same thing with Azure and, uh, Oracle Cloud. So this is moving forward without standards being in place.
And I think, uh, a lot of this is just gonna be up to the organization to decide how to pace their cloud journey, how do they want to optimize their workloads? And we're not gonna be, you know, waiting for, you know, multi-cloud standards necessarily to make this happen. Yeah, I think that's a really good point.
And and to be clear, while I think that the debate around global standards may still be raging, I tend to agree with both of you that that's not what we should be waiting for. When I say standardization, I mean in inter organization, um, or intra in intra organization, right? So, so inside of your organization, build some standards, right?
We are going to, we are going to provide these services and this is how we're gonna provide those services, right? And, and we are gonna deploy these types of devices and they have to have these capabilities. So building your own standards, your own policy framework inside so that you're building your network to a standard that can then be modularized and automated the way that the cloud providers do.
So then you can provide services instead of networks, uh, to your organization, right? And I think that's a big piece of this is this service oriented architecture, this idea that no one actually cares if it's a VLAN or a VXLAN or an MPLS tunnel or a GRE tunnel at at, at the business level. They just don't care.
And so what you need to define is what are you actually providing to the business? And then if you do that in a really standard way, you can actually change things out underneath, and no one even knows if you switch from GRE to IPSec or from MPLS to vxlan. In a lot of cases, you can do that without anyone knowing, as long as you haven't defined that in the product you're providing to the organization.
Yeah, I think the ch the challenge there though, I mean, it sounds great and I, and I would love to live in that world because I'm trying to create that world, you know, by, you know, going across teams to create standards or plan documents or whatever you wanna call it, to say, let's collaborate together to agree on a thing, you know, in advance of doing that thing. And the challenge has always been, if you're in the network engineering team, you might be fighting fires every day and you don't have the luxury of time to sit down and, and plan. Now we all would assert that that is necessary to do, like, we need to sit down and plan something to a, to the nth degree so it goes smoothly.
But what you have is, you know, your arch nemesis, captain chaos and his cohort general disarray running around going, well, I know I told you this thing, you know, last week, and that's what was the plan, but you know, we just found out that vendor X, you know, slipped their schedule and now we can't have that product and you know, blah blah burp. So, you know, there is, there is this rational element of like, if you're going to try to create standards and if you're gonna try to operationalize your environment much like you would try to do it with the cloud, there's planning involved, you know, there's, there's a deep set of planning that's involved that I don't know, that every organization has that kind of cultural intuition to, uh, to execute on. I think a lot of it comes down to the fact that organizations also accrue a lot of technical debt over the years that they're forced to deal with.
And I, I hinted about it when you come to things like vMotion, but I mean, we saw that when we transitioned a lot of these things into the cloud in the first place, specialized hardware that must run in a certain way on a, on a device or networking protocols that are critical to the operation of the environment that probably should have never been configured in the first place. I'm looking right at you, GLBP or you know, ISSU or a bunch of other things that we've come up with as solutions to make a more resilient, responsive network when what we should have done was push back against the app developers and go, no, that's not how we build things. How dare you?
Because to kind of, to Chris's point, standards aren't the solution to this problem. We have standards. The problem is we have too many of them, we have too much choice.
We have too many ways to, to skin this cat to, to use the common metaphor, what we need to do is take away choice. If there's only, if there's two ways to do something, you don't know which way someone's gonna pick. But if there's only one way exposed on the dashboard, I promise you I know which way you're gonna pick and I can plan appropriately.
Yeah. So should we get into a situation where as administrators, we are creating artificial boundaries that basically kind of set our people up for success by telling them there's really only one way to do this? I, I agree with what Chris said is that it has to be a service oriented point of view.
If you can agree on here is the service, here's what matters to me in the, as a business stakeholder, and I can, I can express that desire through, in the case of cloud operations through Terraform at the very low level or through some other mechanism, you know, for on-premise or maybe even the same, you know, I've seen people experiment with trying to use Terraform to provision, you know, the low level network, um, doing it from that point of view. I wouldn't use the word constrained, but really, um, curate that experience so that it's designed at that time to be what people want rather than saying, you know, here is the kitchen sink of all the various options. Go tell me what you want.
It's like, you know, I spoke with you, you told me what you wanted, here's what you wanted. I built what you wanted. Here's the, here's the, here's the execution model.
I think that that's what the cloud providers in a way have done very well with their execution model. They said, you know, here are the things that you need from cloud networking. Here are the ways that we're gonna provide it to you.
The cloud networking, you know, the different providers have a different way of ex exposing those widgets, but at the same time they're like, here's how the app, the app people can take advantage of the network and they gave them a way to do it. I think probably when we're talking about on-premise networks, is it possible that the network engineers just have a hard time saying no? Yes.
And I think that, uh, pivots with, uh, increase prioritization of building cloud native applications. Uh, that's something that I think is a clear trend. As we know many applications that have been on premises for years, even decades, you just don't port them directly to the cloud.
They, in many cases have to be, you know, re-engineered in order to, you know, actually, uh, operate efficiently in a, uh, cloud implementation. And, uh, I think that, uh, parallels that excellent point is that, uh, the conversations I've been having over, you know, the last two quarters at least is it's about the workload optimization. And I know I'm gonna risk slings and arrows from Tom, and a lot of it's about, uh, AI workload optimization.
And that quite simply, I think is an excellent example of here is a megatrend technology that is basically compelling organizations to simply get the best of both worlds as best as they can. That is okay, what needs to be on premise or not in the cloud. And that is, you know, on device AI inferencing is certainly something that immediately comes to mind, but it has to also be orchestrated and optimized with AI training in the cloud.
You know, the heavy lifting training that goes on with the AI clusters, you know, running, you know, hundreds, thousands of GPUs and so forth. So I think this is something that is going to continue to drive the conversation, you know, why on premises, why cloud, how can we combine them IE hybrid cloud in an optimal way? And I think we're also seeing, you know, key players like HPE with GreenLake, uh, Lenovo with true scale, uh, Dell with AI factory and, uh, apex, you know, basically literally, uh, positioning themselves to, you know, make this happen.
And I think that's going to be a, a big difference maker in the competitive landscape. So Ron, let me respond to that because you've illustrated the point that at least I, I'm trying to dry that perfectly because AI is a disruptor. Like I cannot go cobble together a network running on a hundred meg ethernet and say, I would like to run on-premises AI inferencing with a bunch of old Nvidia voodoo cards, 3D FX video cards.
It just doesn't work, right? Because you're sitting there going, no, you must deploy this hardware, you must deploy these workloads. No, you cannot run this data center off of a 10 K-B-A-U-P-S anymore.
If you want to do this, you're gonna have to bump it up. And disruptive technologies like AI or any other kind of thing that have very stringent requirements fix this problem because either you're going to run it the way that they tell you to run it, or you're gonna have to go buy that compute time or whatever from someone who has that capability if, like, if it's in the cloud like that. To Jeremy's point, like networking people don't like saying no.
So if we show up with a pile of token ring cards, someone's gonna say, well, I think we can try to make this work. Because they don't want to try to tell the people, no, no, I'm sorry, you're gonna have to actually invest in a, an ultra ethernet network. Or you're gonna have to put InfiniBand down here.
Like, like the, to me, that shift of there is a technology that people want to utilize, which means we have to orient what we're doing around that technology that is the biggest driver for operating these networks in that manner. Because if you try to operate an AI inferencing network like you would operate a traditional on-premises network, you're gonna waste millions of dollars doing it because, you know, little things like the, uh, the, the tail end latency on a workload, well, you know, in a, in a traditional network, I don't care when the packets show up, like if they're, uh, uh, 500 milliseconds late, it doesn't matter to me. Well, except when you're running, you know, 10,000 GPU cluster, if every, uh, packet is 500 milliseconds late, even that's an order of mag, like several orders of magnitude late, that's causing everything to rack up time.
'cause those machines are sitting there idling drawing power while they're waiting on everything to return. So you kind of almost have to, um, create more structure around them in order to get them to work properly. And I think choice is gonna matter.
I think that's a, a great point, Tom, yes, they're gonna be situations, scenarios where, okay, they're going to have to invest more in say, ultra ethernet in order to, you know, scale heavy lifting GPUs. However, I think there's also going to be choice. So that is right sizing the workloads that is using small language models that can use say, Intel and a and d CPU on a distributed basis so that you don't have to really do that, you know, expensive cutover in order to succeed, you know, in a gen AI world.
And I think that's just gonna be part of the spectrum of what's going to evolve, what's going to happen in terms of the decision making. I think something that helps with this potentially is, and again, going back to this idea of services, is, you know, taking that even a step further where, you know, your network infrastructure team, um, should be treated and should treat the rest of the business as if they were a service provider, right? I mean, a lot of these challenges we have with on-premises networking have been solved in service writer networks.
And the reason is because they define a set of products and they build to those products, and then when a customer needs something else, they define a new product, a new service, and they figure out a way to build to it. I, I think you're right. I think Ron and Tom both said, you know, the problem is, you know, network engineers who are, are afraid to or unwilling to say no, I think if you change the model a little bit and look at this as a service that's being consumed, I mean, we're looked at as a cost center in most cases anyway.
So like, let's take that all the way and look at this as a service. We're providing, uh, what is the service we wanna provide? Okay, you want to build out an ai, you know, data center, great, here's how, you know, here's the network we can provide for that use case and allow the team to focus on the other bits.
And then that also allows potentially the outsourcing of networks. I think that Jeremy's right, that like not every organization is staffed for this model, um, but maybe that's okay. Maybe you know, big enterprises that have the, the, the funding and the staffing to build out an internal service provider can do that.
And smaller organizations, maybe they should be looking to partner with, uh, outside partners who can come in and operate this network for them in a, you know, very cloud-like manner where they can consume it, consume these products, consume these services, uh, at, at the, you know, at at a lower price. It would cost 'em to build out an entire team of developers, system engineers and network engineers and network architects to build the thing the way that it really should be built. I, I think you just highlighted a really important point, which is, you know, if, if an organization wants to benefit the operational efficiencies and all the good things that come with cloud, you know, cloud operations for the network on the on-premise network, there's, there's a human cost.
Like, you know, the Googles and the Azures and the Oracles, they have teams of people, they have software engineers building web UIs and control systems and, you know, alerting and monitoring systems and just all this infrastructure to give the, the end customer that experience, that self-service experience that allows 'em to, you know, poof, stand up a network or poof, stand up a pod or whatever that they're trying to do, and to look at an organization and enterprise organization and say, well, we want the same thing, we want that same experience and not have, you know, an army of people. You know, creating such a thing I think is, is a misstep in ex in expectations. Like people, you know, you people will expect something, you know, they go to the networking team and they say, Hey, I want this thing.
And the networking team will say, well, when do you need it by? And they're like, I need it by tomorrow. And I'm like, do you really need it by tomorrow?
I mean, maybe, you know, what about next week? Oh yeah, next week would be fine too. You know, and there's, I think there's still that subtle, you know, there's gotta be some pushback and negotiation between organizations about setting expectations, because if you want on-premise networks to operate like cloud networks, I think there is a operational infrastructure cost that has to go into that.
And if an enterprise wants that, they're gonna have to pay for it one way or another. And I think that that maybe is where we need to get back to this, the realistic expectations of what we can deliver, whether it's an on-premises network or a cloud network. Once we reconcile the differences between those two things and we can actually deliver that to people, then they stop asking for those crazy, you know, corner case type things.
Um, you know, one of the biggest complaints that I can remember from the service oriented architecture back in the day was the reason why everybody wanted their network to operate like a cloud does, is because they got tired of asking the networking team to provision something and it taking three weeks to get it done. What they didn't realize is, is that when that capability is put in your hands, it really does take two weeks to make that work for whatever reason. And they're just impatient because they think that things should be as immediate as clicking on a website when it's actually not.
But that's what they get with cloud networking. They, they get this instant gratification and an assurance of a quality deliverable. Like that's, that is cloud, you know, that that's the cloud service offering.
That's what people get out of the cloud and everybody wants that for on-premise infrastructure. I mean, who wouldn't, right? But there co it comes at an enormous cost, right?
And I don't think people really stop and think about, well, what does it take to deliver that kind of experience? And do you have the resources to not only, you know, build that infrastructure but maintain that infrastructure? It's like people who are on, you know, the automation journey.
Once you write a script and you put that script into production, well guess what? You gotta support that for, you know, how long it could be 10 years, you know, that you're supporting that thing. And, and that is a, a cost, it's a human cost, it's an engineering cost.
Um, so these benefits that we look for, they have costs. And people should try to think about those costs before they jump to, gosh, I want a thing. It's like a five-year-old wanting to eat ice cream for breakfast.
You know, it's like, you know. No, That's a good point, Jeremy. I think that's, uh, also putting a spotlight on the economic factors.
Yes, the technical challenges are there, there are many innovations that are helping it along. You know, for example, bringing the AI to the data, like we're seeing, uh, with the Oracle heat wave and AI implementation, I, that allows customers not have to go to an external third party vector database in order to get those capabilities. And I think it's also underlying the fact that when it comes to cloud economics, people have, uh, the conventional notion that, oh, this is automatically gonna create savings, not necessarily gonna do that in terms of the actual opex and co CapEx.
So why would people wanna do, you know, more cloud implementations? It's because of these other factors. Improving the user experience, improving the workforce experience, also, uh, being able, uh, to get time to value and other factors that can actually improve overall business, uh, outcomes.
And I think that's something that will continue to, you know, drive the decision making. How do we make the best of both worlds on-premise and the cloud? I don't think it's an either or proposition.
Yeah. I, I once had this conversation with, uh, a presenter who was kind of giving this time horizon view of the, the death of the on-premise data center, like, 'cause everything was going to the cloud. And I said, you know, after he presented, I, I said, you know, and at the time I was working at a startup company that was building software to build data centers.
And so I was mildly concerned by his prognostications and I said, um, you know, do you really believe you know that it's all gonna go to the cloud? He said, no, absolutely not. There's always gonna be, you know, a minimum viable data center, something that has to be on-prem for whatever reason, it's connecting to bespoke hardware that needs, you know, you can't put it in the cloud or there's legal reasons like we discussed.
And even managing just a small data center, you know, in these, in these times is as complicated as managing a big data center. I mean, you have the complications of the protocols and the interactions of, of all the, the bits and bobs. So there's, there's still a cost to that, right?
So even if, if people are looking at, well, what can I put into the cloud and what can I keep on-prem? There's still the complexity of running an on-premise network. Well, as you can tell, the answer to this question is a little bit more complicated because there's a lot of things that come along with just simply changing the operational model of what we do.
We have to examine what our workloads look like, how our developers operate, uh, what our outcomes are supposed to be and, and what we're dragging along with us along the way. And until you have answers to those questions, you shouldn't even be considering the operational model of your network. Until you get all the facts, you probably should keep doing things the way that you do them right now, but it doesn't mean you shouldn't look to the future because one day very soon, you may have that opportunity to turn a key and make everything more cloud-like, and that should make your day a little bit brighter.
Thank you for listening to this episode of the Tech Field Day podcast. If you enjoyed this discussion, please subscribe on our YouTube or your favorite podcast application so you don't miss any of the episodes. And make sure you leave us a rating and a review that helps everyone in the community understand a little bit more about who we are here.
This podcast was brought to you by Tech Field Day, home of IT experts from across the enterprise and part of the Future Room Group. com/podcast. Thank you very much for listening, and we'll see you next week.