80. Private Cloud is not just Self-Service Virtualization – Tech Field Day Podcast
Private cloud is not just virtualization 4.0, self-service VM deployment doesn’t fulfil the same need as the Public Cloud. This episode of the Tech Field Day podcast features Mike Graff, Jon Hildebrand, and Alastair Cooke. Private cloud has evolved from simple virtualization to a more comprehensive, cloud-like experience, emphasizing the need for on-premises infrastructure to offer the same developer-friendly tools and APIs as public clouds. Some application repatriation is driven by cost concerns and enabled by rise of technologies like Kubernetes and OpenShift for managing containerized workloads. A unified control plane for hybrid cloud environments is vital, as is accurate cost accounting for on-premises resources. Enterprises will search for a hybrid approach where developers can deploy applications without needing to worry about the underlying infrastructure.
Transcript
Private cloud's in the news private cloud is amazing, but private cloud is not just virtualization. Find out what the differences and why you might need to manage your costs to make sure that everything makes sense and you don't end up with a $40 million bill for a new AI data center. Join us on this Week's Tech Field Day podcast.
Welcome to the Tech Field Day podcast, where we bring together a group of it technical experts to discuss a single idea about key topics in the industry. This podcast features a variety of perspectives from members of the Tech Field Day delegate community, and is often recorded in association with one of our events tech field as part of the Futurum Group. And this podcast is also published on our sister company site, Textron tv.
On this episode, as we head into our next Cloud Fields event, we'll be discussing how private cloud is not just virtualization four oh before the discussion. Let's meet who's on the panel today. Mike.
Hi everybody. I'm Mike Graff. I'm, uh, currently infrastructure architecture director over at Dolby Laboratories based in San Francisco.
Um, my role is around cloud operations, cloud governance, as well as just kind of general infrastructure architecture for my organization. I'm glad to be here. All right.
I guess I'll go next. I'm John Hildebrand. Uh, you can call me an independent contractor at the moment, doing a variety of different things for different companies.
0 maybe, uh, depending upon where we're at right now. And I'm Alistair Cook, an event lead here at Tech Fields. I, with a long background of teaching, uh, VMware training courses, and then AWS training courses and all things data center infrastructure and compute infrastructure.
And as we're heading into cloud field day, we noticed that there's quite a lot of on-premises infrastructure in our cloud field day. And so my focus for Cloud Field Day has always been on the enterprise reality of hybrid multi-cloud. And so some of that cloud is on-premises, but it's not just virtualization on-premises.
It's not good, good enough to just say, I can deliver you a virtual machine and there's an API to deliver a virtual machine. That's not what I consider to be a real cloud. Um, for me on my paradigm for thinking about clouds comes from the AWS trainer background, but cloud is about enabling a developer to build and deploy applications rapidly and to make changes to those applications rapidly.
And the ability to, to then do that on different platforms, be there on premises or on public cloud, is really where I see that hybrid multi-cloud story turning up. Um, Mike, you had some thoughts around what hybrid cloud really means, what, how it's useful and how it's different from just having virtualization? 0, right, as a, as a premise.
I mean, for, for me, when I think of, you know, kind of the legacy approach where we were using something like a traditional hypervisor like VMware via on premises infrastructure, it was not really managed like a cloud, right? The cloud paradigm of, you know, rest APIs and, and being able to call those to manage your infrastructure, you know, some of the tr legacy vendors have adopted that. But I think what we're gonna see this week from oxide and some of the other vendors is more of that like, let's manage it like it's a real cloud and, and give the same kind of experience to the builders that they, they're getting from the cloud providers.
And we saw some sort of parallels to this as, as one of the ways that we see private cloud happening is cloud providers actually extending their reach down to on-premises with, and you mentioned, uh, AWS outposts and, uh, Azure, uh, hybrid cloud, HCI, whatever they're currently branding it. And of course Google's anthos. So the major cloud providers definitely want to have their, uh, solution their tin in one form or another on your, in your data center and using the same APIs to deploy, but it's not necessarily what we're seeing customers deploying.
John, have you had some experience with people deploying things beyond virtualization as an on-premises cloud? Yeah, um, I mean, we're getting to the point right now that applications exist outside of the VM construct. So, you know, containerization is definitely starting to become very popular.
I mean, if you think for the most part, many applications that were born in the cloud and due to repatriation, they're moving them back into a data center at this point, uh, like Netflix as an example, um, they basically have been moving that application, but developing it still as if it were in that public cloud mantra, if you wanna call it that. So yeah, we're seeing application shift all over the place, mostly for dollar sign reasons. I think the pandemic put us in a, the open wallets of the pandemic are no longer open anymore for the cloud bills, and we have to have a serious look at where the money's going as far as, uh, where the resources are being put for these applications that are key to businesses.
Yeah, I think people really wanna manage, like I said earlier, they really wanna manage these on-premises estates the same way they've been managed with cloud. No one wants to go back to the old way of doing it. And your point about Kubernetes or serverless is, is is really spot on.
I think John. Like people are wanting to move away from these monolithic environments. They wanna be using containers, they want to be able to manage those containers in the same way that they've been managing 'em in their clouds.
I wonder, like the repatriation thing is interesting to me because I always feel like, um, you hear a lot about it, but I wonder how much it's actually happening beyond the people who have those massive, like seven figure cloud bills. Um, I don't know what your experience has been. Well, I could basically state that I've worked with some service providers that have done quite a bit of repatriation for customers that are out there.
Mostly because, well, let's be honest, Broadcom has rocked data center infrastructure for the most part, uh, with whatever you want to call it. Um, but a, the, the point is, is that folks don't associate virtualization just with VMware anymore, and many other stacks exist out there to be able to provide you that developer like feel for application development and application deployment. Shoot, even for the most part, uh, if you think of something like Red Hat OpenShift as an example, you can do both virtualization and containerization and put the workloads next to each other because as we know, any cloud-based workload latency is the killer.
So the the smaller, the smaller the distance, the quicker the application's gonna respond. And don't even get me going on. Uh, I, I know Steven will kill me for this, but mainframes in large enterprises, they still exist in data centers.
And if you want to access that historical data that is in those things and put a new application spin on that latency is going to be the killer. So to lessen the latency you put, you put these things in your data center. From this point on, I absolutely, we're seeing much more discussion in particular of Red Hat OpenShift virtualization, not just Red Hat OpenShift as a Kubernetes distribution.
And, uh, I think that is absolutely being driven by people's desire to have an alternative to renewing their, their VMware licenses if they're not using the entire VCF suite, which of course is Broadcom's plan is to only care about large organizations using the entire VCF suite. I said from the beginning, uh, one of the parallels I'll see in this is, um, some of what we're seeing from both Morpheus and oxide, and they both are providing a cloud automation layer on top of something for Morpheus. It's on top of the KVM distribution that they started showing us.
In fact, at our cloud field day at, uh, uh, uh, uh, a few months ago or prior to the acquisition by HPA, I actually thought that Morpheus and oxide together would be awesome because then you would have this layer that makes your hybrid cloud from both on-premises and, uh, public cloud homogenous. And I thought that was something that would be great for oxide. Turns out it's gonna be great for HPE and HPE would wanna, uh, give you that, that homogenous experience across your cloud and, uh, hybrid cloud environment.
So that's very much what I'm expecting to hear from, uh, our friend Brad Parks as he comes in and, uh, presents that cloud vision from HPE. And I think that resonates with companies to be able to use those same tools to be able to choose where this application belongs, whether the cost that we see as we're putting applications into the public cloud, when you've got that far more variable cost, uh, and per gigabyte second of, uh, compute resource, it's, it's gonna be a higher cost than the cost per gigabyte second of compute resource in your own data center. Just that you only pay for what you re you tend to only pay for what you use on the cloud versus, uh, on premises.
You are buying it upfront. So, and there's definitely a cost dynamic in there to work out where is the right place to put this application. Had customers going back at least 15, 20 years wanting to have that conversation as being able to have a single place where, uh, your internal resources can choose this particular application has these characteristics and this dataset 'cause that also there's some governance requirements and some data sovereignty.
Um, having a, a single location where my architects and application developers can just, I hear the criteria that I, I, I have for this application, put it where it belongs, not having to care whether that's on premises or on which of the public clouds. That vision is still not quite here for us Now. And I think, like you said, I think one of the keys moving forward, especially with the Morpheus acquisition and HPVM essentials that they're most likely going to be showing us, uh, during the HPE, uh, portion is going to be that commonality of the control plane, basically, where, where the developers put the rubber to the road, so to speak, and as long as they can get the developers to buy in.
Um, because at the end of the day, I mean, it's not exactly AWS it's not exactly GCP or Azure, but you, you get to control everything and still give that developer the experience that they're expecting for your business, essentially because what the developer works on, business outcomes come from it. So giving them the keys to the kingdom, um, it, uh, it still satisfies the needs, but you've gotta have that control plane and without it, you're not going anywhere with, that's why you've seen so many different stacks fail. Although one could argue, um, open stack has kind of zombie ffi and brought the, brought itself back from back from the ground, uh, with these Broadcom discussions.
But you're, you're, you're seeing this diversification, which meaning as long as the developers can access what they need, that's all that matters. Yeah, and I think this brokering thing that you're talking about at Alistair, this notion of like, you know, almost making, making the developer not necessarily have to worry about where the workload is running. Like it could go in the on-premises estate, it could go in the cloud.
Um, having a broker to make those decisions, you know, either via human intervention or via automation is I think where we want to be headed. I know that's something that we're, things that we're working on in my day job around that. Right.
And I think the missing piece for the on-premises side is you, you talked about cost, right? It's very clear and easy to see what I'm gonna pay in the cloud, um, to run a workload up there. Engineers, and at least in my company, tend to view that on-premises stuff as free, right?
'cause it's already bought and paid for, right? And there needs to be a little more education and kind of, um, changing that mind shift around this notion that it's free when it's on-prem. Because I do think being able to identify what those costs are and kind of doing, being able to compare on a workload basis, what it's gonna cost to run it on-prem versus in the cloud is, uh, something that we need to, we need to get to.
Yeah, absolutely. The, there's a whole collection of the infrastructure pieces and infrastructure governance that, uh, developers shouldn't need to care about. As John says, their their job is to write features that improve the business, uh, within the application.
Uh, they shouldn't need to care that maybe there's some governance constraints, maybe there's some latency challenges to making their application work correctly. If it's, um, on premises mainframe data being extracted out to cloud developers tend to not want to know about that, uh, unless they're forced to have a, uh, deep experience of a, uh, developer team who were taught about bandwidth, but were never taught about latency. And so they optimized their application to run on minimum bandwidth.
Uh, but they did that by sending thousands of tiny requests, uh, which meant that latency absolutely killed the performance of their application. So the, the more we can make these things, uh, part of our platform and maybe platform engineering is a, a topic we might wanna kick into a little, uh, make these decisions, either automated or as you say, might sometimes human decisions still still play a pretty big part in this, but wrapping around governance and performance designs and all of those elements that developers don't need to particularly think about or shouldn't need to think about, but are absolutely vital for the success of the businesses. This, uh, minimizing risk essentially, and a lot of governances around minimizing risk.
And so these, these things should be codified and they should be managed by policy, and they should be automatically applied if was leaping back to virtualization and minus six, wherever we're at with that, Ian, as, as the virtualization number, uh, would be leaking back to it the days where that was not even thought of. And as you're saying, Mike, the, uh, internal cloud has zero, zero incremental cost for the next workload until you get to the $50,000 virtual machine that requires another virtualization host, uh, the hundred thousand dollars virtual machine, the $40 million virtual machine when you need to build another data center. And having some accounting for those real costs, uh, is one of the things that we've often not seen in on-premises.
Well, I, I don't want to turn this into an AI discussion, but AI is, is is forcing the conversation yes. About costs in a data center, and it's doing it to a lot of the components that you said would take, you know, developers would take for granted because it was free. Powers not necessarily free.
0 days all, all over again where we've gotta worry about these things and they have to be factored into the cost of the overall application development cycle at the same time. So score one for AI for bringing that back to the forefront yet again. But even, even if you may not believe in the long-term viability of AI inside of a data center, it's sparking some conversations based off of what a corporation would necessarily own within those four walls.
And I think AI is also very much the trigger for those $40 million new data center discussions, because when you look at the power delivery and the cooling delivery in older generation data centers, we we're seeing an environment where you can only put two, uh, AI hosts with massive GPUs in a rack, because that's all the power you can deliver into that rack. And that's all of the cooling you can deliver. Uh, you have the cost, the incremental cost for the new data center to be able to accommodate dozens, hundreds, thousands of, uh, GPUs is gonna be massive.
And, uh, I think, I don't recall if it was before the, the recording, but uh, neo clouds came up as, uh, that that's a business opportunity for people to deliver you a GP as a service or AI as a service. That is, again, your own tendency on somebody else's cloud, and again, brings more of that hybridity. And hopefully we're gonna get the same kind of, uh, APIs and tool sets for pushing applications out onto neo clouds that we're using to choose to push them to existing clouds or to hybrid or our on on-premises clouds.
Uh, once again, things keep changing so fast, it's hard to, for us to get to that nirvana state where everything just works, We'll never get there. It's always gonna be inching our way closer asymptotically, but we'll probably never actually get there. Um, I think your point about the data center, again, it kind of ties back to the repatriation conversation as well, right?
It's like, for most companies, does it make sense to build data centers to run these things for, you know, unless you're a very large LLM creator, right? Does it make sense for you to be building this in your own environment? Or even though it may cost you more on a monthly basis, you, you eliminate a lot of the risk.
You, you're going with a actual cloud provider versus building this in your own data center? True. But I do argue, well, um, let's take the EU as an example.
You know, they're big on sovereignty data. Data has to stay in a particular location. So unless you can privatize those links to those particular public instances like, uh, open AI and those particular components, you're not gonna be able to get to a lot of enterprises to specifically adopt those without having to put something on-prem.
Uh, or they're going to have to build their own MSP, which, well, let's call it what it is, it's their own private cloud at that point. So they're, it's six and half a dozen, uh, of, of another. 0.
And I think one of the, the things we do see is that there's a diversity of choices for large to medium organizations. Absolutely, Mike, some are building those on premises, some are, uh, stuffing their data centers full of liquid cooled racks of servers in order to be able to get that density. And, uh, now the, one of my local companies here, it was a fertilizer producer, and they create sulfuric acid or use sulfuric acid in, in that process, and that generated a lot of excess heat.
They generated their own power on site from their own waste energy. Um, and that has a parallel we saw at AI infrastructure fields that we saw neo clouds Having a couple of sort of points of difference to John's point, data sovereignty, sovereign clouds having an, a cloud that has only a presence within, let's say France to be compliant with the French requirement that you keep all French business data on French soil, uh, or that, uh, having more of a green view of being closer to that waste energy. Uh, we saw at AI infrastructure field that placing data centers close to where there's waste gas being burned off at a, uh, natural gas extraction.
Uh, there are a variety of different use cases for these, these, uh, neo clouds. But we also have seen tools for building a cloud-like infrastructure on top of your own hardware in your own data center. And, uh, Rafa comes to mind as, uh, being at the last AI infrastructure field.
They're showing us that automation for building up that multi-tenant infrastructure that's consumable as a service either inside your own enterprise data center or as the way the neo clouds are building these things up. Uh, we keep coming back to needing a unified way to access the different types of resources that are available. It's no longer just, we we're going to be a cloud first on one particular cloud, and all we have to know is that one cloud, cloud, cloud, cloud, um, now it's becoming much more, we can be in a hybrid and complex environment.
We're gonna put some things on, uh, Google because they're better for maybe, uh, the DeepMind, um, team has produced better AI tools for us. We're gonna put other things on Azure because well, Microsoft knows, uh, active directory, uh, intra ID far better than we do, right? We'll be spread across multiple places, and that's huge amounts of complexity to manage, but we're also still finding that there's a lot of use cases where on-prem makes the most sense.
Well, that was kind of the promise of Kubernetes, right? Was that you, you would, you would not have to worry about where it was running. You could use the same model wherever you were, wherever you were deployed.
Um, I think people are seeing as we get, grow more maturity that that's, that has its own layer of complexity and expense associated with operating that, right? Yeah. So much expense on the management side, but you know, it still, I go back to layer one.
Um, there's still a physical infrastructure that has to run all that stuff at the same time. So, you know, you're, you're, you're talking about making standardization, uh, across your vendor portfolios and things like that, that again, large enterprises and MSPs are gonna, are always having those discussions to be able to figure out what pound for pound, what they can get the best out of each dollar that they invest in those particular devices. I would also wanna just circle back to one of the things that we talked about earlier with the repatriation and, uh, whether it's, it's real and Mike's comment about whether just the very large organizations are doing that repatriation to on-premises, because one of the elements I definitely see is the bigger the bill, the more incentive there is to optimize that bill, right?
If, if you're spending $10,000 a month, uh, you, if you save 10%, that's a, that's a grand a month. That's nothing to be sneezed at. But if you're spending $2 million a month and you can save 10%, it moves at a much bigger needle.
So, uh, I do see some cases where people choose not to optimize because the cost of analyzing to optimize, whether you're running on premises or in the cloud, the, the cost of analyzing to optimize is greater than the possible return. You don't wanna get stuck in that situation where you're spending more money trying to optimize the system than you can save from the system. It's one of those, uh, interesting challenges is you change, go.
What I've seen, I'm pretty involved in the finops community, and what I've seen is, is it's moving, the discipline is kind of expanding beyond cloud. And now there's a lot of discussion of like, well, how do we do finops for that on-premises equipment? Or how do I do finops on my SaaS estate?
Right? There's this notion of like, we've had a lot of good success with building those concepts for the cloud use, and how do we take those same things and extend them to the on-prem infrastructure and make sure that we're using it optimally. Um, I think the AI and ML is an area where repatriation has more, um, traction just because of the frightening cost of running these things in a regular public cloud.
Yeah. Uh, we definitely see that transition from experimentation in the cloud to a maturity where you realize the ongoing cost, the month and month out cost for all of your use cases is going to start multiplying that, uh, that cloud because you're paying for everything you use. That's the joy of the cloud, right?
You, you pay only for what you use, but the, the terror of the cloud is that you pay for everything you use. But that gets back to my point about that gets back to my point about visibility of what it actually costs to run things on prem, right? Because I think, again, there's this assumption that it's, it's kind of sunk cost or it's, it's, once you pay for it, it's free.
But that doesn't take into account the cost of the power to cool the data center that is running in or the cost of the electricity to drive the servers or the cost of the facilities. People that are, you have to have to operate that who may or may not be great at running data centers, right? Um, so those are all things that they're very fuzzy and hard to quantify compared to a cloud bill where everything's spelled out at the individual line item.
So I think there just needs to be a lot more work in that area and recognizing what those actual costs are of the on-prem. Yeah. And that's a continuation of a discussion we had when Cloud was new, when people were, uh, you know, it organizations were trying to be the department of no and saying, you can't shift it to the cloud and saying that that pennies per hour that you're getting is not comparable to what we're spending millions of dollars per year on, on premises.
Well, I think it's probably time for us to wrap this up because, uh, all of us need to get ourselves ready to travel to Cloud Field Day next week. So thank you all for joining us today on the Tech Field Day, uh, podcast. But before we go, where can people connect with each of you and maybe carry this conversation on Mike?
Yeah, so of course I can, you can find me on LinkedIn, um, where I go by my, my given name, Michael Graph. Um, but, uh, you can find me there. I also have a blog that I run, um, called Cloudy Advice.
So if you wanna check me out there, you can, you can connect with me and love to carry on that conversation. Yeah, you can find me on LinkedIn, um, pretty active these days, especially on the, since the independent contractor portion of the, of the day job. And I've actually made the jump from X to blue sky.
com, uh, as far as the username is concerned. And of course, you can find both Mike and John's profiles on the Tech Field Day website, particularly if you look at the cloud field Day 24 event. I, of course, am Alistair Cook, the event lead for Cloud Field day 24, and you can find me on all kinds of social media and around the web, either as Alistair Cook or as Demi Tess nz.
Thank you so much for listening to this episode of the Tick Field Day podcast, and if you enjoyed this discussion, please subscribe on YouTube or your favorite podcast application so you don't miss an episode. Give us a rating. Nice review as well.
That always helps to get us in front of more people who might benefit from this conversation. com/podcast, us on text on tv. Thanks for listening, and we will see you next week.