60. Cloud Economics and Scaling Smarter in the Age of Data Abundance – Tech Field Day Podcast
Keeping every application and every scrap of data on the public cloud becomes very expensive; we need to improve our cloud economics. This episode of the Tech Field Day podcast features Vriti Magee, Mitch Lewis, and Alastair Cooke. The belief that data is the new oil has led many companies to retain every piece of data they generate, often in object storage on public cloud platforms. The continuous growth of this data leads to a growing bill from the cloud provider, often with no clear plan in place for recouping the value of the money spent. Generative AI requires training data, which is another reason to retain everything; again, there needs to be value returned to the business. New designs for cloud applications must include data management and managed retention as key criteria. Sustainable, honest designs that enable business change are vital for delivering value back to the business.
Transcript
Is the cloud costing you too much? Have you put too much in the cloud? Do you know what you've put in the cloud?
What data, what applications is it the right place? Our cloud economics and scaling smarter in the age of data abundance vital to your organization? Join the Tick Field Day podcast and we'll find out.
Welcome to the Tech Field Day podcast, where we bring together a group of IT technical experts to discuss a single idea about key concepts in the industry. This podcast features a variety of perspectives from members of the Tech Field Day delegate community, and it's often recorded in association world with one of our events. Tech Field Day is part of the Futurum group, and this podcast is also published on our sister companies website, tech Strong tv.
On this cloud field day episode, we'll be discussing cloud economics and scaling smarter in the age of data abundance. But before the discussion, let's meet who's on the panel today. Hi everyone.
I'm Ruti McGee, a technology consultant specializing in enterprise architecture, particularly with cloud AI and security intersect. Hi, I'm Mitch Lewis. Uh, I'm a analyst with Signal 65, which is, uh, part of the future arm group.
Um, we do all sorts of, uh, performance benchmarking, um, and, uh, third party product testing and validation. And I'm Alistair Cook, an event lead here at Tech Field Day. I'm the event lead for Cloud Field day.
This week we're talking about cloud economics, and I've always said that the wonderful thing about the cloud is you pay only for the resources that you use and that the terrible thing about the cloud is that you pay for all of the resources that you use. And so we want to think about scaling smarter in the, the age of data abundance where we've generated huge amounts of data, we've stored that data. We've often not really known the value of the data we've stored, but we've had to store it because unknown value is supposedly infinite value.
And yet, when are we gonna extract that value? Every day we keep this huge amounts of data on our, um, cloud storage. We have to spend more money on it, more and more and more money as we accumulate more and more data.
Where's the value coming back out of that? Mitch, are you seeing anything, uh, happening along this space about either making storing data more intelligent or extracting data intelligently? Yeah, I mean, I think, you know, you, you kind of hit the nail on the head of, um, you're gonna pay for what you use.
So, um, I think it's important to, uh, understand what you are using and why. Right? So I think it's still a challenge.
I think there's been, uh, handful of solutions, you know, trying to get at this in terms of, uh, data visibility. You of, you know, can you tear this off to, um, you know, S3 or, or something, uh, you know, glacier or something that's gonna be cheaper. Um, if you know what it is, can you actually delete it?
Uh, which is a, a concept no one wants to talk about. Um, and then other tools too that are, uh, you know, kind of more in the finops space that actually look at, um, trying to understand your spending. Um, but I think it's still a big challenge because people just wanna store everything.
Um, and there's this, you know, kind of concept of the, the data iceberg where, um, you know, you have all this data and you kind of know a very small amount of, you know, what it is that's actually useful. And then there's this much larger amount of data that's kind of under the water, um, and you're storing it. Uh, you don't know what it is, you don't know if it's even useful.
Uh, and then you get these planning cloud costs. Is your designing solutions that actually get deployed out on public cloud, is data volume an accumulation of data a a big consideration as you're designing these things? Absolutely, and, um, you know, in my work with strategy, it's quite interesting when you're talking to CIOs about cloud economics, it's, it's almost have has the cloud.
It's not that the clouds failed to deliver, but the value stories become harder and, and to defend when costs scale faster than the impact. And we're not in those early days of the lift and shift days anymore. And so the real pressure really is on optimization, right?
Sizing instances, managing those egress fees, consolidating storage tiers, automating policy-based controls. So it's no longer nice to have it's baseline expectations. That idea of cloud ops seems to be, be very much stronger than I, I certainly saw it probably three, four years ago.
The idea that, uh, that there is a, a corporate level requirement to look at how we are spending money on cloud, which aren't just pouring more and more money into the cloud, we're actually looking for, uh, getting value back to business from it, understanding where there's waste, um, where there is resources that isn't delivering the value we want for it. And also where there is insufficient value because we're not actually buying enough resources, and this is a double sided thing, we can always reduce our cost by using less, but then we maybe get less value, um, the idea and synopses to get that attachment between value and cost so that we are getting value for the money that we spend. That's right.
And I personally am an advocate for not just, you know, we're not just dealing with a technology change. It's, it's a strategic reset and, um, and a mindset shift. Uh, so not just kinda trying to juggle through the tooling that we've got and, and, you know, the obvious pressures of budgets and the fragmentation of data and, um, also major shifts in the vendor landscapes, like the VMwares, uh, transition from under Broadcom.
So that's had a ripple effect through the market, and interesting to see, uh, the broader decisions that are made around that. I think there's also, you know, opportunity for organizations to kind of stop and, and, and puzzle a little bit and think, you know, actually ask, does this belong in the cloud? Um, right.
I know, I know this is the, the cloud podcast, so maybe I shouldn't say that, but, um, you know, the, the cloud costs money as you use it. Um, and it's great for some things, and I think, you know, a while ago there's the idea that we're gonna put everything in the cloud, it's perfect for everything, and, um, and then it gets expensive, so you need to pump the brakes on it, right? Um, I think one example of this that I love, I love to bring up whenever I can, um, but is ai, right?
So, um, there's certainly some reasons to put AI in the cloud, and there's some reasons, you know, maybe not to, and cost is, uh, definitely one of them is it's gonna be, you know, one of your most expensive potential AI applications because of, you know, really, you know, big GPU compute instances, uh, potentially huge data volumes, things like that. So, um, I think beyond just, you know, sizing instances and, and knowing what your data is and, and where it is in the cloud, um, another way to kind of manage your cloud spending is thinking about, you know, should I even put it in the cloud in the first place? And one of the things I, I think is, is vital in our cloud field day events is the perspective of hybrid and multi-cloud.
As you say, not everything belongs in the public cloud. Some things really should be running on, on hardware that you own inside your data centers because it's more cost effective where you have a, a flat line load to place that somewhere where you, you don't have to, uh, pay for the possibility of near infinite scaling that you get on the cloud and you need enough scaling to accommodate your workload. You don't need, uh, if your workload isn't, isn't scaling by a factor of 10, well then you don't need a factor of 10 with headroom, and you don't need to be paying for it.
Uh, of course, scalability isn't the only reason for going up onto the public cloud. And I think, Mitch, you hit a really important one with ai. Uh, we saw in the AI infrastructure field day last month that there can be hugely expensive infrastructure to build foundation models and to do a lot of training, but when we get to maybe some fine tuning, maybe some retrieval, augmented generation, uh, this inference stage when we're actually getting the business value can often be run on a, a much smaller kind of device.
Uh, we saw a couple of solutions of using SSD as a tier for the, the memory, uh, in A GPU, a much more cost effective way of running a large language model that you probably aren't gonna see in the cloud for a while, but you could possibly deploy even out to your edge locations right now. So I think that there is a whole lot of other solutions for AI that are coming along, and I've thought for a while that just more and more resources, larger and larger data centers isn't a, an infinitely scalable solution to the massive scaling we're seeing with ai. They need to be other solutions that assist us with that scaling.
And I think one of the central, uh, elements of our, our premise here was about the massive growth of data and Mitch, uh, that transition from storing your data on your EC2 instances or your, your, your virtual machines using block storage towards using object storages is an obvious one, but even then, that object storage, you're gonna be paying for your, your capacity consumed every month forever until you get rid of that storage or those items on the storage. Uh, I think there's, there's some real challenges about knowing what to retain, what you're allowed to delete. Uh, and I think this is gonna be a really important part of Cloud finops in the coming months and years.
Some organizations are already there, uh, some are already facing huge costs for the S3 storage that they've just been bucketing away for years and years. I think there's some real challenges around data volumes. Yeah, I think absolutely.
And, and that's why you really need to know, you know, what the data is, um, you know, and how important it is, how likely are you to need it? Because, um, yes, there's S3 and then you get into, you know, your glaciers and your, your deep archive, um, things that can really help you, you know, on the economic side. Um, but if you do need that data, it's going to cost you more to get it out.
So if you're constantly, you know, uh, paying those fees, uh, then you have a problem. So it's, I think it goes a lot into, you know, knowing what your data is, um, and then even can you get rid of it, uh, which is something no one wants to do. Uh, no one ever wants to delete anything.
Um, but I, I think, you know, it gets to a point where maybe, maybe you should in some cases, um, and it all goes back to, you know, what is data? Well, your, you know, data is, is nothing but kind of your ones and zeros. Um, but when you want it, uh, you know, data can kind of turn into information that can be valuable.
So, um, you kind of need to think about it that way. Then you can understand where can you move it, uh, where can you place it that's gonna be, uh, cost effective. Um, silos is what comes to my mind first when I think of the word data and the challenge for, um, unifying our fragmented data and, um, you know, we've built incredible cloud systems, but often we've just ended up with silos.
Everyone's get got their own little cloud, that they're cozy in different clouds, different stacks, limited visibility. And, um, you know, one way of looking at it is unified data doesn't mean everything in one bucket. It means consistency in the metadata, access control and governance, regardless of where the data lives.
And I think that consistency in that governance needs to be set up real early. I think this is one of the, the challenges of the early adopters of cloud that have poured lots of data into, um, maybe into, into a, a data lake or a data warehouse and otherwise have, have poured data into an S3 bucket without that good metadata, good, uh, governance around it without really knowing what's there. And this is gonna be a sign of maturity for people as they're designing solutions on public cloud and equally on, on private.
I mean, it's just as difficult to delete things off your on-premises, uh, storage as it is to delete things off, off the cloud. Uh, getting that data governance in early and understanding that the value and the retention requirements for the data you're storing, I don't think that's a, an entirely new, uh, story. It's just that it, the pace of innovation, the diversity of types of data that we have retained in the cloud has led to challenges because we were very fast to retain it and not necessarily as fast to put the categorization and the governance around it.
Uh, it's, it's really something that needs to be done up the front rather than being retroactively headed afterwards. So maybe the best time to consider or reconsider how we treat data, not simply as a system output, but as a strategic asset in its own right, and find those data owners if you can. And then kind of the, the flip side to, to, you know, all of this in terms of like, uh, you know, tiering your data off or, or deleting your data, that's not useful.
I think the other thing that now everyone wants to do is, uh, not only hold onto everything just because they do, but, um, hold onto everything just because what if it's helpful for AI training? Um, right. That's, that's the other thing, um, that people are really thinking about now with all this data.
So, um, kind of just causes a, a bigger headache, I guess. Um, but no one wants to get rid of anything because it could be useful, uh, for some AI training, uh, some AI model down the road. Uh, so kind of how do you, how do you reconcile that also?
And the, um, one of the, the things that immediately comes to mind is that we'd often hold multiple duplicate copies. And this hits at some of Rich's point about the fact that we get siloed solutions, we get the right tool used to solve an individual problem, but not necessarily the right tool to, to provide good data governance and, uh, microservices architecture. Each microservices responsible for its own persistence, if it's requires any, that leads to that fragmentation of data, but also duplication of data, paying for multiple copies of the same data in, in multiple different places.
I think, again, there's a, um, a data management asset management element of, uh, designing your applications around the idea that we should only be storing this piece of data once. And when we start looking at taking live data and copying it into our data warehouse at, at our lake house, then we're inherently creating additional copies of data. Maybe that's gonna come back and bite us some more as well.
Maybe we need to be working with a single unified copy of our data, which is used for both reporting as well as transactional basis to minimize that long-term cost of retaining everything, or multiple copies of everything over time. Used to be easier when we used to just store numeric data, and we used to be able to do things like summarization of numeric data, but as soon as we started storing large amounts of text and then images and then video and all of the messages that we have in our various chat-based tools, we use Slack extensively within the Futureum group. Uh, those pieces of data have value, and they're much harder to summarize, much harder to condense these things into a smaller but usable unit.
And Mitch hit another point that's really important around the cost of retrieval. If you have stored all of this data without good governance, and it's sitting out on Glacier Deep Archive and you now want to retrospectively add governance, well, you need to actually retrieve that information to know what it's, and that's really expensive when you're using deep archive compared to the cost of leaving it where it's, so there's quite a lot of barriers to retrospectively adding this governance to, to your solution. You know, a lot of this conversation is, comes back to the need for, um, you know, data management, and that sounds really, you know, straightforward and obvious, but I think, um, a lot of times it gets kind of pushed off to the side as, um, you know, there, there's certain things that it, organizations are gonna say, you know, they must have, uh, you know, certain performance metrics, maybe certain, uh, you know, recovery objectives and data management kind of gets put into this bucket of, well, that would be nice to have.
Um, so it doesn't always get worked in. Uh, the other thing is I think data management is a really loose term, so it's hard to fully wrap your head around, you know, what does it mean? There's a thousand different companies out there that say they're doing data management.
Uh, some of them are, you know, more, uh, really more data protection, or some of them are more security, uh, some of them are more, uh, you know, govern governance there. There's all sorts of things, but I think there's a more encompassing, um, uh, superset maybe of, of data management that, you know, maybe it does do some of the data protection, but also does some of this, uh, you know, metadata ca cataloging, things like that, that help you, uh, understand what your data is. It does some of the archival or, uh, data lifecycle things, um, maybe some of the, the fin finops stuff.
Um, and, you know, I think with cloud, with, you know, hybrid cloud, multi-cloud, um, data management does become pretty important. Uh, and when you look at the, you know, the economics and the cost, um, I think it starts to move the data management piece from that kind of nice to have add-on feature to something that that really is a, an important kind of checklist, need to have ida, And I think that's probably where we are headed, right? Cloud economics as a capability, not just a cost center.
So I'm thinking, um, a way to sharpen focus, improve our planning and building infrastructure that's not just scalable, but it's sustainable and it's, and it's honest, and we're held accountable for those, um, big decisions that we make. We don't often hear the, the term honest attached to, um, to descriptions of architecture. And I think it's an interesting concept of the being clear about what it is we're doing, why we're doing it, and what we're expecting to happen as re as a result.
Uh, that's a, a really interesting, uh, characteristic to, to want to have in, in your infrastructure design and, and your data handling and your business and in general, uh, yeah, vastly undervalued concept that I really haven't, hadn't seen, hadn't thought of as, as being a vital part of designing infrastructure and designing, uh, solutions. Um, I think, you know, maybe, maybe one thing too that, um, can help with, you know, the, the cost issue is the data management issues. Um, I'm gonna sound like a broken record here, but, uh, ai, right?
So, um, if you can actually automate some of this away, um, with, you know, some, some intelligence that can actually look at your, your infrastructure, look at your, um, you know, data or data, data silos and, and kind of, uh, maybe help you figure out what's going on, um, you know, maybe that's a future path too, that can, that can actually, you know, help, um, create kind of like a intelligent data management solution. Yeah. Just one last thought from me.
Um, I think enterprises need to consider architecture that can pivot without rebuilding from scratch. So potentially that could mean favoring vendors with open ecosystems, walkability, and trusted transparent partnerships. Um, you know, trust and honest complex words.
Uh, we often put our KPIs first to progress our careers. Um, but yeah, I, I think that speaks to personal passions and, and, and, um, just getting the right people in the room often Understanding who was going to be the right person, who was the right stakeholder, who was, who was the right, uh, owner of data, owner of that, that, that long term. And I think your point about being able, being in an open ecosystem and being able to then use that as part of a, a pivot mechanism, understanding where you are choosing to become more locked to one particular solution and consciously making the decision that the value of being locked in that solution suits our business, suits the compromise of being, being locked, and other times choosing not to, to get locked.
Maybe ch choosing a higher cost to achieve an outcome, but to have flexibility in the future for changes. Uh, one of the things we do know is businesses will require change over time, and often that change will be unpredictable, that we will not know this year the changes that are gonna come in our business next year and require us to reevaluate our assessments. Uh, one of the things that we talk about quite a lot is as you're designing a solution to document not just the outcome of your decision, but the logic of your decision, you know, the inputs, why do we make the decision to choose this particular vendor?
Because if those inputs change, we might want to change the output, we might wanna change which vendor, but if we dunno which inputs led us to choose this vendor, we might not be able to move away. We might move away in a way that's inappropriate because we've forgotten that there was one crucial factor that meant we needed this particular vendor solution. And that, uh, having that knowledge that that, uh, source of truth about how we made decisions is often really important.
Yeah, I think just that whole idea of, you know, being intentional, um, is important both with, you know, selecting the vendor, but also, um, you know, we started this kind of talking about, you know, scaling, right? So are you scaling, um, really because you, you need to, it makes business sense, it makes economic sense. Are you scaling, um, because you're in the cloud and, and you can, uh, because you don't wanna delete any of your data.
Um, uh, so you know what, what you need to kind of like look at why you're scaling, um, and being intentional. So thank you all for joining us today on the tech field, they podcast. Before we go, where can people connect with you and continue the conversation ti Well, I'm on X and my handle is ti McGee, um, and LinkedIn.
And, um, I'm really excited to share my thoughts as we move into dialogue with our special vendors. com. Um, I'm also on, uh, Twitter, um, and on, uh, LinkedIn as well.
And of course, I'm Alice De Cook, and you can find me as Demi tes nz on a whole bunch of socials. com site, also on, uh, some of the tech strong sites and equally on the Tech field day do com website. So thanks for listening to this episode of the Tech Field Day podcast.
If you enjoyed the discussion, please subscribe on YouTube or your favorite podcast application so you don't miss an episode. And do consider giving us a rating in a nice review. This podcast was brought to you by Tech Field Day, a part of the Futurum Group.
com/podcast or view us on text on tv. Thanks for listening, and we will see you next week.