Storage Intelligence with Google Cloud
Manjul Sahay, Group Product Manager at Google Cloud Storage, presented on Storage Intelligence with Google Cloud, focusing on helping customers, both enterprises and startups, manage their storage effectively for AI applications. These customers often face challenges in managing storage at scale for security, cost, and operational efficiency, particularly with small and new teams. A key problem is the exponential growth of storage due to the influx of new data and AI-generated content, making object storage management increasingly complex.
The presentation highlighted the difficulties in managing object storage, which often involves billions or trillions of objects spread across multiple buckets and projects. Many customers resort to building custom management tools, leading to significant expenditure on management, sometimes 13% to 24% of the total spend. Google Cloud aims to address this by providing storage intelligence to help customers manage storage, meet them in multi-cloud environments, and reduce or eliminate storage management challenges.
Google Cloud introduced Storage Intelligence to address storage management challenges, offering a unified platform with features like datasets, which provides a metadata repository in BigQuery. Also offered is the ability to identify public objects and batch operations to act on billions of objects, as well as features like bucket migration, all designed to streamline storage management. With a free trial offered, early adopters like Anthropic and Spotify are already experiencing benefits in managing and optimizing their storage infrastructure.
Presented by Manjul Sahay, Group Product Manager, Google Cloud Storage, Google Cloud. Recorded live in Santa Clara, California, on April 22, 2025, as part of AI Infrastructure Field Day. Watch the entire presentation at https://techfieldday.com/appearance/google-cloud-presents-at-ai-infrastructure-field-day-2/ or https://techfieldday.com/event/aiifd2/ for more information.
Transcript
My name is Manju Saha. I have the good opportunity to be in the same panel last year and talk about what we are doing in storage to help our customers who are building AI apps. So, coming back today with, uh, some more of exciting features and how we have helped some of the customers take advantage of these, uh, opportunities.
So when we talk to our AI customers, there are two kinds of customers. So there are the enterprises, which are building new AI apps, uh, a security camera company, which trying to, you know, make sure that AI can now do person detection and, you know, pet detection or maybe an automotive company, which is trying to do the new kind of a self-driving, um, application. So there's a degree of, you know, there's a set of, um, enterprise companies and of course there are new startups which are building new models.
So in some cases going into inferencing. So we see both kind of customers. We have had some great wins for Google Cloud in both categories.
What I will talk about is once you provide the best storage, networking, compute infrastructure for these customers, how do we help them manage from a storage point of view? In many cases of enterprise, even figuring out what data is available to get started with, to take advantage is a challenge. So thus, that's what I'll talk about, storage intelligence, which is the product for helping customers with managing storage at scale.
So all these different AI customers, both the enterprise one as well as the new age startups, all of them have a lot of storage and usually managing storage for security, for cost and for operational efficiency is a huge thing for them. Many of these have small teams, very new teams, which are trying to move very fast and make it happen, and that's where they have a little bit of storage management issues. So, last, let's talk about why is storage management a challenge for these customers first, which I'm sure all of you'll agree with, storage continues to grow.
This is one service where, you know, the bytes keep increasing. Of course, we don't complain about it. That's great for us.
The bytes continue to grow. Yeah, and with ai, there's a lot more of storage. One, there's a lot more of new data being brought in, like for manufacturing defect detection.
These are like high risk cameras taking pictures, you know, at a very fast frequency. And all that lands up in cloud storage or in the other case where the AI generated content is kind of getting created, images, videos, and of course text. So all this is just accelerating the storage usage of customers.
It's growing very fast. Now, the scale of object storage, object storage is where mostly customers by default store all this training data. You know, that generated content, the scale of object storage makes the management tough.
And what I talk about manage, you know, what's the scale? Let me pull it out and show it. Many customers will have maybe thousands, tens of thousands of VMs, millions of Kubernetes bots, but they will very easily have billions of objects in cloud storage.
Some cases tens of billions in vs in trillion object customers. And all these are usually spread across multiple buckets projects. And at least in case of AI companies, a very small set of developers or admins who are trying to manage all this and go very fast.
And AI customers are increasingly multi-cloud as well. You know, many of them are looking for the latest GPU or availability of G-P-T-P-U across these cloud. So they have some data, maybe in another clouds, a lot of data in Google Cloud.
How do you manage across them? And when you combine these two, many of these customers end up trying to stitch together homegrown management tools like scripts and, you know, make it work, put it in a database, how do you make it all happen? And then in our studies, we have found they end up spending 13 to 24% of spend just on management.
Like think of that number 13 to 24% extra just to manage it. So this is the context of storage management. This is the problem we want to solve.
We wanna make sure our AI customers can move very fast, whether they're enterprises or new age startups, meet them where they are in their multi-cloud environments, if that's how they are kind of positioned and help them make this storage management problem go away or make it very simple and, you know, shrink the problem for them. So this is the problem. Again, storage continues to grow.
Scale of object storage is extremely compared to other services in the cloud, and it's not easy to build custom management tools. So the product we have introduced to help customers And guy courier Yeah. Future in research, uh, this has, this has been a interesting topic of discussion lately all around uhhuh, like DevOps and it, although my colleague here, Mitch May Yeah.
Which you're welcome to do. Um, why, why custom management? I mean, I, I think a lot of custom built are custom built because it's fun.
Yeah. Not because it has a good business purpose. Um, it, you know, I'm not saying people are resistant to things like low code or whatever, just that, um, there's less of a le less emphasis than maybe there should be on off the shelf services or off the shelf right.
Packages. So, so why, why highlight that as, as one of the problem, um, number one and number two, uh, what's your perspective on, on that? Yeah.
Are people doing custom built more often than they should? As I'm sorry. Yeah.
Saying. Yeah. So look why that's a problem is, you know, it's very, you know, always fun to build it the first time, but then for the next three, five years when you have to manage it and maintain it, yeah, yeah.
It's not fun. Like I'm at a customer at cloud next, then You're stuck with it in a certain sense. Yes.
Yeah. I'm at a customer at Cloud next, which they're trying to put the metadata of all, I guess like 15, 20 billion objects in a Postgres database. And they're saying now we are struggling with, you know, scale or somebody has to do the version upgrade and it fails.
And, you know, this is a company which is trying to do something in the healthcare with AI and images and they have a lot of data. So yes, I think somehow, you know, many of the customers organizations, they have grown with cloud tools as it has matured and there's a lot of custom tool, but in many cases it doesn't work. Sometimes there's also perception that hey, it all takes, you know, a couple of weeks of an engineer.
So it's, you know, way more cheaper. But the, in the end, if you look at the three year life cycle cost of everything adds up. So, and it Doesn't take two or three weeks, two or three.
It doesn't take the two or three weeks, two or three months, I think. Yeah. Just one.
Yeah, definitely. But just 1, 1, 1 thing here Yeah. Is I think also a lot of times everybody feels like the problem they're addressing or what they're discovering is unique, right?
They don't feel, or they start to try and find out what's out there for it. And that's a difficult process. So then they're just like, oh, well, we'll just make it, I I, yes.
They go straight to we'll make it themselves. But John, Yeah. It's having spent a fair amount of time in the DevOps space, uh, it's that most of that scaffolding and custom built is not product defined in an organization.
Yeah. So it has no home. Yeah.
So it becomes basically legacy immediately and as people move and it just becomes a real cost burden that you just don't want to be in. My question though is Yeah, have you published the research on those numbers? Um, let me, that would be very, That would be very helpful.
Yes. 'cause I totally agree with it. Yes.
Yes. Okay. I'm pretty sure there are some industry publications on it.
I'm not sure if we have published or not. Okay. But, uh, since we put it on our cloud next deck as well, it should have A reference, a resource There.
That's a good idea. Yeah, we are spot on. I'm like, this is exactly, you know, one of the problems we kind of wanna solve for customers also, Um, just as, yeah.
Just as a quick follow up, um, because you, it is a large number. Yep. But I would also wanna understand when you say ops and management, is that like, if we look at it from a technology business management point of view, like TBM council, right?
Is this truly, um, is there an IT spend component? Is there a quote unquote cloud spend component? And then what's the associated or attached labor?
So those three legs of the stool. Got it. Um, understanding how to suss out where parameters are would be really helpful.
That's what the research will help. A good point. Hi, is this, hi, this is Mitch also.
Yeah. Future, another connection point is, you know, increasing adoption platform engineering, they're also starting to look at the storage parts of this as part of that scaffolding, right? Of maybe they're becoming an owner, if not data, data ops folks.
Well, that might be something to, if you aren't, haven't looked at that yet. Yeah, no, no, good point. I mean, many of, at least the AI customers, we have seen very small teams trying to move very fast.
Mm-hmm. And they also don't have the time or the intent skill to be an expert in just storage. Exactly.
Now, many of the enterprises, we have been around for a long time with a lot of, you know, large IT teams, they can afford to do it. Maybe they came from a VMware or some other era where there were these specializations, not with the new new edge companies. And in the new world.
And that's where it's even more important. Like, hey, we don't need someone who's a pure object storage expert. It's too difficult, it takes too long.
So the product which we introduced at, you know, cloud next, just about a couple of weeks ago, hardly two weeks ago, I guess now, is storage intelligence. And this is our unified single management product to manage cloud storage at scale. Now, I wanna emphasize unified, because the approach we have taken is very different from our two big name competitors.
We have not gone and launched point features many times. Somehow our competition will have point management features with some degree of overlap. And, you know, customers to figure out, should I use this, should I use that?
We said we'll have one single management product, of course, with a couple of features integrated to offer an end to end journey. We want customers, when you Say, okay, so I'm confused here, this is a storage management for your object offering. Yes, that's right.
Okay. So maybe I'm, I'm totally ignorant on your storage, your object storage capabilities. Don't you have one single management panel for your cloud storage?
We do. We absolutely have a single management panel or console for it. Okay.
This is more about beyond that, for example, and I'll, I'll show a little bit of and talk about it. Uh, like metadata is a huge thing in AI workflows. Like, let's take a simple example.
I'm doing something in healthcare with ai. I have a million of images and I just need to know this was model run number 174, or was it, you know, for this particular, let's say, you know, DNA sequencing all this metadata. Okay.
So what you're doing is adding a bunch of functionality on top. Sorry, I didn't introduce myself. I'm totally based withum.
Okay. Um, the, so what you're doing is adding a whole bunch of, um, data management capabilities within the cloud object, the, the, uh, yes. Cloud storage capability.
Absolutely. Okay, got it. Okay, Thank you.
So the base management, which is what we call like more of, you know, primitive spa management are available, part of the product free available all of them, including the management console. This is more about managing at scale, which generally is a different kind of a problem set for customers. So our sharing, like this is in our case, we have done a single management, uh, uh, unified management product offering, all these different features.
And the basic idea is analyze your storage rate at scale and take actions. So couple of features here. And the whole goal is eliminate manual storage management practices and free up the customer's team to do more business focused problems.
Build your AI apps faster. You know, get your in training and inferencing to run faster through this. Now let me briefly talk what some of these capabilities too.
Data sets is a very unique and powerful capability for customers. It takes all the metadata across your billions of objects. Puts that in our BigQuery, our data warehouse snapshot every 24 hours.
So if you're running, you know, model training runs and you wanna figure out after two months this particular model maybe had a bias, another issue, let me trace back which of the objects were involved in training. You could do this data set to enable that kind of a functionality. Or for many customers, if they just want to discover, oh, where all do I have, you know, uh, let's say images or objects with let's say cars, because I want to do a training related to that in a security camera context, you could do that.
So it offers you a huge repository of all your object metadata. And then it is set up across whole organization for a customer. These are some of the cloud constructs.
When customers get to do it, sometimes they struggle. Like the basic concept of organization and object storage is bucket objects in a bucket. If we enable it bucket by bucket, it doesn't work for most customers.
They have thousands, 10 thousands of buckets. We enable let them do it with a couple of clicks, right? In our level enable for the whole thing.
But this is dataset features. Another one, you know, we introduced, which was very interesting for customers is in cloud. And I think this, we had a very interesting discussion last time in the same panel or some of you like there are multiple ways to get an object to be a public, open to internet or private.
There are multiple, those are the seven things. What we have done is we will check against all those seven and come back with a definite, is it public? Not public?
And why? Because we keep hearing from customers data exposure, accidental is a huge concern for them. And in case of ai, this is the most valuable ip, the model, its in many of these things.
So it's even more interesting for, so again, very differentiated, unique capability available for customers very soon. And we just announced this feature And uh, yeah, obviously it's an attribute. Yes.
Then there is the, uh, sorry, JK through with Nexus tech. Uh, there's this attribute, but then there's the concept of like, what is the user experience of interacting with the attribute? So, uh, are you gonna cover what that might feel like or look like for someone that's trying to understand, oh my gosh, everything is, uh, is public.
Yeah. Versus this is a discrete subset of things that are public that I need to examine further. Sure.
Um, is there anything that you're doing in the metadata to determine what relative risk might be for the is public object, Um, at this point of time for us is a little bit difficult to understand the risk. But what it'll say is, Hey, we found object, which is named as let's say, um, training slash model run important slash you know, object one jpeg and it's public. Mm.
Hopefully with the name or metadata, that's the most common way customers look at it. They can look at it and say, okay, gosh, this should not have been public at all. And then we provide a policy called public access prevention, which is a single click and you turn it off.
So that's the way they will do it. Um, we do have some other advanced services in Google Cloud, like service data production, which can look within the content to understand the risk and highlight that. So yeah, that's another way to kind of look at it in this spot.
So Kimberly again, okay, so this is public is an object attribute. Yes. That's within the insights data sets that you've taken a snapshot of and put into BigQuery.
Yes, that's right. Okay. Yeah.
So our idea is all the object better data, put that in BigQuery, a data warehouse. Got it. And then along with the name and the location, which are, you know, common things customers look at, we will do this computation against seven checks and add this extra attribute so they can use it for data exposure management.
Okay, I'll keep moving. Other than new capability, again, within the storage intelligence, this is for the action which we introduce, it's called Batch Operation. You know, to your, to your question, Aliya, zero to no code act on billions of objects.
Again, from the outside in, look simple. But when you have to go and update a billion objects, it's not easy. Things fail.
Some things, you know, succeed, some things don't. How do you make sure that it is performant, automatic scaling, retries, all that we have taken care, updating encryption key, or you want to add an object? Metta, again, let's go back to that AI example of healthcare.
You know, a customer in healthcare trying to do AI training on DNA sequence and images, how do you put, you know, a metadata called these 500 million images were used for training run number 998. And this is what something you could do highly scalable in a a short period of time. Okay?
The final capability in the action is, is how do you move a whole bucket from X to Y region? Again, you will hear from my colleagues some great opportunities about using anywhere cache and these other features to make your data available multiple place. But sometimes for whatever the reason customer says, this data which is in US West doesn't make sense for me.
US East is where my TPUs, GPUs are. Or from a sustainability point of view, this is the best region I wanna go to. How do you move the whole bucket with just two commands while preserving the attributes?
So extremely interesting again for customers. So again, you know the theme you will see across all of these, these are a little bit down in the detailed weed capabilities, but that's what they struggle day two to day 10. You know, when they bring in the data and they start working with AI training and planning and all that, you need to do this maintenance here.
Sometimes things happen. You have to move data, you have to, you know, put policies, make sure it's not public. And that's where we have found customers.
If they use custom homegrown tools, it takes them lot of time struggle, search on Google, try this doesn't work, let's open up out case, find out all that. And we won't eliminate all that and provide these ready-made tools and for customers. And again, the whole thing is packaged into one single product which are available.
All of these, and they're integrated. For example, you can take data sets to produce a list of objects which match a certain criteria and then use batch operations to take action on them. So that's the integrated journey, very differentiated from our competition.
You will not find them integrated the way we have brought this to the market. Uh, hey again, I appreciate the arrow showing it going from where it's going to, uh, where it's to where it's going. Um, is there any kind of a policy as code where you could, um, present, uh, you know, a geopolitical overlay right this direction, but never this direction?
Right. Um, is that, is that something in terms of sovereignty that's also also baked into this? Yes.
So, uh, that's like more of our base management capability mm-hmm. Is called custom or policy where you could say, um, in my, in my account, in my customer scenario, I'm only allowing you to use US regions. You cannot go anywhere else with your data.
So that's the way you would implement it. Those are base capabilities available and yes, customers can implement to make that Yes, you're right, they're Automated. Right.
You're saying that's what you mean by policy, right? Correct. Yeah.
I mean customers have to of course make the choice very. Yeah. Yeah.
But yes, exactly. It's not just a tag is what I'm saying. Yes, exactly.
Yeah. And this also reiterates I'm a, I'm, this is a shared responsibility discussion we're still having of, right, of course. Yeah.
And we also have some pretty advanced, you know, controls from a security compliance point of view. Uh, access transparency for example. If any of the Google folks are using or accessing data for whatever reason, giving some of those details, those are available as well.
But he makes a good point. Um, if I'm using some automation to do this, then I'm gonna have sort of grander um, education and and policies. The question is, can you build policy into this solution itself?
Like in his example? 'cause you were saying is basically an IAM problem or solution. Um, Well it's something which is like an iamm built but for cloud storage.
So the usual model would be you put your policy first and then if you use this feature, Right, but it's not gonna be a user driving this type of process, it's probably gonna be an automation solution. And I probably need that. I don't want to have separate every automation solution for every sort policy.
I would like a policy built into this. And I guess I'm not getting a clear answer. Is there a policy, like to his point of his question, can I build policy into the this bucket relocation process?
Yeah. You build policy separately and first and then this tool or you do manual, all of them will, you know, adhere to the policy. So if somebody wants to do a policy, uh, do a transfer.
We we're Not connecting on the question. 'cause you're saying it goes back to the authentication policies. It's not the product itself.
Oh, It's built in. It's not an authentication policy. It's more like saying in my customer scenario, don't allow any data in Europe.
Let's say once you put the policy and even somebody wants to try to do this tool, it'll block say it doesn't mean the right, But some, I know that gets back to, and I know waste too much time here, but sure. Like there'll be lots of different policies. This group can move it here.
That group can't, this group can only move from here to there in a large enterprise. Yes. This kind of data can't be there.
Got it. And I guess that's where, I guess this is the last time I'll ask. Sure.
Can I build that kind of policy into Not I am. I was looking for, yes. Awesome.
So Jim KY neuro defective computing. Yeah. And another question, this is a little esoteric, but it happens uhhuh.
Um, let's say someone deletes a CSV file, some images, whatever that have been used to train or uh, do inferencing and a model because of GDPR, right? That data has to be removed from a model, right? For example.
Yes. Would this product as well as the last slide that you presented, would you be able to actually say, oh, we've got an issue specifically because someone has requested this, it must happen under GT PR. Right.
Uh, it's a great question. So before you jump into that module, maybe in the initial time jump to the demo Yes. For the, uh, showing something related to what you're asking about.
Excellent. But finding exactly what objects have metadata associated with them to do a particular action. And that was my next question.
Yeah. Excellent. So what you see is in fact the management console, you know, which you asked about previously.
This is the standard gold cloud storage console. And right here you have the storage intelligence, of course, step one, like I was saying, very easy to enable it. Couple of clicks enable it for the whole project or for the whole organization, was simple for administrators to do it.
And once we have done, you get this data and insights, data sets, again, one or two steps, what do you want to configure? And once the data is there, I'm showing you one of our looker dashboards where all the data flows in. So look at that, how many objects I have across how many objects, how is it, you know, broken down by different, uh, applications, storage classes.
And you could have custom metadata here as well. For example, training run, one training, trying to, like here we have put something like, okay, these are warehouse images. Product blueprint, again, depending on the customer scenario, appropriate thing will, you know, show up.
So look at that growth charts and then break down. And once you know that you could essentially kind of break, you know, zoom in into any of them, put any conditions, filters, and locate at what exactly you're looking for. I wanna find out only the product blueprints, which in certain category, in a certain region you could do that.
So this is the way it kind of flows all your metadata available in one place. And then after that, you know, you could do that, appropriate actions on that. So with that, I'll quickly go back last two slides and kind of finish it over here.
Uh, we have had some great, you know, success with some of our leading customers. Anthropic using storage intelligence to manage 85 billion plus of their objects and, uh, optimizing their storage infrastructure. Um, they've been a huge, uh, proponent and, you know, found success with this product.
Spotify one, again, of our largest customers using storage intelligence to again look at their data infrastructure, storage infrastructure and you know, make sure they make the right decisions. I'll skip some of this. Eight of the top 10 cloud storage customers are using storage intelligence, though it's a very new product.
We were in preview for a couple of months, but again, just to indicate this is, you know, how the value prop has resonated with customers and we just offered a free trial for our customers at cloud next for a great way to get them started. So I'll pause at that and very happy to take any more questions. I got a little bit rushed on the demo, but hopefully I communicate what I meant.
Yes. In the Spotify example, you had non-compliant. How, how do you determine non-compliant?
And then we follow on that question, can I look for objects that don't have metadata? Um, that's from an audit perspective, that's what I'm really interested in. Yeah, Let me answer your second one first.
I'll come to the first one. So every object has some standard metadata, which is like inbuilt, Right? But like, so no organizational metadata.
Yeah. Sounds like it's system of record data or it's, yeah, Correct. So that of course as a custom data metadata, customers have to feed in most of them when they are writing the object, that's itself the call, the API or the program will have these extra custom metadata, like whatever is your AI application will write.
But yes, even if you don't have this custom metta, the rest will come through the the But I, what I'm asking is, can I look for objects that don't have metadata that don't or don't have organizational metadata? Yes, you can. Okay.
Alright, good, good. Yeah. Alright.
Uh, sorry I missed, your first part was about Spotify. Oh, You had Spotify who had noncompliant and compliant. So I, Again, looking at their policies, for example, we should not have any data in this location or anything, which is for example, a financial customer, not Spotify.
So That was their determination? Yes, their Determination, Not storage intelligence. Okay.
Sorry, I add another question there maybe. Yeah, I see the number there. This says two $2 50 cents per million objects.
Yes. Okay. Is what the cost based cost is.
Is there additional cost for using BigQuery or is that all included here? Yes, uh, so we take care of loading the data, but if some customers are running BigQuery queries, they will pay for the slot charges. So that's come extra.
So there's a standard fee for the product, which is that two and a half dollars. And then depending on the feature that are one or two additional cost items, sum up pricing page. Okay.
But yes, bk, yes. Is there someplace to get documentation On this or? Yeah, it's, it's public.
So if you just search for cloud storage intelligence, I'm happy to share back, uh, you could find it. It's a public offering. So all the docs are there.
Yeah. So, uh, part is from Bitcom. I, uh, so you started with, it's all about storage management now you're adding that intelligence in there as well.
Is that only on the data on in the buckets of your cloud or is it, we live in a hybrid world, so yes, there are also other buckets, uh, where I can have data. Yeah. Or maybe even on-prem.
Uh, do you have integrations there with, uh, your own on-prem, uh, solutions or other? No. Right now it's only on Google cloud storage.
We don't extend into other clouds and all, but a little bit of intelligence I ran out of time. Uh, but it's there in the slide deck is what we have introduced at next is our vision here is where again, with customer approval and permission, we will look into the object content to detect some of these, uh, what is their insight and help customers then make automatic determinations. So that's a vision where we want to make the storage really smart.
For example, if an object has, let's say, uh, a car or a truck driver, we will make those determinations and help customers make those. So that's where we think more of the smartest or intelligence piece will come through. Yeah, because you were talking for instance about, uh, medical imaging.
Yes, I have a few customers, but, uh, in, in that sector. But they say, I'm not gonna put my medical images in a public cloud. It's gotta be in my data center uhhuh.
So maybe you should think as well for some solutions there then. Sure, Absolutely. I mean eventually, you know, as we ma mature it on this platform, yes.
The idea would be take it to our other platforms as well. For sure. Yeah.