Intro to Managed Lustre with Google Cloud
Dan Eawaz, Senior Product Manager at Google Cloud, introduced Managed Lustre with Google Cloud, a fully managed parallel file system built on DDN Exascaler. The aim is to solve the demanding requirements of data preparation, model training, and inference in AI workloads. Managed Lustre provides high throughput to keep GPUs and TPUs fully utilized and enables quick writing and reading for checkpoints.
Currently, many customers leverage parallel file systems (PFSs) like Lustre on-prem. Google Cloud Managed Lustre makes it easier for customers to bring their workloads to the cloud without re-architecting. It optimizes TCO by maximizing the utilization of expensive GPUs and TPUs. The offering is a persistent service deployed co-located with compute for optimal latency, scaling from 18 terabytes to petabyte scale, with sub-millisecond latency and an initial throughput of one terabyte per second.
The service is managed, where customers specify their region, capacity, and throughput needs. Google then deploys the capacity in the background, providing a mount point for easy integration with GCE or GKE. The Google Cloud Managed Luster service has a 99.9% availability SLA in a single zone and is fully POSIX compliant. The service integrates with GKE via a CSI driver and supports Slurm through the cluster toolkit. It also has an integration built for data batch transfer to and from Google Cloud Storage.
Presented by Dan Eawaz, Senior Product Manager, Google Cloud Managed Lustre, Google Cloud. Recorded live in Santa Clara, California, on April 22, 2025, as part of AI Infrastructure Field Day. Watch the entire presentation at https://techfieldday.com/appearance/google-cloud-presents-at-ai-infrastructure-field-day-2/ or https://techfieldday.com/event/aiifd2/ for more information.
Transcript
Product manager on Google Cloud, I lead our manage luster offering. So this is something we announced just a few weeks ago at Google. Next, happy to review this and kind of give you the rundown of how this helps accelerate AI l workloads.
So as my colleagues have presented today, the A IML pipeline is unique in the fact that it, it is very, very demanding. Across the board. You've got your data preparation model, training, inference, and then delivery.
And within each item you have different requirements. Uh, for the data preparation, you've got this potentially massive petabyte scale requirement model training and checkpoints and validation. You've got this high throughput requirement.
And then for inference, you need this low latency. And with our portfolio, we've been developing kind of the solution to help customers through each and every of one of those checkpoints. So specifically here, what we're, what we're introducing is the lust our offering.
And I'll, and I'll go a little bit into what that is in a, in, in a bit. But on a high level, what we're trying to do here is to solve for kind of the main points of areas we're, we're focusing on is this model training, thank you checkpoint checkpointing, and then potentially also inferencing. And with that, the model training luster will provide this high throughput to help kind of keep those GPUs TPUs fully utilized.
And as far as checkpoints, the ability to, to write and read quickly that these are kind of the sweet spots of what, what we're positioning luster for. So as we move to the general category of parallel file systems, uh, as part of the play in our total portfolio, right, the currently on-prem, many, many customers leverage PFSs that's short for parallel file system. And in particular, many of them use Luster.
So this, this is an ideal play for us. What we're trying to do here is make it easy for customers to bring their workloads from the on-prem to the cloud and not have to rearchitect any of their workload. So this is the kind of the value add of, Hey, let's leverage managed Luster.
It's been out in avail for the past 15 plus 20 years, and essentially many, many customers are leveraging this today. So what does A PFS do for the overall workload? Essentially, as I mentioned it, it helps kind of drive these really demanding workloads that are requiring this high io, high throughput ability to read many millions potentially of small files and saturate those VMs to be able to utilize their infra infrastructure to the maximum.
And within this kind of, uh, overarching, uh, proposition is the ability to optimize your TCO, right? The, at the end of the day, as more and more customers are coming from on-prem to leverage Google's ai, hyper computer, they either for burst into the cloud or migrating workloads from the on-prem. Uh, one of the key items you're going to constantly be kind of keeping tabs on is the cost.
And one of the highest drivers of cost is the GPUs and tpu. So anything you can do to help keep that in check is really, is really key. Um, and it's not to mention the fact that at the end of the day, you want to reach your output as quickly as possible and, and the more in the most performant way as possible.
So here I I have a chart basically showing what this means. Example, right? You've left side, you've got a baseline storage offering.
Um, and what that means is essentially through these epochs, it's taking, it's longer than the right side, which is a more performant offering. And why is that? Essentially the compute is waiting longer and it's not being maximized.
And then, so as you, as you optimize each run and you're able to maximize the, the compute for each overall, you're gonna bring your total time down. And so essentially that, that goes back to bringing your overall TCO as a whole lower, and, and that, that's really key. And we see a lot of customers, uh, kind of making sure and, and calling this out, right?
'cause at the end of the day, you, you are trying to, uh, be within certain budgets. So happy to, and officially announced here as, uh, for the third or fourth times, because two weeks, we did it officially at next actually. But, uh, Google Cloud manage Luster, what is it?
Right? It's fully managed parallel file system built on DDN exascale. And I'll, I'll go into who did DN is in a bit, but essentially what we wanted to do is bring this powerful, highly leveraged, uh, file system that's been well known for many years on-prem and offer it as a fully managed offering on Google Cloud.
And historically, while luster is really, really powerful, it is not the easiest to manage and having it offered through a managed service on Google. We basically abstract that from the customer. Kind of the big high level value ads is what are we offering here?
So first of all, it's a persistence offer. Um, and it's going to be deployed and co-located with the compute so that this is to optimize kind of that, that performance and make sure you're getting the best latency possible. Secondarily, uh, as far as a service, it, it's scaling currently from 18 terabytes all the way up to scale.
And we're offering this sub millisecond latency and an overall, initially with an overall throughput of one terabyte per second of overall read throughput. And going back to kind of the partnership, right? Why, why we, we kind of decided to go in to partnership with DDN, as I mentioned, this is a fully managed first party offering.
So essentially what we did is we went and partnered with DDN who gave us their, their version of, uh, a version of exoscale that we then took and applied to run on our GCP Infra Infra. And what that means is we get that tried and true proven, um, offering that they've been, that drives kind of the largest, some of the largest, um, clusters in the world as far as on-prem. And we're bringing that expertise and product over to the cloud and, and allowing our customers to take benefit of that.
And within that partnership, right, we, we get not only the, the the instance or software itself, but we also get the know-how and, and the knowledge base of the DDN resources that have been doing this for several years. And that's another, another key value add for our customers. So diving deeper into kind of what, what the offering actually entails, um, as I mentioned, the key value prop is it's managed.
So what what does that mean as a customer, you tell me what region you need it, how much capacity you're looking for, and your overall expected throughput. And, and essentially the throughput scales to the capacity at one gigabyte per second per tip. So essentially what, once you give those kind of inputs, what we'll do is we'll deploy that capacity for you in the background, um, and you literally go into co console, you spin up, create the instance, and it'll give you your mound point.
And it's simple as that. And then as far as utilizing that mound point, um, you, we are compatible with GCE or GKE, and I'll speak to a little bit about that in a second. But essentially, um, luster itself is, is a really, it's really, uh, performant, but that does, that doesn't come magically.
There's a ton of infra pushing that performance in the backend, but that's all abstracted from the customer. And I'll have somebody from Didi and join me shortly to walk over kind of the technical details of our, the architecture and kind of a little bit, give more into luster itself and what's that kind of secret sauce and magic behind it. So Dan Yeah, Jack Paul, our paradigm Technica, I actually, um, wanted to go back on, on your previous slide.
One of the things you had on the previous slide was, uh, uh, you mentioned about zones and zone, I can't remember the exact Yes. You said, right. How does that play with the, um, zone availability that we were discussing the last session?
The, the availability else everywhere schema? Good question. Yes.
And it, it, it is a very different but kind of value prop and it doesn't actually play into it. It's because it is tied specifically to that zone and kind of the, the, I think what you're referencing kind of was the anywhere cash where Yeah, thank you. Yes.
That's, that's specific for, for object storage. Okay. Where this is, no, this is a purpose bill kind of deployed, um, storage offering in that particular zone and cluster.
Mm-hmm. So if that answers your question. Yeah.
So it does. So if you want to take advantage of the luster, you have to, you have to tie yourself to that specific zone where the luster is, and you can't avail yourself of the performance benefits of the anywhere cash. Correct?
Correct. It's two different, I mean, there are two different storage offerings, but you're right. Yes.
E essentially, um, what's gonna end up happening is you're gonna, you're gonna deploy this instance or create this instance in that one zone to take advantage of those GPUs or tus, right? That you're getting. Mm-hmm.
Um, if you do by chance have other GPUs or tuss in another zone, yes, you're right. You would need to create a different instance for that. Gotcha.
Uh, guy Coer rum, are you also guiding customers as to model type, um, uh, you know, AI type and uh, and size and so forth for this? 'cause? Um, I would imagine it's not simply linear, it's not just the biggest models in every type.
It's No, it, great question. Essentially what we do and how we're trying to address this, and it is a good point because models are all all over the place, right? So what we do is, as, as you first come into and, and express interest in leveraging luster, we'll have these kind of discussions with you, understand your workload, and then appropriately to that, we will tune kind of your instance to best accommodate that.
So it's, it's a very well known file system within the h HPC world mm-hmm. Where there are high precision Yes. Sometimes very atomized data sets, like all those sorts of things that lends itself to lend themselves to high parallelism and correct.
And, and, you know, very low availability. That is not true across ai, especially across the AI life cycle. So you, you are working with the solution.
DDS has doing this for years. You're working with the solution, correct? Correct.
Customer when it comes to ai, Correct. We're, we're, we will tune it to the best that we can for that specific use case based on the, the feedback from the customer. But to your point, right?
It's, it's hard to be a one, it's not a hammer, right? You have to kind of, there is finesse in that. So kind of let, let me walk through this.
Um, persistent offering is mentioned. 9. So in a single zone, and, and that goes to kind of your question and we do offer this fully, it's another value add, right?
It's versus our object stores, Hey, we're fully s compliant. So, and, and that goes, that speaks to, again, customers with their on-prem workloads, it's very easy for them now to move that here. So from, if your 99, your three nines yes.
Availability, if I wanted to get to a higher nine with you, what would we, I need to do? Um, honestly, right now, there isn't a line of sight to go past that just because the nature of the infra that requires to run the, a high performing offering like this. There, the fact is there is x number of, uh, nodes to power this type of performance and hardware fails, right?
So we are providing this, this SLA but that to accommodate that. And, and that's why the kind of a line of sight past that is unlikely, but it doesn't mean it's not something we'll look towards in, in the future. Yeah.
One quick comment on that, Kimberly is like this, this is a persistent offering. So you can have your data there and have confidence is there for three nines of availability, but if the zone goes out, then obviously you have to think about other things, which is one reason why when you doing checkpoints, excuse me, doing checkpoints, you can periodically write those back to cloud storage. And we have bulk APIs to extract whatever data you want for subset of training for different workloads.
But then also if you have to do a checkpoint restore, it can be from cloud storage because that has 11 nines of durability and has at least a regional offering, if not multi-regional offering. Yeah. And I would make sure that gets brought up Yeah.
The kind of conversation about it, because No, That's a great, While the HPC doesn't like necessarily need high availability, they, they haven't built for that. 'cause it's costly for all the reasons that you said. Yep.
You also need, you know, that having the customers knowing that that's what they have, the architecture of what they would do is kind of gives them some reassurance there. And that is a great point. As to Sean's point, I'll, I'll, I'll speak to this and I'll go back to, to our GKA implementation, but we, because of this, we've built in a data batch basically transfer API within our, we handle kind of, we handle this transfer of data to and from GCS.
So knowing very well that most of the data does live in GCS, A lot of these customers, they leverage GCS, right? They put their petabytes plus scale on there and, and kind of our, our, our proposed kind of usage is, hey, you, you know, you wanna, you're gonna run these models or you're gonna run these workflows, leverage this, it's internally built, it's gonna manage it for you. You don't have to worry about it.
It's gonna, and it's gonna kind of batch transfer that from luster, bring it, I mean from GCS, bring it to Luster. And then when you're done with, with whatever you're gonna do, to Sean's point, what he mentioned is if you're writing checkpoints, right, you first to optimize for performance, you write it to Luster, but eventually you go spit that back to GCS, that doesn't need to live in luster, especially as historicals, you don't, that doesn't need to be retained in luster. Um, so that's another really strong kind of integration that we've built from the GetGo with our luster offering.
The second one I wanna mention is the GK. So not only we, we also, right off the bat, we fully integrated with GKA. There is a CSI.
Yes. Questions on that. Just going back, um, so I had this question originally, and I was gonna wait and see if you addressed it, but, um, how, how do you manage consistency on your managed luster?
I mean, what, what is your consistency model? It's fully, I mean it's, I I believe for your, it's fully consistent as far as, are you talking about transferring back and forth to GCS? Well, That was gonna be my second question.
Okay. But my first question is that if you actually have a distributed, uh, parallel file system, uh, how consistent is the data? Is it, is it eventually consistent or is it instantly consistent?
It's eventually consistent. Okay. Okay.
Okay. And then what is, what is the consistency between VCS and managed luster? That would be that question.
I mean, it's, it, it will be also, it will follow that kind of same eventual consistency. But the fact is, the fact is, is this data transfer job that's running, it's gonna monitor that and it's gonna ensure that consistency. If not, it's gonna, it's gonna point that out that, that that particular data transfer had issues and you, you would be, you would see visibility into that.
But the, the agent, basically what this consists of is, as you see here, we spin up kind of these agents in the background and it, and it's, and it's doing that transfer for you. And, and we're monitoring that and we're, we're, we're passing kind of that end state to you with that. You're, you're gonna see any errors if there were any.
Okay. And so the, the follow on question that has to fall out of this is, um, it's how do you handle, um, simultaneous or multiple, uh, consumers accessing the data at the same time As far as on luster itself? On luster or between luster and GCS?
Well, do you have, do you have a locking mechanism or We, So event, how, how initial state is, it's very clean. Um, you're using, there isn't this kind of back and forth. Essentially you're, if GCS is your system of record, that's where you're gonna go.
And that's, so there isn't this kind of, uh, conflict resolution between the two. If that answers your question. We, we can, we're, we're kind of almost, we're running into time limitations, so we can have follow up questions after.
I did wanna cover one last point before handing it up to Yvonne, who will go over a little bit deeper into the technical, um, architecture and kind of luster itself, technicalities. The last point I wanted to make is this GK integration. We do have this unmanaged driver initially with a managed driver coming up very soon.
So right off the bat, we do have this integration with GEA to run your workloads if you're leveraging that. And we are also integrated with the cluster toolkit. And so we do support s lm So I'm gonna go ahead now, and as I mentioned, hand this over to Yvonne.
Yeah. Just one quick comment before Yvonne goes in. So the question about the bulk transfer and which is consistent and so forth, like that bulk transfer is a operation that you do.
Once that's complete, then all the clients are hitting managed luster as a PFS solution. And so that's very different. There's no in, in-between state that there would, there would be inconsistencies.
It's an either or. Once it's complete, you progress with managed luster. If you wanna write checkpoints back, you can do that as well.
And you'll know when those checkpoints are successfully written back. But it's not a combination of the two. It's very different than anywhere.
Cache cloud, storage, fuse and object storage. This is, think of this as a second option for primary data for AI workloads. Okay, well, I, and I mean, the second part of my question went to the multi-tenancy aspect of it.
Can there be multiple tenants actually accessing the same chunk of data? I'll yvo, Yvonne, you wanna talk about that, but like At some point Yes. I don't think I can fully speak about it.
So luster itself, yes, it does support multi-tenancy, uh, for a specific case on Google. I, I think, Uh, within Google it's all single tenant. Yeah.
Right now it's a single tenant. If They, if you have multiple projects or want to have multiple individuals and different groups within the company, yes. But it's all within an individual company that spins up and manage luster instance.
It's not a multi-tenant service that Google is hosting five different companies within a single instance. Okay. Yeah.
So it it's up to the consumer to come up with a, um, contention mechanism. Yep. If, if, if, if needed.
I, I don't, I, I feel like this is addressing, uh, a sort of a zone of of use cases that, um, where this level of, of high consistency is not. There's a lot of batch, there's a lot of, not batch really, but there's a lot of jobs running as opposed to a lot of users. Yeah.
Agreed. Accessing and, and a lot of, Alright. So, um, there was a, uh, hi, my name is Yvonne Pat Dubi.
I'm director of engineering at DDN. Um, and my team is partially behind this service and great offering. So glad to be here.
Uh, I'm gonna talk about the, uh, um, there was a great question about, last is known at the, uh, HPC world for 25 ish years or so, uh, whatever 90 start, start 1998 or so, give or take. Um, but great question about known in HP in HPC world. What are the advantages in ai?
So I'm gonna focus a little bit on this. And the first advantage is, uh, it's the actual parallel file system design that brings the biggest advantage. So, um, many modern, uh, object storage file system offer, you know, access to the objects, but they focus mainly on the, um, you know, synchronous, um, kinda, uh, workload.
But the, uh, parallel health system allows the simultaneous access by hundreds of, or thousands of clients, uh, in a random, uh, fashion, uh, access to the multiple oss, multiple OSTs, and therefore, um, random access to the file, random workloads. And it brings the, um, minimizes the, um, GPU down, uh, downtime. So the, the last thing you want in the AI workload is your GPU sit in idle.
So the parallel file system design in its nature was the access to the, uh, by the multiple clients, by the, uh, um, you, you want the data to flow, constantly flow. So that's, that's probably one of the biggest advantages that we have today in it. Um, so high throughput and low latency io combination of these two things gives you the, one of the cutting edge advantages there.
6 gigabyte a second. This is 700, uh, percent efficiency post storage node there. So we don't obviously have it today on Google, but we are thriving to go there.
Um, this is the use case on-prem, but at some point, Google will will use the very similar technology and will constantly working on accelerating and improving it. So this high throughput and low latency is one of the key things that Luster brings, uh, to the Google infrastructure. Google, uh, Google world, uh, and will continue working on that.
Uh, the Exoscale Exoscale platform itself is linearly scalable, and the luster is linearly scalable. So you add more nodes, you add more capacity, you add more systems, you get better results, uh, which is crucial for training models, especially if you have billions of parameters, uh, and, um, multimodal data sets. So, um, very complex to manage if you're doing it on your own.
But, uh, cases like GCP where you have the managed offering, um, that headache is taken away from you. Um, so other key aspect is the, uh, separation of the, uh, uh, data and metadata services. Uh, one of the things is, um, um, big in the AI world, world, uh, if you have text files, you train in your model on the, uh, on the text or small images, uh, the small IO is a big, big problem.
If you can separate it between, uh, data and metadata, you're eliminating this bottleneck. You are minimizing this headache for you, uh, object storage, um, or non-parallel file systems out there. They don't have where that separation is not there.
You will run into some problems, you will run into some bottlenecks. Yes, there are pro ways to solve it, but the luster with its, uh, current architecture solves this problem for you today. And, um, uh, it eliminates this, uh, eliminates the, the, the bottleneck.
So faster metal metadata lookup, um, you look up the file, you prefetch the file name, you prefetch the permissions, you prefetch those things, um, with certain techniques, um, you'll get a better result right there initially. Big thing in AI world. Next bullet point, checkpoint, uh, luster provides the way to do asynchronous checkpointing, um, which again boils down to better utilization of the GPUs.
Uh, the, um, current luster checkpoints run about 15 times faster, um, uh, and, um, simplifies your recovery. It's, um, uh, and again, back to the parallel file system design itself. Um, when you are on a large training models, when you need to write terabytes of data or checkpoints, you can write it more efficiently.
Uh, object storage. Um, yes, Azure coordinating is great, everything's great, but think of it like you need to type, uh, when you need to write, um, possibly two or three times more data in the checkpoint. So imagine when you writing a terabytes of data in the checkpoint, um, you don't have time, you, you, you spending cycles doing it.
Um, but if you can do it with the last hour today, um, and write it faster and more efficiently, that's, that's saving your, saving your time, providing the faster time to do the, uh, to achieve your results better. Lopez Info Advisors. Sure.
15 times faster than what Exactly. Comparing to the other, um, let's say our competitors. I don't other architecture Things, but your competitors who are also doing Checkpoint, uh, everyone's doing checkpointing.
Yeah. Um, so I, yeah. So other competitors, other, um, object storage file systems or object storage solutions, I would say not necessarily file systems.
Yeah. Uh, big thing is the, uh, support for GPU direct storage. If you bypass the CPU memory, um, and you write directly to the GPU memory, it's always faster.
It's always better. So Luster does support zero copy RDMA for full, and it has support, full support for read and write again versus some other, um, uh, solutions out there. So it reduces your data overhead, transfer overhead.
So saves you, saves you time, saves you money. Uh, in the last but not least, um, back to the metadata intense workloads, the distributed namespace support and the, uh, perfection of the, uh, file attributes. Um, again, it does save you, uh, time to, to do the, especially for the small io back to that, um, progressive file layout and file level redundancy.
Uh, you can do certain performance optimization techniques. Uh, this requires some of the knowledge, uh, of the last health system. But with the, uh, certain techniques, you will eliminate the hotspots in the read heavy AI workload and, uh, you will get some benefits out of that.
Um, and posix compliance, uh, you can migrate your existing workload today, now without any headaches and run it today. You don't need to change your application, you don't need to rewrite your application.