The latest in high-performance storage, Rapid on Colossus with Google Cloud
Michal Szymaniak, Principal Engineer at Google Cloud, presented on Rapid Storage, a new zonal storage product within the cloud storage portfolio, powered by Google’s foundational distributed file system, Colossus. The goal in designing Rapid Storage was to create a storage system that offers the low latency of block storage, the high throughput of parallel file systems, and the ease of use and scale of object storage, while also being easy to start and manage. Rapid Storage addresses customers’ limitations when working with GCS, where they often tier to persistent disk for lower latency or seek parallel file system functionality.
Rapid Storage is fronted by Cloud Storage Fuse, aiming to provide a parallel file system experience with insane throughput of 20 million requests per second and six terabytes per second throughput on the data side, while maintaining the benefits of cloud storage like 11 nines of durability within a zone and three nines of availability. The sub-millisecond latency for random reads and appends, achieved after opening a file using the Fuse interface, makes it significantly faster than other hyperscalers. A separate data stack enables this predictable performance focused on high performance, where GCS emphasizes manageability and scalability.
Rapid Storage, available as a new storage class in zonal locations, leverages Colossus in a particular cloud zone, bringing it to the forefront of cloud storage. By accessing cloud storage directly from the front end, with only one RPC hop away from data on physical hard drives, Rapid Storage minimizes latency and physical hops. Integrating gRPC, a stateful streaming protocol, and hierarchical namespace enhances file system-friendly operations and performance. While currently in private preview with a target for allow-list GA later this year, users are encouraged to sign up and provide feedback.
Presented by Michal Szymaniak, Principal Engineer, Google Cloud. Recorded live in Santa Clara, California, on April 22, 2025, as part of AI Infrastructure Field Day. Watch the entire presentation at https://techfieldday.com/appearance/google-cloud-presents-at-ai-infrastructure-field-day-2/ or https://techfieldday.com/event/aiifd2/ for more information.
Transcript
My name is Miguel Shimanek. I'm a principal engineer here at Google Cloud, and today I have a distinct pleasure to talk to you about rapid storage and how it is powered by Colossus, our foundational distributed file system at Google. So if you have joined us at Google next, uh, just two weeks ago, then you would have seen the announcement of, of rapid storage yourself.
Essentially the discussion was going along the lines of can we have a storage system, which is offering a latency of block storage, throughput of parallel file system, ease of use, and scale of object storage. And it's very easy to start and manage. Uh, and there are some good reasons why these particular aspects.
Basically we are seeing our customers in GCS to, uh, work around all kinds of performance limitations by tiering to say persistent disc to have the latency of block storage and then offload completed files to to GCS once it's all done. Uh, we will also see customers look for parallel file system functionality, like even for last that was just discussed. Uh, and obviously we would like all our customers to be, uh, happy using object storage, our GCS product, which is excellent to manage content at scale and to, uh, grow pretty much infinitely in all directions.
So that was our goal when designing, uh, rapid storage. And we have basically announced it just now two weeks ago, uh, at Google Max. What rapid Storage is basically a new zno storage product made available within cloud storage portfolio.
So it's fronted by, uh, cloud storage views. That's basically, we really want you to be able to use rapid storage as a, as a parallel file system. That's also why we make sure that it has this insane throughput of 20 million requests per second, six terabytes per second throughput on the data side, but at the same time, it's still a zone or bucket and a bucket in a GCS sense, and that's where you'll get all this goodies from cloud storage itself, 11 nines of durability in a zone, three nines of availability in a zone, and then most of other cloud storage features, uh, right outta the box exactly because it's integrated into, uh, into GCS proper.
Probably the most exciting piece though is the latency, which is sub millisecond, uh, for random reads and appends after you open the file. So the fuse aspect is coming to picture here as well. If you are using, uh, a zal bucket through this file system interface, then you have a separation of open and read, open and write, open and append.
Uh, that's where, uh, this distinction is important. After you open a file, that's where this huge latency wins and throughput kick in. And the sub millisecond latency is obviously many times faster than other leading hyperscalers.
It's actually 20 times faster than our own product right now in a single region. And you can see this test here on the right hand side are the actual performance test for small scale region append, and it's roughly half millisecond in either case, also very, very predictably. So that's another important thing.
As you know, it's very predictable performance. There is no zigzag like effect and you don't need to worry about that. So in terms of positioning of zonal storage and rapid storage in particular, uh, versus everything else we have in cloud storage in DCS, uh, that's how it looks.
Rapid storage technically is a new storage class and it's only available in this new zone out type of a location. So when you create a bucket as opposed to choosing in region, dual region or multi region as a bucket location, you can also now choose a zone. And once you choose a zone, the only uh, storage class available in the zone is indeed rapid storage.
So even from this picture, you can already see how rapid storage is kinda sitting on the side of, of the rest of GCS, and that's for a good reason. Like rapid storage is essentially, uh, separate data stack that we have where we, where we have focused on high performance, uh, rather than the, you know, manageability and scalability and accumulation of data. That was one of the best parts of GCS.
So yes, where, okay, so we've got all this, but you guys also talked about the cache, whatever, of course, I can't remember off the date anywhere. Cache cash anywhere cache. So position that, where does that anywhere Cache is not on this piece.
Where do you put anywhere cache on there? Yeah, So anywhere cache can be in any other bucket type to accelerate reads from any zone, uh, that you want. You can have your regular bucket, which is growing in GCS as it as nature intended, and then you can put anywhere cache in front of it to accelerate reads for your VMs in a particular zone.
It's especially useful, uh, in multi regions because that's where you also accelerate cross region data transfers, which are typically a source of bottlenecks and and slowdowns. Yeah, so Kimberly, you can think of anywhere cache being used on any of the red buckets region, multi region, dual region with rapid storage. There's no need for anywhere cache because the latency already delivers in the same zone.
Understand. So when, why would I choose one versus the other? How are you positioning with a client?
So that will depend By just asking a dumb question, everybody else understands it. No, no, it is a good question because there are definitely trade-offs. It depends upon where your data is residing.
If you wanna extract data from like a multi-region regional bucket into rapid storage and use that locally, so that could be similar like a subset of your training data. If you have very high throughput needs of the six terabytes per second, or if you have a need to really scale the QPS rapidly, rapid storage will do that differently. You can think of rapid storage as US engineering cloud storage object storage to be similar in what you just heard about apparel file system.
It's not a full pex, but it's trying to minimize the typical latencies that you would see in an object storage bucket for AI workloads. Okay, But so with this then you are, I don't wanna call it competing with luster file system. So how, where do I position then luster file system on this one?
Is it still faster or is it multi region or? So like most things it depends, but right now go manage luster rapid storage is in private preview. It's gonna be generally available later in the year.
So this is a little bit more visionary things that we're laying out and, and having people begin to try and test. But depending upon your AI workloads, you may be well suited for rapid storage. If you're designing like for object storage and, and that type of latency, or if you're coming from on-premises, it's gonna be the easiest to go right into managed stor, managed luster luster and use PFS or if you're coming from another hyperscale, the same thing.
Um, and the other aspect is if you need and how you've written it, if it's ICS versus the rapid storage with GCSF, the file level capabilities, that's not a full file system, it's file like attributes. So there are some nuances there based upon your workload. Okay, thanks.
What's the durability on this, on this one as opposed to the other? The red ones? Uh, durability is higher and availability is higher as well.
So higher Than no higher than the other. Okay. So you have, so the other ones are going to be, this is sitting in one zone versus the other ones are going to through the other zone.
So your durability on them is like, I think it's the 11 nines is kind of what I was, what my peers just said. Well, the durability is different than zal availability. Okay.
Right. So both managed luster and rapid storage have redundancies built into the hardware for failures that would happen within that. Okay.
But if the entire zone goes away, that's where the three nines of zonal availability applies to both of them. Yes. And you exceed that zone of availability.
If you go to region, dual region, multi region, that's where you get into the four or five nines. Okay. Yeah.
So that's the critical thing to, to watch out for that zonal installation is almost by nature, potentially more susceptible to failures because there is no redundancy that all the other, uh, bucket locations give you. Uh, so like I said, it's uh, a little bit of a different beast versus the rest of GCS. So let's take a quick look how exactly we put it together.
So like I said, it's um, exposure of Colossus file system that we have internally for many years and that we are using for many of, pretty much most of our products, uh, at Google. And now we are bringing it to, uh, to the cloud, the whole power performance throughput, uh, all the right metrics for others to use, uh, in GCP, uh, obviously the rest of GCS is also used on top of is is built on top of Colossus, but this is a federation of many colossus clusters spread that spread out all over the world, metros, uh, continents and so on and so on. So what is happening here is that Colossus in a particular, uh, cloud zone is brought to the forefront of, of, uh, cloud storage.
You're basically accessing cloud storage directly from the cloud storage frontend, uh, which is visible here on the right hand side. The bottom part of the increase how cloud storage frontend is the Colossus client. And that Colossus client is talking directly, uh, to Colossus storage node.
So you are literally one RPC hop away from data on physical hard drive, uh, and the entire stack now, including also the customer VM running GCS views is enclosed in this one cloud zone. So it's all tightly co-located, uh, for low latency in the network and for minimal number of, you know, potential physical hubs, uh, involved in this communication end to end. Uh, and now if you're looking on the right hand side of this column, we have GRPC, which is the streaming protocol fronted by fuse and then terminating in cloud storage front end.
We say it's a stateful protocol because it keeps track of, you know, what kind of files you have open. So as soon as you open A-G-R-P-C stream to a particular file, that's essentially where file open happens. It happens on the fuse side for, for your application.
It also happens on the Colossus side. That's where Colossus client calls open in its own metadata and, uh, resolves the location, for example, of the hard drives you will be talking to. So once this open happens in GRPC streaming, you are all set up to be just two RPC hops away from the physical 12 drives in Colossus.
And just to complete this picture, um, the fourth element of rapid storage here is hierarchical namespace. That's basically where we make sure that also the cloud storage metadata is represented in a file system friendly way. So if you have heard about hierarchical namespace from last year, it's essentially, uh, a new type of a bucket where we present bucket contents in a hierarchical way where you can actually anatomically rename folders a automatically rename, uh, objects exactly like you would expect it to work, um, in a file system instead of simulating this operation slowly and gradually, uh, using a flat list of of flat bucket storage.
Uh, so basically all three pieces in this picture are file system friendly. Exactly. So when they come together, there is no artificial uh, simulation.
There is no artificial degradation of performance across the stack. It makes it sound like a cloud storage fuse is hosted on rapid storage. Yes.
So cloud storage F fuse is mely a representation or presentation layer for your vm. Uh, So, so FUSE is accessing this. If you're using rapid storage, accessing the object stores, you're using those, Yeah.
So Fuse you can use for any, any type of bucket. Ah, it is different in rapid storage is that you have this GRPC streaming and with GRPC streaming the separation of open and read and append, that's where it kicks in. You are basically overlaying open fuse with open the stream, resolve the colossus metadata, get me the hard drives ready, and now you are ready to go with this insane throughput end to end.
Um, alright, so let's take a quick look at how Colossus works internally. This is going to be obviously a very high level overview. So on the top right, uh, part of this picture you can see Colossus curators, which are basically metadata servers hosting the mapping of files to where they are, uh, and colos custodians, which is basically a whole bunch of maintenance workers which take care of transcoding garbage collection, all kinds of usual operations known from distributed file systems.
And at the, at the bottom you have a massive fleet of disk servers, uh, which store actual bytes of your files. And then on the left hand side you have, uh, colossus client libraries in the cloud storage case, it's basically a part of the front end of the serving process and it talks to Colossus curators for metadata operations. And once, uh, it knows where to go for data, it goes directly to the disc.
So that's once again, an important principle that Colossus client library is just one hop away from the physical data on a physical disc. And then, uh, it's not always just the client library in a serving process. It happens in most cases.
But for example, if you are running a persistent disc, then internally it's also using Colossus, and then if it's a GCE VM with persistent disc attached, then the Colossus client library, which you can see here at the bottom of the left hand side, is actually running directly on the network interface. So that's actually pretty neat and slightly magical how you can have close client library on the nick and then it benefits from the same performance, uh, improvements again separating metadata and data and having just one hop away from the physical hard drive underneath. Very quick question.
What is, uh, titanium nick, just to be sure, Titan titanium nick Offload things on itself? Yes, that's exactly the offload. Yeah, so I, I don't know exactly which uh, titanium hardware we are using here, but we can figure this out if you ask me after the session.
Okay, Thanks. Alright, so let's see how, uh, the whole story of open and append separation works in practice. So here we have an example of BigQuery for a change, and here BigQuery in order to append the block of data to a particular filing, Colossus will first talk to curator exactly to resolve the location of couple of disks it needs to talk to.
Then it'll get the block of data to append and then it'll talk directly to the three Ds in this picture, uh, to append the data as needed. So much of the intelligence really of, of the magic or colossus, uh, is uh, is actually, uh, enclosed in the client itself. This is basically where the heavy lifting happens as well.
All this intelligence of flipping and switching between metadata and data and making sure that the performance is right is done because of, uh, heavy investments in how client library works in Colossus. So that's basically, again, open goes to curator, you, you get back this file handle, you file handle in GRP. In GRPC sense becomes part of the context for the, for the GRPC uh screen.
And then you can have this direct disc access from Colossus client and indirectly also from Fuse behind GRPC. So obviously a natural question at this point is what do we do with, uh, conflicting writers? If you open the same file for append, uh, then we would like to have some assurances of things not going sideways.
And that's where, uh, Colossus basically assumes a single writer, uh, mode for all files and tracks the versions of who is now the owner of the file that can be appending to that file, uh, directly in the data, uh, stored on the disks. So here in this picture you can see how uh, first you are, we are writing from the client one, uh, and that's the green path, say, uh, and it's writing to three different disks. If there is something wrong, the client gets disconnected, somebody else takes over.
By the virtue of calling Colossus curator and Colossus curator responding to open, uh, part of that execution will be also to mark all these data notes that now a different version of the file is being, uh, used. So that client too can start appending the, uh, contents to the file as before, but client one, if it suddenly comes back, uh, will now be rejected because all these data notes now know that the version one is no longer the one that should be accepted for updates. So you're basically versioning the handles and and versioning the, uh, data fragments and it's all managed by, uh, curators during this open file call.
So we also mentioned the al next place. Yes. Quick.
So, so all of this sort of feels like it's another, it's a sort of very fine tuned version of a parallel file system, right? Where you're just paralleling access to the individual discs, right? Uh, you could say so, yes.
Okay. Yeah, so you know, it's not a secret that Colossus itself, it is, its itself a parallel file system that we just use internally. It's not fully ics because we also have this convenience of writing our applications specifically for the semantics that Colossus offers.
Mm-hmm. And Colossus was designed with performance in mind and not necessarily with full coverage of ICS because Right, we don't need it all. But you're right that, you know, the manifestation of Colossus, uh, is essentially propagating the same sort of properties also to the cloud.
Yeah. Okay. So just Jim Rinky from ZDC, just so I can kind of wrap my head around this, I get you going through specific mechanisms to make sure that old version of file, new version of file, what kind of files would constantly being hit like this?
Would it be log files? Would it be the files used for checkpointing typically? Would it be possibly brand new files coming into the file system from say, uh, an IOT system?
Just can we have a practical Yes. Session of that? Yes.
Yeah, so I probably was not emphasizing AI enough in the stock, but that's because Colossus is used for all kinds of purposes inside Google. So we have Spanner database running on top of Colossus. We have, we have big query running on top of Colossus Media.
Streaming is running on top of Colossus, so files can be anything you want, whether it's, you know, just fine grain transaction logs or large volume videos coming in from live media or things like that. Okay. So it really is a multipurpose, I would even say beast, which is powering all of Google.
And in some sense, by the virtue of that, you can be sure that if it's good enough for everything at Google, it's probably going to be good enough for you as well. Right. But not necessarily just for Checkpointing, right?
I Think Checkpointing and right throughput of Checkpointing is one of the selling points here. Okay. So if you have this insane throughput of bytes coming in from Checkpointing, then that's also an option.
Okay. We just don't necessarily differentiate this versus everything else with massive data coming in. Yeah.
Thanks. Thanks for bringing it down to fourth grade level for me. Sure.
Thanks. I needed that. Alright, um, yes and then yeah, we talked about hierarchical namespace as the metadata piece inside the GCS.
So in the middle between Fuse and Colossus itself, there is still GCS and all the features of GCS this time represented in a al main space. So that's something we have launched already last year. It's available in all locations, not just in Zno.
Uh, and it does offer this file system experience instead of just flat list of objects in a bucket. It is especially attractive if it's fronted by GCS views because that's where you actually can see this, uh, logic of having the metadata level operations for say, renamed folder, which is flipping anatomically rather than going through this gradual process of rewriting all the objects involved. And it also happens to have higher performance for, uh, higher throughput for, uh, new buckets, which is where, uh, we are actually positioning it as a, as a performance advantage on the metadata side.
And if your workload is doing, uh, things like folder rename, which some workloads do, uh, in essentially to advance the stage of the computation, then the performance improvements are actually quite massive. Like we have seen from some customers, uh, reporting how renames are blazingly fast this time and the entire workflow can be 20 times faster too. Exactly because of how, uh, the dependence of, of the rename, uh, is there.
And it constitutes a big part of execution time. So if big part of what you're doing in your workflow is actually renaming, uh, folders in order to stage the next step of the computation, then on traditional block storage, uh, cloud storage, you would have to go and rename object by object, which takes a lot of time in h and s or in iCal namespace. It's essentially a single atomic operation in metadata, which is instantaneous and that's how it accelerates the entire execution of such a workflow.
Alright, so that's about all I have for, for today. Uh, here is a bunch of references for rapid storage. We have recently published a number of blogs, uh, blog articles about rapid storage and high performance storage for AI in general.
Uh, and then if you haven't seen those videos from next, they are very well done by my product colleagues. Uh, and if you have not seen them, uh, you should definitely go and do that. It's really well made presentation.
I was there at next, it was the first time. Unbelievable show and hopefully we'll see all of you there next year. And yes, if you are interested, there is the signup page for private preview.
Sean was mentioning earlier how it's now how Rapid Storage is now in private preview. We would love to, uh, see you sign up and use it and give us your thoughts after, you know, playing with it yourself for a little while. I love the fact that we're getting this back to back and, and half the time, half of it's going over my head right now, but um, with Fuse I thought there was a capability to do renaming, uh, the name spaces with Fuse.
Was I mistaken that? Did I miss that? I'm like, I'm trying to go back to 9:00 AM I haven't had enough coffee.
Yeah, so, so Fuse as a, as a file system interface, it does expose the file system, API and one of it is rename, it's rename now depend, yeah, depending on whether the bucket is actually hierarchical or flat, different behavior will kick in. Either you'll be able to, uh, use the folder level A PS in H and S and then it's an instant flip of your directory rename. Or if you realize this is actually a flat bucket, it'll go through slow renaming.
So the functionality will work, but the performance will be drastically different because you are now adding all this overhead of per object operation and that's where you can probably spend a lot of time. Okay. You can think of Fuse as a client, it makes the, it makes the bucket look like a file system, but the bucket doesn't actually act like a file system.
Har giggle namespace actually changes the innards of the bucket and the structure to make it more like a file system. So the combination of the two is what gives you the best performance. So as Al was saying, if you use the rename and F but you don't, you have a flat name space in the bucket, it's gonna do it a rename by one object by copy and delete at a time.
If you do it when the bucket has been optimized with hierarch equal name says, now it actually has that structure. The rename does the entire folder all at once. So it's a combination of the two that's most effective.
Uh, fuse picks it up from this, from the structure. Okay. Yes, you don't, you don't have to do anything with Fuse.
And uh, between the two of you we've seen this private p preview reference a few times. So in terms of timelines towards more general availability, uh, when you get the private preview attached to this offer, what's the timeline look like for, for end users that are not part of a private preview? When, when, when would they expect to see something like that?
Just generally appear as accessible to them and their services? So that differs between different features that we are launching, depending on the complexity. I think for rapid storage we are targeting later this year for at least uh, allow list GA and then later full GA as well.
But to be clear, anybody's welcome to come and like sign up and, and get interest in the preview. Okay. So the sign up here for private preview is literal.
Go sign up for it. Yeah, Go sign up for this and we'll love to hear your feedback about it.