Learn About Scality RING’s Exabyte Scale, Multidimensional Architecture with Scality
Scality’s Giorgio Regni presented at Cloud Field Day 23, focusing on the Scality RING’s exabyte-scale, multidimensional architecture. Scality’s origin story stems from addressing storage challenges for early cloud providers, such as Comcast. They found that existing solutions weren’t meeting the demands of petabyte-scale data and the need to compete with large providers. The company’s core concept is “scale,” and their system is designed to expand seamlessly across all crucial dimensions. This includes capacity, metadata, and throughput, allowing them to scale each of these components independently.
Regni emphasized the RING’s disaggregated design, highlighting its ability to overcome common storage bottlenecks. The architecture separates storage nodes, I/O daemons, and a connector layer, enabling independent scaling of each component. He shared impressive numbers, including 12 exabytes of data currently in production and 6 trillion objects stored, with customers having billions of objects and applications using the system. The presentation also contrasted Scality’s approach to that of competitors like Ceph and MinIO, highlighting differences in metadata handling, bucket limits, and the flexibility of the architecture’s scaling capabilities.
Finally, the presentation covered the multi-layered architecture that supports various protocols, including S3, a custom REST protocol, and file system connectors. The architecture is based on a peer-to-peer distributed system with no single point of failure, supporting high availability and replication across multiple sites and tiers. It can manage different tiers, such as Ring XP, the all-flash configuration, and long-term storage. Scality RING also offers multi-tenancy and supports usage tracking, allowing customers to build their billing systems, with the overall goal of the system being an infinitely scalable storage solution.
Presented by Giorgio Regni — Founder & Chief Technology Officer, Scality. Recorded live in Millbrae, California, on June 4, 2025, as part of Cloud Field Day 23. Watch the entire presentation at https://techfieldday.com/appearance/scality-presents-at-cloud-field-day-23/ or https://techfieldday.com/event/cfd23/ for more information.
Transcript
Hello everybody. I'm Giorgio, Winnie, CTEO of Skeleton, the co-founder. A lot of familiar faces, some new people.
So nice to meet everybody. Uh, today we're gonna zoom on the tech. Uh, before that we will show, I'll show some numbers.
And, uh, I wanted to start with giving the story of skeleton a little bit, the origin story as, as I call it. Uh, so how do we get our names? So I think Paul talked about it at the beginning a little bit.
Uh, we used to work with large service provider at a previous company. So people like Comcast, Cox, cable, orange, KDDI, SoftBank, and, uh, go back to 2008. And they had, they had issues with storage.
And so we asked them, what's your biggest concern? Well, we do compete with Google and, uh, um, Facebook, Google, at the time, Gmail was one gig per mailbox, and Comcast was 20 meg. Mm-hmm.
Comcast was using financial grade sand, you know, like a large beefy, uh, purpose built server. So the same tech you would use to store financial data in banks was used to store pictures of cats and pull on jokes, because email was that in 2008, right? So you have to do something different.
So they look at what Google was doing, what Amazon was doing. Uh, S3 was started and it was all cheap servers, uh, cheap drives. Let's put all of them together and give you an infinite pool of storage.
So they say, we want the same thing, but we have no developers to build it. So it wasn't our business. Uh, storage wasn't our thing.
So we looked at should we, we sell somebody else? Look at some of the tech we could, we use. And nothing, uh, was there.
We looked at a lot of solutions, uh, ton of tarps at the time, uh, trying to solve that problem. So then we said, should, should we build our own? And I was, no, I don't wanna do that.
When we looked at it and we looked at some papers, like dynamo papers, we looked at all this, it was very interesting to us. And finally we decided, 'cause we keep hearing about no Singapore failure, scale out, infinite pool of storage, um, ability to scale in any dimension. And that sounded great to us as French engineers.
So we decided to jump into this and study the POC. The POC was successful. Uh, we signed a service provider in Belgium called T and we jumped into the storage business.
Um, we had a lot of ideas. We didn't have a name for the company. And it's very frustrating when you have ideas and all the names are taken.
So what I did is I generated an algorithm based on all the characteristics that customer wanted, like infinite, uh, ocean of storage, uh, find some names so he find a thousand names, try to make them spell and available, and who is searches and stuff. And we finally ended up with 10. The company voted and scar.
Uh, one. So why scar? Because care, this is our definition.
The ability of a system to expand seamlessly across all critical dimensions. So not just one dimension, all the dimension that matters. And in the presentation from before you heard about more metadata, more capacity, more throughput.
These are all different axis. And, uh, we know how to scale in all these different axis independently. So this slide coming next scares me with too many points.
So I'm not gonna go for each of these. Instead I'm gonna show some numbers, and then I'm gonna show some examples from, uh, customers. And also contrast to, uh, competition Used to be ity ring elements were all similar.
I mean, they were all doing the same functionality. It seems like you've segregated. So it was always disaggregated.
So, um, it was day one. We had the, uh, storage nodes talks to the network. It's the only thing it does.
Then we had the IO demand only talks to the disc and they communicate via a fast communication internally. Think of it has in the, of a fabric, but inside the machine, uh, all, uh, shared memory type of thing. And that was already day one inside the box.
So some customers decided to split that afterwards, right? Uh, we also had a connector layer, uh, on top. And we did that because that's the way we think.
We think, uh, um, microservices, right? And then as we went into issues in production is, oh, it's a good idea. We can move that service to another machine.
And that's when we got the idea to make that more easier to manage, right? But it was already disaggregated day one, some numbers. So, um, in production today we have 12 exabytes of data, but maybe 15 talking now, um, Amazon is 400 trillion object.
Um, sorry, I'll talk about object after. This is just the capacity that's, uh, around 500,000 drives. So we know a lot about drive failures, know, we know how to change a drive, when a drive is gonna fare what's happening on the platform.
This is the number one issue that people have. You know, so network is a big issue. Uh, drive failures happens all the time.
Uh, of course, all our customers, I might say there's at least 10 drives that fail every day, right? But the system is, uh, constantly trying to rebuild and fix these kind of issues. Um, we have around 6 trillion objects stored on the platform.
AWS is 400 trillion. So I'm like, maybe we work 3% of AWS 3%, but the market doesn't agree yet with me. Um, a single customer, we talked about it today earlier, has 300 billion object.
Um, then there's a, we also work with a lot of intelligence agency and another one has 150 billion object. These numbers, there's no limit to how many objects you can have on the same platform. Uh, the meantime to failure is five years.
What I call failure is, um, uh, losing service for more than five minutes. Mm-hmm. Okay?
And so a customer, we get one event every five years. That's, uh, on average. Uh, so it's always a another way to say that.
Um, one customer has more than 2000 application, what we call an application is an S3 client that we don't know about. So some of them are ISV partners, some of them are just homegrown application and there's a lot of homegrown applications, uh, talking into the system. Um, our customers, our service providers, a lot of them.
So then we have end users. And if we count all the end user, we are around 700 million user using our product every day. We don't know they're using it, but it's actually ity in the back, uh, performance wise.
So a single cluster, that's an example. It could do more. But 60,000 S3 operations per second.
I'm talking about full operations, you know, of the entire get put delete in four. So that's 12 million IOPS on disk. If you were to count the actual disk, uh, operations.
Um, 80 gigabytes per second rate is two sites. So it's actually 160 gigabytes per second, uh, across two sites. How do you do that?
You have to ize everything, you know, 'cause TCP has window size issues, you know, uh, if you were to do one connection at a time, you would get a very low performance. If you were to do a big object one at a time, it would be very slow. So we speed things up and try to make it, it as pilot as possible over the wire.
Typical latency, uh, we use hybrid flash and drives most of the time. So flash for metadata, dries for capacity, but gives you between two and five milliseconds. Latency, what I call latency is time to first bite.
So if you have a hundred megabyte object, it's gonna take longer to copy the entire object. Uh, we recently released a full flash solution and now you're looking at microsecond la, right? Uh, and people need that today, especially in ai.
So where Do you get all this data? Uh, are you actually able to look in the customer environments? Yes.
Are you just asking them what they're saying? No. So some customers, we, we can't.
So we know the contract, how much capacity they have bought, uh, other customers. We actually get feed from the customer and we use that for stats. Uh, let's look at some examples.
Now, we had a lot today there. Uh, some of them, some more. So look at this axis.
We call the number of applications. Aurelian talked about this bank that has more than 2000 applications. Um, that means what does it mean?
You know, it's multi-tenancy. So none of the apps c the other apps data, they're all separated. Uh, we have QOS so creative of service.
We know that one application is not gonna impact the other one, right? They have limits into what they can do. Uh, we also do all via utilization tracking.
So, uh, this tenant is consuming 10 terabyte. They should not go more than 20. All of this we track in the system.
That's what we mean by multiple apps. It's actually the multitenancy in the back. And a lot of our customers want to build their own customers, even if they're internal or external.
So we do all the usage tracking. So you consumed, uh, a hundred gigabits per second. You burst it to one terabit per second.
All this data is fed so that you can do your own internal or external building. Another example is the number of buckets. So typically you would, you would think that you don't have many buckets.
So most of our customers do one bucket, one customer. So orange has 45 million mobile subscribers. So that means 45 million buckets.
Wow. Uh, when we design the system, we never thought somebody will abuse buckets this way, but actually we all do. So we had to design a specific bucket, metadata engine that can scale to these numbers.
And when we started, you know, we, we had this distributed ring system, and the metadata was not as scalable as the ring system. So you could design this crazy system that scales automatically, but the metadata was not at the level. So we worked a lot to make it as redundant failover replication, uh, as the actual engine vendors care.
So Amad is an airline service provider. They, they do all the booking, editing, boarding passes, this kind of things. So the customer of ours, they have to log everything for compliance.
And for us, that means a petabyte a day of logs. So we ingesting a petabyte per day, and these are not big files. So it's a lot of operations on the platform.
Can I Ask a question? Please? Yes.
Um, disaggregated architecture, is that microservices or different Microservices? Yeah. Yeah.
The microservices in a disaggregated, there's a notion of efficiency where you're not coping with data around, you know, so it's logically disaggregated, but there's some optimization in the data path. But yeah, it's kind of microservices. Um, so this is, I'm gonna try something.
Okay. So the market is very noisy these days. Everybody does everything.
Everybody is the best object storage. Everybody does all the AI that exists on the planet, even though we don't see the actual project. But I digress and everybody is the best.
So I wanted to compare to other systems. So in my, I dunno, le let's see what happens. So let's compare to a sef.
So I'm looking at, uh, actual documentation online. Okay? So we talked about we can scale two hundreds of billions of object per bucket.
So the actual official self documentation, say if you have more than a hundred K object, new bucket, you will get performance degradation. So, so that, that works for some customers, that doesn't work from others. And that things, that requires a lot of work to go past these kind of limitations, um, simply because we don't have a metadata database.
So they have to do a lot of things in a non-scalable way. But that's one difference, right? Another difference on safe side.
Um, we talked about how it's, uh, the coed. We have a discs, we have the, uh, network demons in our system. We have a connector layer.
Actually the connector layer is three different things that you can scale independently. So that allows you to do these things where a customer, uh, iron Mountain, you could see that he had only 36 S3 nodes and 120 storage nodes. So it's all fully decoupled, sizeable.
Where you want that helps you scale to exabyte of capacity. Yeah, we don't have one customer where at that scale we need the exact same number of storage nodes that we need of connectors is always a difference at, at that scale. Uh, let's look at, at mine.
So we told you many customers have one bucket, one end customer. So orange was was 45 million, uh, bucket. In mine, you would go to 500,000 buckets per cluster but's a but that's a good number.
But there's a limit into how many buckets we can create, uh, in our design. There's no limit to how many you can have of anything. Um, the fact that, uh, a lot of scenario for being a lightweight and easy to deploy in one machine, on one container, uh, what what this means is it's one big stand, uh, monolithic component.
Uh, and you don't have this flexibility of how you mono scale it. They don't have a database either. So an object is a file, meta data is a file that means listing is gonna suffer because you know, you're doing directory operations as opposed to, um, database indexed lookups like you would do if you have your own metadata engine.
So I dunno if it's useful. I just wanted to contrast some differences. I'm not saying something is bad or better, I'm just saying there's difference.
It's not all the same. Um, the layers, so many different type of applications connect to the KT system. Then you have this, uh, connector layer.
So XDM is our internal, when you talk about lifecycle management, tiering expression as the XDM layer doing this, uh, it can also talk to the cloud, talk to Google cloud to AWS to glacier, both the, uh, express service for example, but also the, uh, cold storage service. So Azure, uh, code Azure warm, uh, and so do some tiering of data to the cloud. Then you have the S3 protocol we've been talking about.
There's also our internal rest protocol, which is much more lightweight. That's what we use internally and we give it access to customers who ask for that level of performance. So AI like is a partner of ours in ai.
They go through the, uh, internal a i, we don't want VSV overhead. Uh, we support file system connectors. We have our own, um, file system metadata engine that's distributed on the cluster and we support S-M-B-N-F-S-T-D-M-I, you know, we all want the TDMI to succeed, but it's an object interface on top of file.
And fuse fuse is when you wanna mount directly on your actual, uh, application and don't go through the NFS or SMB overhead. So all the system connectors are, are on the same, um, namespace. Then you have the actual storage hardware.
And as we discussed both layer scale, uh, independently, you need the management interface. What, what Do you mean by q Layers are scaling independently. So that connector layer can be on different servers, can be of on virtual machines if you want to.
Oh, okay. And then you can add as many as you need if you need. Right.
And the storage nodes vis are the actual servers with flash or drives or QLC flash. Uh, you can scale them without having to scale the connector layer. So you need more capacity.
You add servers at the bottom, you need more client performance, you add more servers at the top. Hmm. Okay.
And so what are the requirements for the, for, for the servers, for the compute? Like, you know, is there like, you know, NVME or like what, what, what is Yeah, So there, there's two kinds for the compute. Yeah.
Some parts of a connector are stateless, but we don't need anything. No need for drives. Just give me CPU and memory.
Right. Okay. Some of the connector layers actually need some, uh, flash for metadata.
So you would do an VME or you would do a fast flash for them. Mm-hmm. Uh, still with a lot of compute and memory.
'cause the connector needs that for the storage node. We don't need a fast CPU, but we need a lot of drives. Yeah.
So we'd find servers with a hundred drives, 90 drives. Uh, you'll find QLC boxes with a hundred with a petabyte in one two U box today. Mm.
Uh, that's that layer, the storage Node. And it needs to be like, you know, node. So we, you can't use like, you know, external storage or anything like that for Oh yeah.
When I say node, I'm kind of simplifying. Uh, the node could be a small server plus expansions in the back, and then you add capacity by adding, uh, enclosures of, uh, jbo. That's one way to do.
Oh, So it can be like external storage. So it can be like an okay. Yeah.
Mm-hmm. But it's a jbod storage. It's not, it's not like a, a storage system.
Oh, no, no, no. It's a jbo. It's not another storage system.
Yeah. Oh, I see. Okay.
Yeah. Um, What about the metadata? Where's the metadata stored?
So in that picture, it's on the storage nodes. I'm sorry, it's on the what? It's on the store nodes.
The, the metadata in that picture, that is something that you can correlate as well. You know, and we have customers with the metadata on different servers than the storage. So it's effectively distributed throughout storage nodes.
Mm-hmm. In some cases. But it could also be consolidated to a couple of George nos.
And is high availability for the metadata, uh, part of that or how? Yeah, it's the same as the data we follow the same Suppose so depending on the customer is a ratio, ratio code requirements, he, the metadata would have the same ratio ratio. Coding requirements Yes.
As the data. Yeah. Metadata is a lot of time replicated.
So you will not go to 12 parties, you will go five. 'cause that's enough if you are replicated. Yeah, that makes sense.
Yeah, That's actually a good point. So like how this can be distributed, like, you know, let's say the connectors and the storage hardware layer, like Yeah. You know, are there any like latency requirements between each other?
Or like, you know, can be Yes. Um, so if you do one ring on one side Hmm. Who will be on a fast network with each other, right?
They will be on a 10 gig a hundred gig. Okay. On a, on a fast network, um, you could decide to move a connector close to the application.
Mm-hmm. Okay. But you need enough bandwidth of a ring, but you don't need the 10 gig.
But I talked about if you, so There are no latency requirements between the connectors and the storage hardware. Uh, apart from your application requirement, if you want one millisecond latency at the application level, you're gonna have to have a one millisecond network. If your application is not latency sensitive, uh, then the connector don't have to be, uh, on a fast level.
So where are metadata storage and then on the storage? Storage On the storage? No.
What is the most common bottleneck you see in the architecture from a customer perspective? Are they not buying Yeah. Enough trucks don't have a good enough network, not enough compute in the connector.
Yeah. So the multi-site would be the biggest problem. Multi-site.
Yeah. They want replicating less than two hours, but there's only one gigabits per second. That's the number one thing.
Say, yeah, we, we can try, but there's no chance. Uh, there's the, um, uh, network failures is, so it's not, but we don't have a capacity. It's that the network is unstable.
What happens all the time. How do you, how do you avoid in the multi-site, like, you know, the, the split brains and stuff like, you know, is there any like, you know, third, Ooh, So I have another presentation for that. Oh, Okay.
We're going rabbit hole. Okay, then Yeah. Yeah.
Just high level, like, you know, there's like something where we, okay, So the ring itself is based on code C-H-O-R-D. Yeah. Yeah.
It's a peer to peer distributed system. Mm-hmm. So none of the nodes have to know about each other.
Yeah. They have to know about five or three to five. Right.
So they maintain communications with to these three to five servers, and as long as they're alive, they think everything's happy, everything's fine. And they don't need to communicate to anybody else. So that really simplifies the network traffic.
You don't have Weakness, Basically. No. Yeah.
There's no single point of failure. There's no, no, that's better than the others. There's no masters, no slave.
They're all part of this, uh, peer to peer system. VI can present at one of the other take field days if you want. I have a slides for that.
Um, any other question on this Supervisor? Is, is is not replicated the management monitoring? It is, it is.
Yeah. So just not Yeah, but that slide Yeah. Should make it better.
Yeah. The supervisor Okay. Is not a single point of failure.
Okay. That's a good question. Um, this is the architecture for our largest customers.
Okay. Um, so I'm not gonna present it yet. I'm gonna go through some simplified version and then build up to this architecture.
Uh, if you have ambition to get it to an exabyte, that's what we're gonna deploy. Okay. So the ring xp, is that just designating all flash or is there actually anything different In It's all flash and you have access to our internal APIs to use Ring X.
So I, I would go back to HP after. So this slide, so if you have one site, you know, I think we, we talked about it. I don't know if I can walk to the screen.
Is it following me? Yeah, maybe. Uh, so you have you different apps connecting to the system.
Then you have the S3 protocol. I'm focusing on S3 right now. We give you one global namespace across the entire system.
Uh, we do connect to your, um, internal, uh, directory for users like active directory, ldap, whatever you have, uh, for single sign-on. Uh, we support multifactor authentication, everything that you would need for security. And that goes into the, um, connector layer.
This is what we call the S3 metadata engine and the drives in your site. And these are the two layers that you can scale independently as we talked about. Now that would be a, a one site deployment.
So in this case, the metadata is on the connector nodes not on the storage nodes. Is that hard? No.
So Val, sorry. Uh, Val is a storage node. The S3 metadata is here.
The connector is actually here. So the connector, So there all storage nodes in that block there. I see.
I got You. The connector once without any metadata, uses the storage servers for metadata along with the drives. Okay.
Then if you add multiple sites, then you have this, uh, component of, uh, lifecycle management that comes into play that can manage transitions, replication, explorations between sites. It uses the same metadata, so it has the same capability of scaling and, and, um, data protection for its own configuration. Uh, and that is all managed into the, uh, global land space at the S3 level.
Okay. So we talked about bucket configuration for, for application. Uh, if you follow the S3 protocol, that's all you need to do.
And we manage it in the back, all the policies. And This is working, uh, on your hardware? Like Yes.
The hardware with the software don't need any external virtual machine or anything to, to bring this functions. No. Correct.
It's Just implemented inside, Inside. So, but You said on the, on the upper layer. Mm-hmm.
That's where you can run it as a virtual machine. You could run it, we don't need it, but if you want it, you could run it as a virtual machine. Yeah.
But, but you can run it on the, on bare metal as well. Yes. Yeah.
Okay. But that's, that's, that's over, isn't it? And so If you want only compute, if you only need compute because you disaggregated the architecture, you would do virtual machine.
It's much easier. Yeah. The connector nodes are split across sites or they're connected.
I, I don't How do the connector nodes, how are they distributed across sites? So the connector nodes, if we move the metadata part, only the connector part, they don't have any state, so they don't need to know about anything else. So you just add connectors as you need more connectors.
So it could be anywhere. Yeah. Could be anywhere Stateless.
Mm-hmm. We love stateless when you build up this to the full end to end architecture where you, you add the hot tier with xp, so the flash would be N-V-M-E-T-L-C, you know, the very fast. Um, there's no drive on V XP side of things.
Then we have a tape support that we discuss with, uh, class as an example, uh, where it could be one site, it could be multiple sites. We can do y applications to two sites. Uh, we can, and uh, we can do, uh, all the policies of expirations through glacier cold storage, API, all that is managed from the same, uh, ring.
So Yeah, go ahead. So, so the ring lt, uh, it's a future feature? No, no, we have that today.
Oh, okay. And ring LT means the long term storage existing, But, uh, so, so basically you have the model inside the software you have that can be distributed as well on this bare metal or hardware part. Yeah.
But as well moved to the vm and this is the middleware that it's connecting to IBM or wherever or hp, uh, tape libraries or need something in the mid DMFA tempo IB. So there is no needed for any additional software. It's just that, and it is doing the tiering automatically.
Yes. Okay. And, and the same way it's going for, uh, glassier, for example, A WSB year.
Right. So you have another model which is connecting to the AWS Glassier, right? Yeah.
And you push the data there. Right. Okay.
And One, one for me, so my understanding right, that the Ring XP ring and ring LT are like different S3 endpoint, which you can access, or the intelligence is happening in that like, you know, in that orange layer. Okay. So ring and ring LT are the same namespace, they're not a different system.
So you Oh, okay. Have the same endpoint. Uh, XP is gonna be a different endpoint because typically you want maximum performance.
Mm-hmm. Makes no sense to go for the same, uh, endpoint. Yeah.
So XP will be close to your, uh, AI nodes, for example, right? On the same fin band network. It's not gonna be shared with, uh, And can you extract that endpoint for the, for the lc, let's say?
Yes. Yeah. They're part of the same namespace, but on different network.
Right. That makes sense. Okay.
Yeah. Yeah, that makes Sense. If you do the XP directly going for our low level API, then you, you are, you're stuck on the machine.
There's no, none of the MULTITENANCY works anymore because you made a decision to get maximum performance. Mm-hmm. Do, can you have a multi, uh, LTEL support?
So in, uh, you have two sites, for example, two data centers, and in each one you have the, uh, LTO, uh, and then you drop the data on the both LTOs in the same time. So Yes, that's, that's something we've done already. Mm-hmm.
Okay. I thought you were gonna see multiple tech of tape, different vendors. This, I don't know.
We could, but Nobody will want to risk that, right? Yeah, but I, I'm saying that, you know, just if you have two data centers and then you are protecting yourself from burning out the HSM solution. Okay.
What's, I have five seconds left. Is it what I'm looking at? What's the distinction between Site zero and, uh, the three replicated sites on this slide?
Oh, because it's a full flash site, it's, it's unique. It's full Flash. So it's not replicated.
It's not, it's, it's there for, typically You're not gonna do any fancy replication on that one. It's all for pure performance. I finished with a, a picture of our ui and am I complete out of time or do I have Yeah, sorry.
I'm done. Thank you very much.