Introducing Clumio, Cloud Native Data Protection by Commvault Cloud
Clumio by Commvault Cloud offers scalable and efficient data protection for AWS S3 and DynamoDB, addressing the limitations of native AWS capabilities. The presentation highlighted Clumio’s features, including a new recovery modality called S3 Backtrack, and emphasized the importance of air-gapped backups for data resilience. Clumio provides fully managed backup-as-a-service, eliminating the need for managing infrastructure, agents, and servers. The solution offers logical air-gapped backups stored within AWS, but outside a customer’s enterprise security sphere, offering enhanced security and immutability.
The presentation emphasized Clumio’s focus on simplicity, performance, and cost-effectiveness. Clumio claims a 10x faster restore performance compared to competitors and a 30% lower cost. Key features include protection groups for granular backup and restore of S3 buckets, based on various vectors such as tags, prefixes, and regions. For DynamoDB, Clumio offers incremental backups using DynamoDB streams, providing cost savings and the ability to retain numerous backup copies.
The presentation concluded with case studies demonstrating the effectiveness of Clumio’s solutions. Atlassian saw a 70% reduction in costs, a faster RPO, and RTO. Duolingo achieved over 70% savings on DynamoDB backups, with the added benefits of immutability and air-gapping. Clumio’s architecture, utilizing serverless technologies and AWS EventBridge, enables automation and scalability. The solution offers options for encryption key management and supports data residency requirements with regional control planes.
Presented by Akshay Joshi – Senior Director Development, Commvault. Recorded live in Millbrae, California, on June 4, 2025, as part of Cloud Field Day 23. Watch the entire presentation at https://techfieldday.com/appearance/commvault-presents-at-cloud-field-day-23/ or https://techfieldday.com/event/cfd23/ for more information.
Transcript
So for the next few minutes, we're gonna be talking about lumio, what we do, how we do things, and, uh, we'll do a feature deep dive on S3. Backtrack. It's a new, uh, recovery modality that we introduced back in November, 2024.
Um, and finally we'll do a quick demo of the feature and then I'll hand it off to Gowin to talk about Cloud Rewind. Alright, so what is lumio? Um, we are now part of the Commonwealth Cloud portfolio, and then we provide simple automated backups for AWS data services, um, but with a twist.
Thanks. So, um, it's, it's a fully backup as a service, fully managed backup as a service offering that's built in AWS for AWS services. Um, you don't need to manage servers, infrastructure agents, none of that in your account.
Um, it brings hands off simplicity, and then it's how does this Differ from metallic then? Uh, yeah. So, uh, we are AWS uh, we, uh, provide backups only for AWS services.
Um, metallic is SaaS as well as, um, uh, the other services as well. Um, and it's air gap and immutable by design. And we have high scalability and performance.
In fact, our restore performance is 10 x that of our competitors. And we do all of this at 30% lower costs compared to, um, the other product. Um, the other products in the market.
I briefly mentioned air gapping, but then, uh, let's do a deep dive and understand why air gapping is crucial for data resilience. Um, if you think about backups, there are three types of backups. One is your local backups.
These are stored within your account within the region that your primary data is. Um, they, these pose the highest amount of risk you can maybe, uh, recover from operational errors, but then if an AWS region goes down or your account gets compromised, there's no way to get the data back. Um, next step is cross account and cross region backups.
Again, these provide some level of isolation, uh, compared to local backups. Um, if your AWS region goes down, you can bring it back up. Uh, if your account gets compromised, yes, you can, uh, get the data back, but then ultimately the backups are still within your enterprise security sphere.
Uh, which means that, um, you know, if your account gets compromised, it's only a matter of time before all of the other accounts in your enterprise security sphere gets compromised as well. Uh, so there's no way to get the data back in that scenario, and that's where air gap backups come into play. Um, these are what give you the highest resilience and pose the lowest risk.
Um, these backups are stored outside of your enterprise security sphere in what we call lumio secure vault. Um, and, um, you cannot access the data without due authorization and access. Right.
And these are immutable by design. Actually, can I ask a question? Sure.
When you say air gapping, is that like physical air gapping that happens in OT environments, or is this this logical? It's logical air gapping. Uh, it's a, and we'll go into the architecture as well where you'll understand what I mean.
Uh, the question is, uh, lumia secure vault, uh, where is, uh, you know, like outside of AWS infrastructure? I understand, or it's inside AWS Infrastructure inside AWS. Okay.
And that's a good segue into, um, the architecture diagram. Okay. Um, we'll go into that in just a bit.
Uh, but then, yeah, think of it as a separate account from your own. Uh, so it exists outside of your enterprise security sphere, but still within AWS. So it's still within the same region as the data what?
No, you can choose again, um, air gap backups can be in or out of region. We let you choose where you wanna store the backup. Yeah.
But, uh, if you use a stream, right? I assume you use S3 because there is nothing really else, right? Mm-hmm.
Uh, it's always across different regions. Uh, if I use a stream, Yeah. S3 Oh S3, like Object storage, right?
Yeah. It's generally by design global namespace. Right?
Uh, but the bucket itself can be in a different region. Yeah, it can be, but it's just a logical endpoint, Logical air gapping. Okay.
Gotcha. Yeah. And the backend for the lumio secure vault is S3?
Yeah. Uh, we store all the data in S3. We've built a, uh, we have a purpose built storage for all the backups that we store.
Alright. So what do customers use us for? Uh, we've had customers use us for, um, operational recovery, ransomware recovery, long-term retention and compliance.
These are scenarios where highly regulated industries like financial services and healthcare industries wanna retain the data for like seven, eight years due to FINRA or HIPAA compliance requirements. And then, uh, we've also seen, um, consolidation of backups across various, um, use cases, um, and of course to optimize costs as well. Um, we provide backups and data protection for various workloads, uh, starting with DynamoDB.
Um, and we also, uh, help you protect unstructured data and data lakes in S3. Um, and we also, um, support regulated data and RDS and then the other apps and services like EC2, EBS, Microsoft 365, and so on. Um, yeah, and this is the architecture, uh, that I wanted to talk about.
Um, on the left hand side, you would see, uh, the customer AWS account. Uh, let's say this is in region one. And then, um, the only thing that they need to do is onboard an account, uh, with us using a cloud formation template or a Terraform, uh, script.
And then the IAM role gets created on their account. And literally that's the only footprint, um, that, that we, um, that the customers will have to manage. No infrastructure, no service, none of that.
And then once the account has been onboarded, uh, a logically air gap, AWS account that's dedicated to each customer is created. Um, and again, this can be in the same region or any other region, uh, for that regional isolation. Uh, and then once that's done, it's completely events driven and serverless from that point on.
Uh, we leverage AWS event bridge, uh, to read all the events and then trigger actions on our end. Um, and we have, um, a bunch of serverless lambdas running that take care of all the tasks that we do. All the sausage making that, that happens behind the scenes is done by lambdas.
We auto scale out and scale in, depending on the workload, the size of backups and so on. Um, and once that is done, we store, uh, the data and the metadata immutably in our environment and we prepare the data for, uh, recovery. Again, we support a bunch of recovery modalities.
Um, you can search or query the data using Amazon Athena, and we support instant access, uh, which leverages S3 endpoints behind the scenes. Um, and you can do a full restore as well. Again, these restores can be cross region or cross account.
Um, and yeah, that's the architecture. Could you, Could you go back to that one slide? Okay.
The end of the slide? Yeah. Thanks.
And while we're doing that, um, something that stuck out to me, 'cause I have to work with the data teams a lot, so, um, in what, or maybe you're gonna get into this, but in what ways does it help with the structuring of like the various data, stuff like that? In, like, in what way does it also get involved with DynamoDB? That I'm really curious.
Yeah. So, um, we protect the data. Mm-hmm.
Right? And then it's readily queryable as well. Uh, but we won't help you with structuring the data.
That's the part I'm curious about. Yeah. Okay.
All right. So let's take a look at RS three offering. Uh, again, S3 and DynamoDB is what I'll be talking about for the most part here.
Um, if you compare S3 to any cloud native backup today, uh, we reduce backup costs by over 50%. Mm-hmm. Um, it's the lowest list price compared to all the other cloud native backup solutions out there.
Um, and the other important thing that I wanted to talk about is protection groups, right? Um, so with all the other, um, backup offerings out there for S3, it's all or nothing. You need to protect the entire bucket or not.
Uh, but then we have something called protection groups that lets you choose the backup and restore granularity, right? You can slice and dice your S3 bucket and then backup and restore only a small slice of it if you wanted. Again, these slices can be based on various vectors like tags, prefixes, regions, and so on.
Um, and yeah, uh, we let you, um, back up and restore, um, through protection groups. Is the, uh, is the granularity, uh, a big part of the cost savings there or is something else factored? No, uh, so it's all in, we do a, um, I mean, like I said, we have a purpose built storage, uh, layer as well that automatically gives you, uh, you know, uh, savings on your backup costs.
And apart from that, our list price itself is lower than all the other, uh, cloud native backup vendors, um, out there. Uh, and then apart from that, on top of that, you can use protection groups to slice and dice your data and protect a single, a small part of your bucket. Okay.
So you get, you get the over 50% off the top. Yeah. And then you can save even more by being selective.
Exactly. It's apples to apples. If you were to protect only that slice using us versus any other vendor, then you would see 50% or more right off the bat.
And on top of that, if you want, you can protect a single slice of data. Yeah. Because you get more flexibility.
Exactly. Okay. Yeah.
Um, all the backend that it's in this dedicated account, which is air gap, it's basically, uh, you as a customer, I will not pay those costs. It's your cost, right? Yeah.
And what about, uh, how your, um, license your product by, uh, this protection groups or by terabytes? Yeah, so it's, uh, by terabytes. Uh, we have two vectors, uh, when it comes to S3 backup pricing.
One is, um, the storage cost itself, right? Um, we charge you two and a half cents, uh, gig, um, per month On the front end, uh, post of, uh, storage or the after the duplication? Uh, no, it's front end.
It's front basically what you're backing up. Okay. Right.
And then the second vector is, uh, the object management fee. Mm-hmm. Um, based on the amount of objects that you're managing with us.
Um, some of our customers manage billions of objects, so we charge you a dollar 50, uh, per million objects that you manage Okay. With Us and no charge like, you know, between the regions or anything like that. Uh, we do have a restore fee as well as, uh, cross region transfer fee as well.
Oh, Okay. So on top Of that, on top of that, yeah. Uh, and that becomes an issue only during restore?
Yeah. Uh, for the most part. So, yeah.
Alright. So, and we are also 10 x more scalable and performant. Um, so, uh, and these are customer anecdotes again.
Um, so we've had a customer that was trying to restore a hundred million object bucket, um, compared to cloud native backup solutions. We were at nine Rs, um, and they were at three days. Mm-hmm.
Right? So that's the amount of, um, performance difference that you would see with, um, lumia. And then, uh, we are also faster when it come comes to continuous backup.
Uh, so this helps you with RPOs, right? So while all the other, um, uh, cloud native backup vendors out there support a daily RPO, uh, that's the minimum. Uh, we support continuous backup, so it's completely events driven, and then we keep scanning your inventory, um, you know, to understand any changes that are happening on your end, and then we'll back it up as soon as that happens.
And then we are also 10 x more scalable. Um, we support 50 billion plus objects in a single bucket. Um, if you were to compare that to the cloud native backup options on AWS, um, they're currently at seven and a half billion objects, and we are at 50 plus already.
Uh, and we are also easier to manage, right? So I want to, um, talk about three numbers here to, uh, hit home the point that, uh, we are much easier to manage compared to our competitors. Um, our NPS is 85, and then 98% of our support tickets are proactively resolved, which means that we have 24 7 predictive monitoring of your backups and restore operations.
Um, so even before you know it, even before you wake up in the morning and see the alert, we've already resolved the issue in the backend. Right? So 95, 90 8% plus of our, uh, support tickets are, uh, handled proactively.
And then finally, um, it's a true set it and forget it, um, experience, right? It's, it's the true Gmail experience is what we call it. Um, uh, an average lumio user locks into the portal once every 37 days.
Wow. Right? So that's, that's how easy to manage lumio platform is.
Um, so if, if, yeah, you have your secure vault and there are multiple customers that are actually, but the deploying backups into that secure vault, they're all encrypted. Yeah. By, uh, who, who, where is the encryption key, Right?
So we have two options. One is we manage the keys, or you can bring your own key as well, right? But all the data is encrypted at rest.
And, um, as far as, uh, national Geographic, uh, data, you know, re regional regionalization, you support that as well. I mean, yeah. So the cual secure vault would be in region.
Yeah. Uh, yeah, of course. Uh, so for GDPR purposes, we have purpose built control planes in various regions.
Uh, we have one in Germany, um, and we have one in US, and we have one in Canada as well. Um, so based on your data residency requirements, So both the control plane and the data are sitting in the, the, uh, national boundary Yeah. Data plane.
You can choose at the time of backing up your data, you can create a policy and select where your backups need to go, and you have complete control over that. Are you automatically, um, doing cross region or is that, is that an option that customer has to turn on or is, uh, The customer has to turn on, uh, And the cost associated with It? Exactly.
Yeah. So at the time of creating a backup policy, you choose, uh, where you, where your backups are gonna go. And, and your continuous backup is based on change tracking, is that right?
At the object level or, um, Yeah, it's at the object level. If you create more versions under the same object, we capture those events as well, and then we trigger backups based on that. And, but this is in the object storage option, right?
Yeah. But when you back up, uh, RDS or wherever SQL is there, right? Yeah.
Then it's not the object driven, it's Yeah. So tracked there, you can select an RPO. Again, our RPOs are as low as one R um, so, So I'm sorry, you, you back up EBS as well as object, is that what you just said?
Yeah. So, But in that case, it's not change tracked. Exactly.
Yeah. Okay. Can we, can we double click on the protection groups?
'cause I think like in the cloud mm-hmm. That's where the, the cost saving is probably like, you know, the most beneficial. Yeah.
So talking of like, you know, dynamo db, like how granular we can be with the Dynamo db. Like, you know, is it like the whole thing or like, you know, can we select like tables? Like, you know, how granular we can go?
Yeah. So Dynamo db, you can select tables, the table. Yeah.
Okay. Exactly. And, uh, you probably mentioned in a previous slide, I, I didn't take a screenshot.
Um, is there any support for the Kubernetes, for instance and stuff like that? Like a KS and Yeah, So, um, EKS would be, oh, Sorry, EKS. Yeah.
Yeah, E ks, the AWS equivalent. Uh, but then, uh, yeah, we don't support configuration backups just yet, but then what we've seen customers doing is, um, they, they use a persistent storage Yeah. Which is one of S3 EBS or EFS.
Mm-hmm. Uh, we support S3 and EBS already. Um, and then EKS, they can, you know, it's, it's basically a fleet of, uh, Kubernetes clusters that they can Oh, okay.
But you don't have the granular granularity of like, you know, pods and like, you know, like, you know, storing particular, like you don't have that? No. Um, yeah.
Within the EKS Yeah. We don't support the pods or any of that. Okay.
We don't support the config, but then if you have a persistent storage layer, we can support Yeah. Like the whole thing, basically. Yeah.
Yeah. Okay. Thank you.
Are you using, um, and I'm not sure what the terminology is for AWS, but where they have, uh, where they'll migrate, uh, uh, idle data objects from one tier to another down, are using that sort of stuff, theory, the backend, backend for this? Or you have all of it's currently active in S3? Uh, it's currently ac We don't use intelligent tiering.
Okay. Yeah, yeah, yeah. So we don't use that.
We've seen that we are more cost effective than intelligent tiering as well. And, and that's because you're compressing the data or, Uh, we do a bunch of things behind the scenes. Like I said, the purpose built storage layer takes care of all of that.
Um, we do compression, um, and then, um, yeah, um, we, we also store it in, um, a preferred tier, let's just say that, right. Uh, for cost benefits. Um, and then we pass the, uh, cost benefits out to the customer as well.
And so how much shared infrastructure it is between your clients who are storing, like, you know, how much air gap am I as the end customer from the other end customer? Yeah, so, um, if you look at this, um, each account is dedicated to a single customer, right? There's no PepsiCo problems, no noisy enablers, none of that.
Mm-hmm. Um, each customer gets their own data plan account. Okay.
Thank you. Yeah. Alright, so, uh, let's take a look at a case study, uh, for S3.
Um, so Atlassian one of, uh, the world's large largest software project management companies, uh, they were struggling with their S3 backups at that time. They had 30 billion objects in a single bucket. Um, their RPO uh, was about four hours.
And then they were taking about 248 days to restore a single bucket. Uh, and then they were spending more than $10 million to backup, uh, all of that data. And then, um, they were also repeatedly, um, failing to scale due to the three, uh, 30 billion object scale that I was talking about.
At that time. Uh, they were trying to DIYA solution using cloud native backup, uh, solutions in AWS uh, and then 3 billion was the maximum, maximum scale that they could hit in a single bucket. Now with Lumio, they were able to bring down their RPO to about one hour, um, and they were able to restore an entire bucket in less than two days, right?
248 days versus two days. Think about that. Right.
And we were also able to reduce their cost by about 70%. Um, and again, uh, they're currently at, um, 60 plus billion objects already, uh, and we are able to handle their backups without breaking a sweat. I'm really curious about that.
RTO, like what attributed to such a huge RTO that Yeah. Clio was able to solve, like specifically I, yeah, I'm wondering, so imagine this, right? When the maximum scale that you have is 3 billion objects, um, you'll have to split the bucket into multiple backups and you need to orchestrate the restore operation from each of those 3 billion object buckets.
It's like, it's 10 of those, right? So, So it's a lot of the manual work that you're, it's removing from the customer. Exactly.
You manage it All. Exactly. We manage all of that and we are, uh, scalable, 10 x more scalable.
Gotcha. But, but you have a limitation of number of objects per bucket as well, right? Um, I thought It was 30 billion versus three.
Yeah. But then you don't have to worry about how many, uh, buckets we use behind the scenes, right? Oh, you automate that process.
Yeah, we automate that entire process. Again, we haven't hit those limits yet. Um, Oh, but this guy, it's not gonna be long.
Alright, so, uh, let's talk about DynamoDB real quick. Uh, with DynamoDB, you don't have a lot of backup options in the cloud today. Um, right.
So, uh, they have something called point in time recovery. Uh, DynamoDB backups, um, come with point in time recovery for 35 days. Again, these are stored within the same account as your primary data.
Um, so you don't get air gap or, uh, the backups are not immutable because once you delete the, uh, primary data, or once you turn off Twitter, you lose all of your backups. Right? You can't go back to any point in time after that.
Right. The only other option for you to use is, uh, the cloud native backup option in AWS. Um, but each of those backups are full in nature.
So if you have a terabyte of data in DynamoDB, you take seven copies, you're paying for seven terabytes regardless of your change rate, right? Uh, which is where we come in. We provide incremental backups using DynamoDB streams and export to S3.
And it's always less expensive compared to full snapshots. Again, uh, your first copy itself is air gap, so you don't need to worry about creating another copy of data, uh, um, for air gap purposes. And then finally, we are unified and auditable.
Um, you don't need to worry about, um, having a separate copy of data for operational purposes and compliance purposes. Um, it's, uh, it can be a single copy that access both, uh, and we also are readily queryable and, um, we have automated compliance reporting. Alright.
So I wanted to talk about Duolingo real quickly. Uh, here, um, um, before We goes down that path, there's a lot of cyber resilience capabilities that were part of the other discussion. How does that apply to the cl lumio backup solution?
Right. So, um, we bring resiliency with backups, right? Um, so we air gap is, uh, I think I would probably let, uh, Tim take it.
Um, do you wanna take the Yeah, While, while you're discussing, I guess, uh, my, yeah, I'd like to expand on Ray's question. Like there's a lot of features you were mention in Commvault Cloud that you have air gap and immutability, but certain things like threat detection Yeah. And, and an IRE that are not, yeah.
I guess the question is like, are there plans as LUMIO gets more integrated into the comm portfolio to introduce those capabilities? Yeah, that's the ultimate vision. Uh, we wanna integrate with the Commonwealth platform, but then we don't have timelines at this point.
Uh, but then what you can do today is, um, LUMIO is API first, right? Uh, the reason that our average user logs into our portal once every 37 days is that our API uh, framework is so robust that you can do almost anything using APIs. We've seen customers use, um, Splunk or Datadog for monitoring.
Uh, and then once their APIs trigger a, um, an event, uh, they'll automatically trigger, um, a recovery mechanism using lumio. Um, so our APIs, uh, can be, um, that bridge between the pre-event and post-event. Um, alright, so, uh, You mentioned before that, uh, you know, a customer that had a terabyte of, uh, um, dynamo DB would if he did normal cloud native backups.
Yeah. Seven times be effectively seven TE Terabytes. Yeah.
In your case it wouldn't be because you're doing change tracking and differential backups. Exactly. Incremental backups effectively.
Yeah, we are doing incremental backups forever, and we'll talk about that, um, as part of the Duolingo, uh, case study. So they were spending more than a million and a half annually on just, you know, DynamoDB backups, and then they were able to retain only seven daily, uh, backups. Uh, so if you were to think about it, they could go back only seven days if they wanted.
Uh, now Lumio came in with incremental backups, forever incremental, and then we were able to save them more than a million dollars, 70% or more savings, um, using incremental backups again. Um, they also got immutable and, uh, immutability and air gap, um, air gap, uh, for free along with it. And then now they're able to retain more than 30 copies at a much lower cost.
Is, is, uh, is the incremental something that customer gets to select, or is that something that's automatic? It's, It's automatic. Uh, Yeah.
The challenge with incremental backups is the restore. Um mm-hmm. You're trying to rebuild the, the structures from, you know.
Yeah. Uh, an original, uh, full backup. I sort is what's the restore time for something like this versus, uh, cloud native backup, which would've a complete terabyte version of the database?
Our reto performance is the best in the industry, uh, Best in the industry. Yeah. Uh, DynamoDB store performance is, is really good today.
Um, I, I can get you some numbers offline, but then, um, that's fine. Yeah. I can get you some numbers offline, but then yeah, we are, uh, we've had customers tell us this again, we are not, you know, telling this ourselves, uh.