Cloud Rewind for Cloud Native Applications an Overview with Commvault
In this session, Govind Rangaswamy presents Commvault Cloud Rewind, a solution designed to protect cloud-native workloads and enable recovery to a pre-disaster or pre-cyber attack state across AWS, Azure, and GCP. Cloud Rewind facilitates in-place recoveries or recoveries into different tenants and regions, offering business peace of mind and the ability to make an attack seem like it never happened. The presentation highlights the challenges of protecting cloud applications, emphasizing their dynamic, distributed nature, rapid change frequency, and significant scale compared to traditional applications. These factors lead to increased complexity in managing and protecting these environments.
Cloud Rewind tackles these challenges by offering a “cloud time machine” and a “Recovery Escort” feature. The tool addresses the limitations of traditional disaster recovery by capturing not only data but also configurations and dependencies. It uses continuous discovery to track configurations and dependencies. Recovery Escort automates the rebuilding of the entire application environment using infrastructure-as-code, simplifying the recovery process by combining multiple runbooks into a single, automated process. Cloud Rewind leverages native cloud services, such as AWS and Azure backup services, to ensure data management flexibility. This enables options for backups, replication, and recovery within the customer’s cloud environment.
The core benefit of Cloud Rewind, as showcased, is its ability to dramatically reduce recovery time, enabling rapid recovery and testing. Customers can perform comprehensive recovery tests with a few clicks, achieving recoveries in minutes instead of the days required by traditional methods. The tool offers extreme automation by allowing for rebuilding or recovering in a different availability zone, region, or account, which further enhances the service’s ability to deliver application resiliency. It also integrates with other Commvault solutions, promising a unified platform for managing multi-cloud and hybrid cloud environments.
Presented by Govind Rangasamy – VP Recovery Solutions, Commvault. Recorded live in Millbrae, California, on June 4, 2025, as part of Cloud Field Day 23. Watch the entire presentation at https://techfieldday.com/appearance/commvault-presents-at-cloud-field-day-23/ or https://techfieldday.com/event/cfd23/ for more information.
Transcript
Excellent. So I, I'm Ami, former founder and CEO of Pan that was acquired by Commvault and now it's Cloud Rewind. And, um, I think I have about 30 minutes or so to go through.
So what I had in mind is to go through a quick overview of what we can do with, uh, cloud Rewind product within Commvault portfolio, and, uh, give you a quick, uh, demo if you, uh, if you have time, uh, to go through a comprehensive view of what, uh, cloud, rewind and Vari helps. And of course there will be questions in between. I'm happy to answer those as well.
So when we look at, uh, cloud application, um, what you see on the right hand side looks to be like an electronic circuit board. This is not an electronic circuit board. This is a real world cloud application environment on an AWS uh, uh, customer, uh, region.
And this application is running across three different availability zones. Um, it has, um, multiple audiences in the backend. Uh, it has about 146 application load balances and similar number in terms of network load balances, EFS system, and you name it, all the networking complexity.
So when you think about a traditional application versus cloud application, it's night and deck. That's a point that, uh, we, uh, kind of explain and we highlight the fact that cloud applications from the get go are dynamic and very, very distributed in nature, right? Even if those applications were lifted and shifted, uh, similar to this particular customers in my, and this is part of, uh, uh, uh, financial services, which is doing about 16 to $20 billion in mortgage payments every quarter, right?
And the second thing that we, um, noticed is that these application environments are changing, um, rapidly, right? They are not sitting idle. These applications environments are being driven, let's say from the latest buzzwords like, hey, driven DevOps pipeline.
They are literally now being driven by, um, co-pilot, um, created code, which means that they are changing the cloud environment faster and faster and sooner and sooner. And those environments are also going through several other ways, um, uh, through cloud, cloud console, or through third party tools. So at the end of the day, uh, the changes to this environments are, uh, much higher on an average, roughly about 50, uh, times a day in terms of the number of changes that go through this environment.
So the question becomes how resilient are your cloud application? The third thing that we notice is that the scale of these environments are also much higher, right? Um, what used to be a simple data center application, if it was lifted and shifted, or even that created directly on the cloud environment, uh, over a period of time, uh, because you can logically create cloud accounts or now separate VPCs, you'll end up seeing lot more services running for your applications.
Even if you took logical ation of dev test versus performance test versus your production applications, there are a lot more services to manage. So the scale becomes another huge problem. So most senior leaders do not even know what they're running, that that's why they're running into cost problems.
But from a resiliency perspective, you are not really seeing, uh, discovering these environments. How can you really protect or how can you really bring resiliency to this environment? Now, the common solution, uh, that's prevalent in the industry is that, uh, if you want to have disaster recovery, or if you want to have a mechanism to come back from, let's say a regional failure or a z failure or some kind of a misconfiguration issue, um, either by the cloud provider, right?
Sometimes, uh, the account itself gets deleted or, um, typical human error might get it off some of the backend infrastructure for those applications that are running, uh, for some time. So your SLA gets affected, but the solution, uh, the best practice is that you have an ideal environment, or at least some of those application infrastructure created and kept title. Now, few problems with that model, the traditional recovery environment, uh, the problems are, first of all, you are, you end up managing lot more services, lot more infrastructure and lot more, uh, resources just because you have the I environment.
Second of all, um, if those ideal environments are kept in there, I initially created, but your production is changing so many times per day, your ideal environment, your recover environment is drifting away from the production, right? So kind of that ideal environment is useless. Third, the most important thing is that if your production environment gets attacked, if the ideal environment is kind of connected to it, that environment will also be attacked.
So in the end, you have not only completely wasted your budget on an ideal environment. Second of all, you also wasted resources, meaning human resources to manage that. And third, you are not really benefiting out of it.
So the solution could be that, you know, protect all of these resources properly, meaning that not only protect, uh, the data of those applications, it could be all your EC2 instances are Azure resources and behind it you'll have to protect all the disks and all the audiences databases and so on. Not only that, but also protect your configurations because when you want to recreate our rebuild after failure, you got to have all those configurations and the dependencies mainly properly protected as well. So what we are saying here is that change your protection approach so that you can recover your applications fully, not just the data.
We give a lot of data resiliency products, right? We have seen lumio, we have seen our Commvault solution that has been around for decades, no. So we know how to protect the data, but the model of protecting the data is, uh, uh, I need to go beyond just protecting the data to protecting the configurations and the dependencies together as a for roughly That way, you know, at the time of recovery, you don't have to really think about several runbooks.
You know, typically when you think about, uh, a security team, um, after a security, in a security incident, they will have a runbook and application team will have a separate handbook, or your application architecture team would have a separate handbook. Your backup and recovery team would end up having separate handbooks. Instead, you can combine all of that.
You can simplify and have one. And typically, um, this advanced cloud users tend to create their own frameworks starting from the backup product, you know, backup up your data first, and then you can build some kind of a backup service on top of it. Have a replication for your data to go from one region to another region.
And then you need to have some kind of a configuration management database. And then infrastructure code on top of the, you put an SLA so that you can really derive an SLA for your application as a whole. And while you do that, like create that framework or a product by yourself, do it yourself mechanism, your hyperscaler is changing the services underneath.
So it becomes really, really complex. Instead, what we are saying is that you could use Cloud Rewind because Cloud Rewind gives you a cloud machine, and that cloud machine has two different vaults. One vault is for all the configurations, all your EC2 instant configurations, the dependencies between the EC2, um, or Azure infrastructure, Azure infrastructure services, past services and so on.
So pretty much all those configurations and the dependencies are continuously discovered and collected from an application centric perspective because these environments are changing, you know, faster and faster. So the continuous discovery process and the dependency management takes care of it. Second, what Cloud Rewind does is it also orchestrates the data, meaning it take your application data from one region and replicate to another region.
And we employ directly the cloud native mechanism. We are using AWS backup service or Azure backup service. In addition to that, we pretty much replicate the data.
So imagine you have the backup, but we also replicate. So you don't have to buy a replication product. You can replicate to another region, you can replicate to another account.
And it's also incrementally. So we give you a choice in terms of data management. You have seen now you have options to backup and keep it in an ad gap copy or backup, large amount of data or backup right there, uh, on the, uh, particular account and across to another account.
So multiple levels of resilience. Third, when you want to recover, you can pretty much offload that work to in fact rebuild the entire application environment. And that automatically is done by Cloud Rewind using recovery scope.
So we have introduced in the market, in the industry rather two important, um, uh, concepts. One is this cloud time machine for your cloud environment. Second is recovery scope.
What the recovery code mechanism does is it automatically combines all your several run books into a single infrastructure code based brand book. And it's completely automatically written for you based on, uh, the location or the region that you want to recover. And it knows the dependencies.
So you don't have to really understand that the dependencies between your infrastructure services, your past services, your Lambda services, rather, you simply offload all of that. So you can go and rebuild or recover by going back in the cloud time machine across to another zone, across to another region, across to another account, depending on your use case. So you don't have to worry about that application ity.
And people tend to kind of use, uh, Terraform Scripts. But the approach of Terraform script is that, you know, just like what we do, you know, we use Terraform internally, you can see on the left hand side it's a busy slide. But the point is that you don't have to really use the Terraform mechanism, uh, for resiliency.
You continue to use that. We co-op very nicely with Terraform. We continue to use that for your production.
But when it comes to resiliency, that combines, uh, disaster recovery and cyber recovery along with managing Drift can be done in a single product across multiple cloud environments. So it is application aware, so can really go after your, uh, SLA or RTO and RTO, right? And that's what customers like this one and financial services, it's not a huge environment.
9 qui terabyte, what they used to do for a single recovery test used to take 72 hours over a weekend with hundreds of people in the call. 'cause they cannot do it anywhere in between the middle of the week. They can't do it.
And recovery tests need to be done often. And now they are able to do that recovery test within 32 minutes because it's all a single click by every one of the tiers. So five tiers of the applications and the largest tier is basically pulling the data out of the RDS and they are using one of the largest Oracle audience instances on these region.
And it can all be done, it completes in over 32 minutes. And that way they're encouraged to do more and more. 5 to three different cloud environment, you don't have to have ts across multiple clouds for resiliency.
You can get all of the time with a single platform. It's almost like the middle layer of your resiliency on top of your cloud account. That way you can offload a lot of that resiliency testing, um, uh, recovery for individual, uh, partial applications or even achieving resiliency across to another, um, zone or another region.
Because sometimes what happens is that when you have done under the, uh, precept of, you know, digital transformation, uh, lift and shift applications end up in only one zone of a particular region. 'cause that's a data center. And they may not have even done anything with respect to testing that application for recovery purposes in another zone or another region that can all be offloaded to a platform like Cloud Rewind.
And that, uh, typical approach, and this is what, uh, has happened with one of the e-discovery firms when they had an, um, encryption attack literally, and that took down 18 Azure subscriptions. They were paying for all the resources, but they could not use, meaning their customers could not use any of that because they, the customer could not get into the environment, the company could not get into the environment because all of the resources have been completely locked up, including their ad cloud, was able to rebuild everything, isolate it from that infected environment, that infected environment, um, was given to the insurance company to go and do, um, you know, forensics so that they can prevent future such incidents. Whereas Cloud Rebuild was able to rebuild everything in another environment so that they can continue to operate from that environment.
So that's the level of capability that we bring in. Uh, let me stop here and see if there are questions. Um, before jumping on, uh, to the demo itself, I, I would have couple questions for you.
So first question is like, you know, speaking of the Drift configuration trips and stuff like that, so, you know, most of clients I'm working with, they're using infrastructure as a code to manage their cloud environment. So how do you manage to keep like, you know, these two in sync? So, you know, I'm managing my environment with the infrastructure as a code and I have the cloud refined, you know, and there is the event and I need to recover exactly what was in the moment before the attack, and I want to avoid it.
I will need to maintain like this repo with the code to manage my environment and I need to manage my cloud refine to keep it up to date. So there's no drip between those two. Is there any sort of a, you know, integration or how, how do you ensure this We have a loose, uh, integration mechanism?
So first of all, in that situation, the number one thing that the customer would want to have is their applications back, right? And they shouldn't be scrambling, um, across multiple depositories of the infrastructure score. That is one of the biggest problems is why people do not do the recovery testing often, even in this infrastructure score world, right?
Um, they don't do it because it's too complex. Uh, you know, one layer taking care of the networking, another layer for, um, let's say past services, another layer for application services within application layer, there may be multiple repositories, they don't even talk to each other. That is the type of, uh, runbook that I explained.
Uh, what Cloud Rewind would do that is that it directly goes to the infrastructure. It doesn't go to the repositories of the customers infrastructures, code repositories, cloud Rewind uses the cloud native API to query all the configurations and the dependencies directly from the infrastructure. So because infrastructure at the point is the single source of truth, not your repository.
So we directly go and collect all that, uh, configuration data and the dependencies data along with the data application data. Hmm. So I would, I I I would tend to disagree on what you said because if there is a attack already ongoing, like, you know, my source of truth should be my repo, right?
Like, that shouldn't be what is going on in my cloud environment because what if already, you know, I'm being attacked and you suddenly are pulling the wrong information from my cloud environment. So then when you will be recovering, you basically recovering with already a compromised environment. That is why the cloud time machine helps, right?
You can go back in time to a cleaner copy, okay? And that, that's the difference, uh, in terms of what Cloud Rewind does, is that you can go back, yeah. At the moment of, um, uh, you are running applications, you know, when you run your application environment, you're running across multiple instances, for example, right?
And you are using, um, multiple load balances. I mean, these are all separate code rep reports. You know, your lambdas are using something else and your, uh, BPC is being deployed across, um, multiple regions or even multiple availability zones using a different report.
So all of that is already being deployed. Now when Cloud Rewind comes in, it automatically detects all of that is we are able to, after getting the permissions to discover read only, we can go in and get, uh, for RDS, even if it's a single RDS for a simple parameter group, you have 589 parameters. So you could save the data, you can back up the data, but if you don't have your RDS along with all the parameter groups as it was before the attack, there's no point in having the data 'cause your RDS is not going to run.
So you not only have to have this audience as it was before the attack, and it could be, um, an hour ago, it could be, you know, four hours ago or even yesterday, got to have the same RDS with the same parameter groups along with the rest of the complications, right? So that I can reproduce that application, as I had shown, is one of the tiers of the application. And even in those incidents, if the region is down for whatever, uh, uh, you know, failure mode was for AWS, you should be able to rebuild across to another region.
And what if your production comes back? I should not be interfering with your production. So I could pretty much do it across to another account completely isolated and are within the same account in a different VPC.
Mm-hmm. So for that, um, I can go to Cloud Rewind and do a cloud connection and uh, when I do, uh, a connection, I am basically doing a handshake between Cloud Rewind and the customer environment, right? And I can simply add a cloud connection, uh, for my AWS, uh, put in the the region that you want to go.
And if I'm running in Virginia region, simply select that I can go to Oregon or I can get go to Mumbai and as official region. That's simple. I imagine before the cloud, before the hyperscalers, um, you'll have to have data centers created pre uh, built with all the service and so on, none of that.
So cloud rewind pretty much takes advantage of all the on demand nature of the hyperscalers. So we include that. Yeah.
Right, right. Sorry. Um, I, I I will have a second question.
Um, and Govin, can we go back into the slide please? Uh, and I'm gonna ask the same question like, uh, in the previous session, so you mentioned the Kubernetes, so how deep you can go into the Kubernetes in terms of the recovery It is, including the configuration maps, so all the Kubernetes EKS or a KS and so on, we can go into the Kubernetes cluster configuration maps. Okay.
Okay, cool. So When you're bringing up, um, a cloud application, uh, there are certain sequences that might be required in order to bring up things earlier than later. Um, is that something you guys provide automatically or how do you detect that, uh, with your continuous learning system?
What we do is that when we discover, we even know when a particular EBS volume was attached to an EC2 instance, because we have been at it, right? So based on that information, see we use several mechanisms for that. One is, uh, uh, you know, knowing the data at which point the EBS volume was attached to an EC2 instance, and also we give you segregation across multiple assemblies.
What we, the concept that we introduced in the market is something called a cloud assembly. And assembly is nothing but uh, set of dependent resources for that particular application. You could even create a cloud assembly for a bunch of audios systems, and when you, this is what exactly happens with that financial services, they keep that assessing tier and your s need to come up first, uh, even within the same, uh, application environment.
If the EC2 instance happened to have an EPS volume that was attached yesterday was yesterday, we would know that. So we would bring up that old one first, and that is all done automatically based on your current infrastructure or whatever the time that infrastructure was captured. So you could go back in time and an hour ago or, um, even a month ago, right?
Yesterday, you can go back in time that gives that time series information along with the theory, right? And sometimes, uh, you may be right in terms of, um, very complex application, you will have to tier separately. And for that, we give you, uh, a kind of a super assembly through the API.
So you could combine, uh, a, let's say this particular application environment, uh, needs to come up first, and you can call that using your API mechanism second. So we will give you that option. Third, we give you a mechanism for you to, uh, attach additional, uh, intelligence for your application, uh, here at the post pre-web proof model.
So you could program whatever the way you want in terms of modifying that mechanism, uh, um, before the recovery happens, or before a rebuild happens or after the rebuild happens. So if we have several mechanisms, uh, around that particular in a cloud assembly, which itself is this a complex piece. So, so the other question I have is where is the data actually, um, backed up to?
Is it, uh, sitting in Plumal secure vault? Is it, is it sitting someplace else? I mean, is it, is it within the region?
Is it outside the region? Is customers actually able to control where that data resides? Ab Absolutely.
This is where you simply, uh, put a single policy, you know, I mentioned this particular, uh, thing about capturing your configurations and the data and the dependencies roughly together. Because if you have a database coming up as of yesterday, and if you have a load balance as of today, they will not talk to each other or they may not talk to each other. So even if you brought up your data securely through any of the walls, it's of no use because your application is not running.
So that is the first and foremost. So once you have that view, you put a policy for the entire application and that, that policy could be something like this, you know, capture, uh, the entire configuration and the data every hour and put it, put the, you know, retain number of copies in the same region or ship the copies to another region or put the copies across to another account that is all single policy, one policy taking care of your backup as well as replication and so on sec. Um, yep.
Uh, there was one other thing that I wanted to answer. Where is that application data? The application data based on your policies is well within your own cloud account or in another region of your cloud account or another account that you have dedicated for that particular application or combination of all of this.
So you have multiple mechanisms with a single policy approach. That way you're not working hard when it comes to that distributed and dynamic cloud application. We make it that simple.
It's very similar to what you saw on Plum Set and Target model. Good. Could you use a solution to, let's say, migrate from AWS to Azure?
Or is that out of the realm of, uh, this environment, You know, people always out of the possibility, right? We don't recommend that, you know, it's not about application, you know, it's about, it's not even the data, it's about the dependency. It's about the service.
You know, when you go to, I mean, all those things are part of the, the picture that req that are required to be migrated when you go from one cloud vendor to another, the data, the application, the metadata, all the structures, all the linkage, Everything exists. Yeah. But, but the issue here is that they are not compatible.
Yeah, yeah, Yeah, yeah. I Agree. That's why we don't recommend, uh, but you know, if you are running, uh, uh, really tightly, you know, controlled architecture, let's say on Kubernetes, yeah, we do all that mechanism a if it's completely well within, uh, that Kubernetes system, because it is running completely on containers, on Linux based images, potentially you could go from a KS to Ks RV Ks to GKE And the, the data backed up.
Is that, uh, compressed and things of that nature? Or is that just pretty much an image of the, the EBS volume or the RDS or whatever it is? Correct.
We don't compress the ddu. That is what, you know, our data management products do. So we give you the option, uh, in, in the market, right?
Uh, we give you the capability to achieve resilience. And in Cloud rewind, people typically go for three to five copies, otherwise it becomes expensive. And if the customer sees that it's expensive, they can offload that copy over to our A GP copy, which is Compass two.
Or if the S3 buckets are really large and numerous billions of objects, we give you the option to offload that to Lumia. So we give you all of those facilities. That way you get to choose the level of resiliency that you want from an application perspective.
That's correct. Yeah. As far as the cyber threat detection assessment and all that stuff that, uh, Commvault other solutions provide, is that also integrated with Cloud Rewind?
Uh, it's integrated through our extension mechanism, right? Um, we don't directly integrate because, um, every customer changes and more and more integration is coming and we, we got acquired about 12 months ago. So okay, integration is done, then all of this will be well within, uh, a single platform giving you multiple levels of that resiliency, right?
But at the end of the day, it comes down to, um, the capability to, um, bring up your entire application. And for that, all you have to do is to go back in time, click a button, and then say, I want to go to cross region or cross account and I want to rebuild it, right? Tech day rebuild.
And all I have to do is to select that. Let me put this one up there and hit recovery and then say recover. And if you are doing tests, simply turn it on like an iPhone toggle and hit that recovery when you do that Cloud.
Rewind knows SF an hour ago, and it could as well go back to yesterday, our last month and so on. It has all that information about your route table, the security groups, the load balance, the group, and um, you know, target groups. You know, nobody knows all of this because the cloud provider has been changing, right?
And, and you could do that not only here, but um, across, uh, let's say if you're running on Azure, you do the same thing. You can go back and then hit the button and then do the same thing for cross subscription on Azure. Or you can go to, um, your, uh, Google and pretty much you do the same thing here as well.
And in all of this cloud then their architecture of the cloud is different. And uh, you don't have to really put in people to manage all of that because it's all done for you extreme level of automation. Because at the end of the day, what people are looking for is a mechanism so that I can bring back the entire take day rebuild with the click of a button.
And that's the code that's automatically gets written because it's not huge, it's about 1500. And this is bringing up not only the application components and the dependencies, but also the data so that your entire application is coming up, right? This is the tech day rebuild deployment happening across one of the region.
Your load balance, this, your VPCs, you don't have to even re rebuild the VI mean you don't have to keep the VPC running 'cause the VPC might be changing all the time and you get it done really, really quick because the platform automatically takes care. So you can see all of these are already built and it's running in the California region here, right? That the rebuild is happening, the resources are coming up in California while the production is running, um, here in the Virginia region.
Great. So that's the overall idea here. Thanks Govin.
So, um, first of all, thanks Govin, uh, Akha and Michael, just to in wrap up, uh, I think, uh, I want to touch on a, a few things that you saw over the course of the day. One of the reasons I'm so excited that we now have, you know, c Lumio and Cloud Rewind as part of the, the Conva Commvault portfolio and, and platform is even before we acquired these capabilities, the, the, some of the things you heard Michael Falo go through around helping our customers not only remain resilient, but especially in kind of their complex multi-cloud and hybrid cloud worlds, is we were already managing, you know, exabytes of data in the cloud. Beyond the obvious kind of integration points that we're starting to do with lumio and Cloud Rewind, some of the stuff that we're taking into the platform goes beyond just kind of the straightforward integrations.
There are things that you can do in the cloud when kind of it's all code and configuration as you're suggesting that it's, it's either really hard to do, or in some cases just impossible to do in a traditionally packaged world. So some of the things that we were learning from these teams is how can we apply some of these mechanics that go well beyond integration and, and bring that to some of the other facets of our platform, is just one example, you know, some of the stuff Govin was talking through. I think it, it, it is a great example of how Cloud Rewind goes way beyond just data recovery and really helps organizations start to recover aspects of their infrastructure and kind of full cloud stacks and the examples he provided.
And so expect to start to see some of those ideas and mechanisms that, uh, get applied, uh, holistically kind of throughout the platform, you know, that aren't just applied kind of in, in these sorts of cloud contexts. In, in addition to some of the integration work that we're continuing to do across both of those products. So, um, thanks not only for your time, but also your engagement.
It was great to get a chance to hear all the questions and, and tackle many of them. We didn't get a chance to actually dive deep and do a, a demonstration of the LUMIO stuff. So we'll follow up and we could do that.
One of those for, for those who are interested.