Seamless Business Continuity and Disaster Avoidance with Qumulo
This Qumulo presentation at Cloud Field Day 23 focuses on delivering business continuity and disaster avoidance through its platform. Qumulo leverages hybrid-cloud architectures to ensure uninterrupted data access and operational resilience by seamlessly synchronizing and migrating unstructured enterprise data between on-premises and cloud environments. This empowers organizations to remain agile in the face of disruptions.
The presentation dives into two main approaches. The first leverages cloud elasticity for a cost-effective disaster recovery solution. By backing up on-premises data to cloud-native, cold storage tiers, Qumulo allows for near-instantaneous failover to an active system. This approach utilizes the same underlying hardware performance for both active and cold storage tiers, enabling a rapid transition and incurring higher costs only when necessary. This is a more cost-effective alternative to building a complete, on-premises continuity and hot standby data center.
The second approach emphasizes building continuity and availability from the ground up. By deploying a cloud-native Qumulo system, the presentation highlights the benefits of multi-zone availability within a region, offering greater durability and resilience compared to traditional on-premises setups. Qumulo’s data fabric ensures real-time data synchronization between on-prem and cloud environments, with data creation cached locally and then instantly available across all connected locations. This offers significant cost savings and operational efficiency by eliminating the need for traditional replication and failover procedures.
Presented by Brandon Whitelaw, Field CTO and Cloud GM, Qumulo. Recorded live in Millbrae, California, on June 4, 2025, as part of Cloud Field Day 23. Watch the entire presentation at https://techfieldday.com/appearance/qumulo-presents-at-cloud-field-day-23/ or https://techfieldday.com/event/cfd23/ for more information.
Transcript
I think the, the interesting aspect to some of the questions asked with the the stories, what makes it really real is the customer examples that are relatable. And for those that are in us in the industry for a long time, it's easy to kind of geek out into our specific domains on, on things. But what I've always loved about being in the unstructured data space is it's really easy to explain what we do to, like my 7-year-old, or my mom, my dad still has a little trouble.
He asks me, what is it again? You do about every other time that I see him? But otherwise, uh, it works out pretty easily.
And, um, it's because it's relatable. It's, it's the more human aspect of data interacting with entertainment or images or, you know, x-rays or any of those stuff that people normally understand is, is data that they interact on a regular basis. Um, so I'm gonna talk about two different scenarios related to how do we protect and maintain data as the life force of driving human achievement forward.
At least that's what our customers tell us. Um, the first one is kind of just the, the standard approach, uh, out there today. This is something that, uh, I worked a lot in over the years, and a lot of our customers still believe that the only options out there today is when we're trying to figure out how to protect data.
Uh, the first thing we think of is, well, how much, uh, downtime do you want? Uh, how much business continuity do you want to interrupt? Uh, and how much do you wanna spend?
Those are usually the variables and they're at odds with each other. On one hand, you could say, I'm just gonna have a, a dumb Bitbucket store. My data, I'm just carrying about if that data is lost.
But hey, if something happens to my primary, I'm gonna be unavailable. I might lose data there. It might be days, weeks, in some customer circumstances.
I've been with months to get that data back up and running business going. Um, and it's interesting, you know, I have a good friend who worked for a travel site who was cis admin, and he was really trying to get the business to agree to build a secondary data center to have as a backup. And they just couldn't justify the expense, or I guess agreed to the expense of course, until the primary went down.
And then they quickly calculated how much revenue lost per minute that that site wasn't running and helping people book trips, right? That's hindsight's 2020. But you know, the dilemma here of, you know, pay for everything twice and park it on standby.
So I maintain full business continuity. That sounds great, but not a lot of businesses can afford to do that. They end up doing the option on the right.
Uh, and then suffering in downtime and unavailability and disruption to the business. One thing that has been a really popular use case for our cloud native services has been this idea of kind of have your cake and eat it too. What if I could have the best of both worlds?
What if I could have an instant hot dr? So customers have a system on-prem, they back that system up into our cloud native cold system running on glacier retrieval or Azure block blob, uh, cold at the lowest cost per terabyte possible. And we have basically the compute on that system running as low as possible, effectively idling, uh, when it's not being used.
Right? When you're not using it, you wanna have the least amount of expense possible. Um, but the neat thing again about the elasticity of cloud is that I can turn that cold system into a hot system almost instantaneously without migrating the data.
Is that the little sturdy, little secret, if you will? Maybe not so secretive at this point anymore, but all the different tiers of active storage in S3 or Azure blob are the identical performance and hardware. It's just financial modeling between the cost to store data and the cost to use data.
So from a file system that's built on that as a primitive, that's really using the object store as effectively a disc and then really intelligently caching in front of it and independently scaling the performance. That's all great, but it means I can park the data on that coldest of tiers. And then the moment you need to fail over and spin VMs up and run the business in the cloud, uh, within five minutes I can go from a gigabit per second of bandwidth to a terabit of, of bandwidth if you want or anything in between, uh, within about five and a half minutes and never migrate the data, never have to worry about rehydrating it, it just completely runs.
Now you might be thinking there's gotta be a catch somewhere. There's a little one, which is again, that financial modeling. The API costs to that GIR bucket or Azure blob cold bucket will be brutal through that period of time, but it'll be a fraction of a rounding error to what it would be to actually build a full, you know, full continuity hot standby data center on-prem.
You know, the one day a year if that, that you actually use it means that I'm only paying for that, that heat when I need it. And the rest of the time it's sitting cold and saving a ton of money. So that's just the, the first approach is again, leveraging that elasticity to get the best of both worlds.
Something you normally, So Brandon, I always thought something like Glacier or something like that was not the same infrastructure as normal S3. Yeah, so, uh, now I lost the battle on the naming rights for Glacier Instant Retrieval. I was not in favor of that name 'cause it is confusing.
But Glacier Instant Retrieval is a online active class of S3 Glacier Deep Archive is the virtual tape libraries that would not be, not be appropriate. Yeah. But, you know, you can reduce the price down from $23 a terabyte down to four by just going to glacier insert retrieval.
And from a total cost of ownership perspective, we're we come in about 35 to 55% less expensive TCO than even an an a cold DR system on-prem, uh, by using that in the cloud. And and you're saying you can go from that scale to top S3 functionality in five Minutes? Well, the nice thing is, I don't have to move the data.
I just, because I can independently scale the compute from the storage, I can just, uh, paralyze the amount of throughput to that same bucket of, of data and get whatever performance you need out of it. Again, from a performance perspective, it's identical to S3 standard, but there's just the API costs. And again, that's an easy price to pay if I have my entire system on-prem down and I need to get everything up and running.
Um, and you're only paying for it exactly when you need to be up and running, not every single day of the year. Just in case, you know, if you're gonna say, what's your ROI on that, you know, completely active hot standby system, you know, the, the one that was on the, the, the right here, uh, well, it's pretty bad 'cause most of the time it does nothing but sit there and burn electricity. But in this circumstance, again, I can turn it from cold to hot in a matter of minutes.
The moment you're done with that, I can bring that right back down to bare minimum idle. Say, I think one of the challenges in the hot standby environments is sort of like a stretch cluster approach. You're synchronously replicating data.
Mm-hmm. And this sort of environment that you're talking is not necessarily synchronous or replicating data between the customer and your cloud. Yeah, I mean, it's continuous replication.
It's not, It's continuous replication. It's asynchronously continu. Okay.
Unlike some of my former companies that actually works really well, uh, across the entire cluster on just the change blocks, as they're changing, being continuously replicated in file environments, it's extremely rare to, uh, to need full, you know, uh, like metro stretch, active, active on the dataset. 'cause you're not dealing with corrupting a database, for example. Uh, or something along those lines.
It's, Hey, if I've acknowledged the file, uh, then I take ownership and take care of it otherwise. But by the way, I'll, I'll show you on my next example how we do even better than what you described in just a moment here. Yeah.
So the question was, what's the impact of our neural cache sitting in front of this, this object here? Well, uh, the nice thing about neural cache is that it allows us to, uh, preemptively cache data, especially on hot systems. It was predominantly built for our hot systems to give basically an all NVME like performance.
Even though the data's sitting on S3 standard or, or, um, intelligent tiering or Azure blob Hot and again, allows us to get to this tremendous performance that normally you could never get out of S3. I mean, the first time I heard that we were gonna build a file system, you know, high performance commercial HPC file system on S3, I was like, oh, that's cute. Okay, wait until you hit the tail whipp latency or the transaction per second limitations or per bucket issues.
But by being able to preemptively cache data based on the heat map of the blocks in the file system, the file naming the file types the user behavior by directory, we basically have a pattern matching system that's looking at, uh, which, uh, caching pattern should I implement per directory in real time? And what's the, what's the active real time inference score of that comparatively to its peers? When should I swap out one for the other?
That's all behind the scenes. The customer never has to care or know, they just care that the, it makes the artist happy or the x-ray comes up very quick. Of course.
And every single time we've come across, uh, a novel application that any of our customers globally come across, we're able to capture that data, create a digital twin, and then train that algorithm on that data through a digital twin regression training that then when done, we are able to, you know, instantly update out to the entire fleet through a non-receptive patch update. Now, every single customer that might use that application's gonna benefit from that. And again, they have to set nothing.
They don't have to specify what application it is. It's by directory always optimizing for. And Where's the cash located?
Is is located in front of your cloud or is it in the customer environment or in both places, Or, yes. So it depends on where the cluster's, uh, deployed. If, for example, if it's OnPrem, it would be on the SSD layer in front of the, then no in front of the hard drives or the NVME layer.
Uh, you know, in front of ql, deep dense QCs in the cloud, it's sitting on instance attached NVME in the EC2 instances that are in front of the, the Azure, you know, hot bucket or the S3 blob, or sorry S3, uh, you know, standard bucket. And that's, you know, fully, uh, globally coherent and distributed across those, those nodes. So as you add more compute and more bandwidth, you add more caching and avoid, uh, the tail whipp latencies back to S3 and anything else in it.
Now on the cold side, the benefit is that, hey, those API penalties are really challenging when really any API cost, there's death by a thousand slices for using object store, right? But if I can cash stuff that's being used, then I can avoid going back and getting hit by that multiple times. And while the typical S3 bucket gets 15 to 20% of its cost per month in API charges, our global fleet average is sub 1%.
So a pretty drastic reduction in those expenses across the board. And of course, yeah, as, as Doug talked about in the cold situation, that also helps reduce those penalties and costs even further. So, uh, yeah, it really allows you the ability to deploy this in cloud of your choice and any particular region that helps with continuity or disaster recovery capability, uh, and then instantly make it hot.
But there's, we can do even one better. This is, I would say, iterating with the elasticity of cloud on kind of the typical approach. The even better approach is to just engineer continuity and availability in from the beginning.
We're gonna show you that, what that looks like today. So I'll start with a foundation first. Our default configuration, that cloud is on the left here.
So this is, you have a, a minimum of three nodes. We can go down to one if necessary for like a, a remote spoke or something like that. But for minimum availability, you have three compute nodes and they're all talking to, uh, a bunch of S3 or blob buckets.
Okay? We take care of that. They're completely committed to us.
How we write the data across those buckets is very unique to optimize on performance and reduce cost. And then we cash in that layer and the clients are sitting above that. Now, while the data is sitting there at 12 nines durability, the availability is a single data center.
Okay? So right off the bat, even in my initial minimal deployment, I already have higher durability than an on-premise system. Uh, even a replicated on-prem system because, you know, your typical scale out NAS with eraser coating is getting between four to five nines of durability on one site.
They replicate it to another site. I'm gonna ignore the fact that it's asynchronous. Let's just imagine it was synchronous.
I would maybe get nine to 10 nines of durability. My minimum config, uh, in this cloud configuration is between 11 to 12 nines, depending on the cloud. So just one instance of cloud native chemo already better than a dual system on-prem, which by the way helps a lot in total cost of ownership comparisons when you can instantly go from two to one in this CIR circumstance, you know, uh, a children's hospital recently, for example, was looking at making the shift 'cause they wanted to be able to get their data, uh, into Azure to take advantage of the, the health AI platform that they're building out.
And having an on-prem doesn't really work. Now, the problem is, is that they have dual system on-prem. If you think about a, a PAX healthcare situation, like you can drive the cost per bit way down on a really dense NAS system that you stretch over seven years.
That can be really, really cheap. But the, the fact was, is that they were having their primary and their dr. And from a durability perspective, this single instance was better.
So they were able to cut that cost down quite substantially, save a lot of money, and put their data in proximity to those AI tools to help them advance their research and help patient outcomes. So we can do even better than this. It's CUAL neural, what's the CNQ Cloud?
Native cumul. Okay. And that's in these three zones, the compute entities in all three zones.
So In this first deployment, the computes just in one zone. So it'd be single zone available, and the data would be multi-zone durable. But the other zones then are effectively Other zones sit there doing nothing they're doing not being used.
But then I can deploy a multi AZ cloud native Qumulo pretty easy. You tee me up perfectly there, I'll pay you later for that. That was perfect.
Which is okay, what if I want to do better? You want the money now? Yeah.
Uh, which is have the compute spread across all three zones. Now, the typical way to get multi-zone available is, believe it or not, in the cloud, kind of exactly what you do on-prem. You buy for a second system and you pay for the data twice, and you behind the scenes, it's effectively replicating between those two zones.
Uh, that's actually not what we have to do here. The data is already multi AZ by default. I don't have to do anything with that.
I don't have to pay for the data twice, which when you get into petabytes and petabytes, if not hundreds of petabytes of data adds up really quick, all I have to do is spread my computer across the three zones. And the computer in this case is your connector and the neural cache kinds of thing? Yeah.
So where the file system lives, the caching layer, the protocol stack, the data services, all that kind of good stuff, and It's running on EC2 or Uhhuh, whatever the components are and what cloud Yeah, Azure, blo, VM or Azure VMs, EC2 on, on uh, AWS. And so for example, if there's a, I don't know, an eeb, you know, a EC2 outage of a particular family or something, and it takes out that, that, uh, you know, Qumulo node, no problem. I can float those ips from one side to another client can see, keep connecting without disruption.
If the whole zone goes down, then I can maintain continuity of that over there. They can spin up new clients on that side if they want connect to the same cluster, no problem. So ultimately it provides really the ability to have higher availability.
And Those zones are cross regions or within the same region? Within The same region? Uh, yeah.
I mean, from a a to, to avoid egress, which is what you hit going from zone to zone, you would want 'em all in the same region. But, you know, if you look at the, the ratings of how those data centers will put together independent power sources, independent networking, triple redundant, et cetera, it's very, very, very high levels. So, you know, taking, having an entire region go down is really, really rare from an availability perspective, from a data loss perspective, it's unheard of.
So again, from a durability, uh, you're, you're much, much better than what you're doing on press. Where's the file system metadata located? Is it back end in S3 or, or is It Yeah, the metadata is sitting both within the compute nodes and on S3 as well.
Yeah. So again, it's by completely disaggregating the sy the front end from the back end of the file system. I can maintain a full continuity of, of the data itself, both metadata and raw data, independent of what happens to the compute.
In fact, we have some customers, by the way, that from a, from just a cost savings perspective, it's a DR cluster or it's an archive cluster. They don't care about the, the cache. It's not a high performance system per se.
Uh, they just shut down the compute at night or when they're not using the system or in between backup windows. Uh, that's fine, no problem. Shut down.
The compute data's still there. When you need it again, you spin it back up, you can replicate, shut it back down again. So you can really, really, really cut the cost down on this quite a lot.
And this is the first file system that at least I've ever come across or done the math on that is equal to or lower total cost of ownership than a high performance on-prem system, but completely running on the cloud and about 80% less expensive than any other file system in the cloud because of this architecture. Like, you know, ference like, uh, architecture or, Yeah. So today, pretty much any X 86 based EC2 will run the system.
Okay. Uh, I mean, there's a very minimal amount of, uh, memory and CPU u again, equivalent to what you run on a little nook, right? Mm-hmm.
It's a very lightweight file system. Uh, our preference is something with NVME, with instance, sta mbme. It doesn't have to be much, you know, our defaults like an I four i four xl.
That's for the K or two xl Yeah, for the cash, yeah, for the caching. But we have customers, for example, that run in AWS local zones where those instances aren't attached and they use a little EBS as the read cache. Mm-hmm.
Uh, we have other customers who, again, as a cold system, they don't care about the caching at all. Good. So You're flexible.
Very flexible. Okay. Okay.
So, okay. And we'll have even more flexibility coming later this year relative to that as well. Okay.
So let's walk through the, and, And those various nodes are all cache coherent uhhuh Globally. Absolutely. Cash coherent.
Yeah, Absolutely. Any client can connect to any node, get 100% of all the data and absolute coherent consistency. Yeah.
I Guess I guess you just can't be mixing like too much, right? Like there needs to be like consistent consistence in terms of like, So technically you can have a mix of nodes in a cluster. It wouldn't be our recommendation.
Yeah. Because usually weakest link will, will go with it. But I'll be honest, the interesting aspect here is that almost all the traffic is north south there between S3 through the cache to the client.
Mm-hmm. Um, we do have logic to whether or not it's faster to grab it from a peer node in cache versus again from S3, and it also considers the cost delta between those two things. Mm-hmm.
Um, when you're talking about a single zone, it's always faster just to grab it from the cache and cheaper. Hmm. When a multi-zone though, you get cross zone networking charges.
So we actually do the math as to whether or not it's faster to go and cheaper to grab it again back from S3 or go across the NVME layer. I have a question here, because you can stretch it globally to outside of the AWS for example, or Azure, right? Mm-hmm.
And put it on the mm-hmm. Intel, no, right? Yeah.
So to have this coherency between the AWS, which is extremely fast, and Nook, which is standing on some BSL connection wherever Yeah. Place, it's, it's, it's a very hard statement that, um, uh, it's a bold statement that the coherency between all those two sites will be there anytime, anywhere. You know, the, the system in my opinion, maybe I'm wrong, will slow down to the lowest, you know, the slowest, uh, link.
No, No. It's, so, it's interesting because the approach that, that Cumul engineering took, I was extremely skeptical. Uh, in fact our, our, our head of product, uh, likes to make fun of me because when I came in, uh, went through a three day process of evaluating whether or not I wanted to come to Cumul, I, the cloud native MLO side, I was like, yeah, okay, I can see that.
I get that you've countered all the things that I could come up with the global namespace, the cloud data fabric. Oh yeah. I'm not gonna hold my breath.
I've seen way too many attempts at this promising a lot and not delivering. What I wasn't accommodating or wasn't really considering was the advantage of building it in the file system under the protocol stack at a very integrated way. Not something that just kind of slaps on top or is an add-on, but something deeply integrated and still maintains, uh, the same file locking mechanisms that normally occur.
So, you know, if a user is connected to this nook and wants to go edit a file, it's gonna go request that lock from where the data is, is housed, where its home is that might be in the cloud, might be the core data center, might be another nook for all I care. Um, but it's gonna, it's gonna go there and ask for the lock it doesn't have, because Locke is always, uh, navigated at the hub of that data for by directory, by the way. It's not all or nothing.
It's by directory. You can have one directory home here, another one home there. They're going both ways, doesn't matter.
But it'll, it'll go back to that home directory and navigate the Locke on that file so that it can maintain consistency on it, not have, you know, cross patterns on it. If you'll do a lot of changes with the small files on the weakest link, you know? Mm-hmm.
Then it'll take just time because of the bandwidth of the connection it could to be synchronized. Yeah, it Could. So, uh, I mean, I think a really interesting use case was, uh, that I was really worried about that circumstance with, we had a large, uh, credit card company wanting to do fraud detection on real time audio calls to their, to their call centers.
So people complaining about things going on, and they built this really cool pipeline in the US and it took a long time to build out and get it running really, really well, uh, where we were accelerating data to feed those GPUs, uh, collecting it from the call centers, all that was working fine. And then of course, they threw the curve ball, which is cool. Now we want to do this for, uh, our data in Japan, 'cause we have a call center there too.
But I don't want to go replicate the, the entire pipeline to the AWS region in Japan. 'cause a, it's a lot more expensive and b and then I have to main to maintain two pipelines. I just wanna have it here.
Can I, you know, have the data in Japan streamed real time through the cloud data fabric into this US West two data center for AWS. And, uh, especially for, from an AI perspective, we are quite worried about the performance on a 200 milliseconds of latency. Uh, but it's up and running today and they're very happy with it.
And the nice thing is that from a, a data sovereignty perspective, I can send it to where that data is never actually ever housed outside of Japan. It is streamed through that US data, uh, AI pipeline, but never, ever in its permanent completeness ever stored outside of Japan. So from a sovereignty perspective, it depends on country by country, but for that circumstance, it was still valid.
Um, so it helped helped them solve that problem through that situation. So, uh, moving along, just for the sake of time here, let's imagine that, again, back to this situation, uh, by the way, June 1st, beginning of the, the hurricane season, so kind of timely here. Mm-hmm.
Unfortunately, uh, probably see some, some bad stories, but hopefully it'll be minimal. And, you know, we have customers that are set up in the circumstance today where, you know, they're looking at, okay, hurricanes coming in, uh, what should I do? When should I fail over to my dr?
Should I, uh, you know, I had a customer once tell me, yeah, we have a DR data center. Oh, have you ever tested it? No.
Why? Well, because if I shut down the business while testing it myself, that's a resume altering event uhhuh. But if actually something goes bad and it doesn't work, then like, you know, we can play the blame game after the fact.
So no, I've never tested the DR. Fail over capability. Well, uh, I don't know if that's a great policy, but Sure.
And actually that is, and shame far from the first I've heard that multiple times. So I mean, the really stressful part here I could even imagine is like, when do I fail over? Do I fail over?
Should I, would it work when I do it? Can I bring it back? I've never tested it.
Or maybe even if I had, I'm still very stressed about this. This, this is a very challenging traditional approach from, you know, your primary to your dr. Uh, and what's gonna work out there.
So what if in contrast, we had this concept of, you know, continuous and constant continuity. So we, we first, again, back to the building, the lifeboat is what if the, the, from the very beginning, I designed my system to take advantage of the highest durability and availability, uh, even in hurricane prone areas by putting the data in the cloud in the first place. Okay, that sounds great.
But there's a lot of customers are very early on in their cloud journey and or their workloads are just not yet compatible with the performance requirements or legal requirements related to getting everything built there. Now, from a data perspective, that's one thing, but running the VMs, the virtual desktops, the application layers, et cetera, sometimes that can be challenging. So by the way, I can pick, you know, it could be AWS it could be Azure, uh, you know, cloud preference, G-C-P-O-C-I, we support all of these.
In this case we'll talk about AWS. So then what I do is I, I build up my on-prem data center as business as usual. I have maybe a high performance all NVME, uh, hardware system running Qumulo on it there.
And I then connect that through cloud data fabric up into cloud native chemo. So now the home for all the data is in the cloud. I'm extending that through a data portal to on-prem.
And this is a bi-directional data pool. It's full read write from either side with strong consistency, uh, 100% of the time. What this allows me to do is I've now put the caching layer where the data's being created, you know, from the surveillance cam cameras across the city from, you know, the imports from the GIS data or from, you know, putting evidence into, you know, at the police stations in, or whatever the case may be, where the data's coming in is here on prem and it's gonna cache that data in that high availability system on-prem.
But the moment the data is written, it's not only going to cache, it's going up into the cloud layer. So block by block, as it's written in, it's going both directions to the cache layer on-prem. And the entire cluster on-prem, by the way, is acting as a cache.
It's not, there's no data that's persisting there, it's just a cache. Now it's a highly available cache. I can have nodes fail.
I can expand that as big as I want over time. In fact, I could even if I really wanted to design that on-prem cache to hold all the data a hundred percent of my dataset, but I'm no longer in this business of replicating data, you know, on a, a certain period of time. It's doing that in real time between the cache and writing it into the cloud instance.
And so sits in, like if I'm the client and I'm running like my own data center mm-hmm. This sits in my own cloud environment and I'm, I'm deploying like the cloud native m like, you know, in my own environment. Yep.
And your VPC and your tenant and your US Now. And that's the previous diagram where you had like all these EC twos and stuff. Yep, that's, yeah.
Yep. And it's, you know, we have a one click deploy from our single pane of glass nexus. Very uncomfortable with running cloud formation or Terraform, but you can also do that as well.
It takes all of six minutes to deploy. Okay. Okay.
In fact, we go to live last year Networking endpoint to access that will be always in the, in the cloud. Yeah. So there's a, either a direct connect or a VPN between these two, but You will always go through the cloud even to your own data center, right?
Nope. So the client talking to the data right here, it, it all has to do is where's the data being created? So in this circumstance, data's creating on-prem, they're going and hitting that cache locally all the time.
Mm-hmm. So They have the second DNS name for a local zone and cloud zone. Yeah.
In fact, the clients OnPrem never even know nor care that the cloud instance exists. We have our own proprietary, you know, protocol that's on a block level streaming that data between them, but all the data is ingressing. So very rather, if ever we would, Let me understand that the DR event is happening and the on-prem customers, the clients needs to change their DNS endpoint to access to the cloud, right?
Or you use GSLB. Yeah, we'll show that through the demo. It's a good question.
Okay. Yeah. You can use the exact same IP address.
So you can suspend a VM in the on-premises environment and store it onto the Qumulo cluster and resume it off of the Qumulo cluster in the cloud, into the cloud. You can have the exact same IP address on the CNQ cluster in the cloud, in the VPC as you're using on-premises. And in doing so, it has the same, sorry, I'm trying to answer his question.
It has the, uh, exact same IP address, the exact same NFS certificate, the exact same mount point, the exact same file structure. You can resume the, you can resume playing the video at exactly where you paused it. Okay.
You know, usually it's pretty hard to move the ip, Sorry, I was in the middle. It Work somewhere because of the networking issues. Right.
And that's why I'm asking if you can have, right. And then potentially, somehow automatically reconfigure the customer, the client's stations 'cause or, or use GSLD and it's just happening in the background. Right.
But if, if you have it in the demo, you know, just, yeah, We'll show you. And so the part of it is also is, you know, you could, you could have this on-prem and cloud. You could have clients here, you could have clients here if you want or use this, use the cloud for, uh, for compute, for bursting compute needs or, or other virtual desktops.
But in addition, we can have also access at the edge. So you could have, uh, a single node on a, a small office, or you could have this little nook, or it could be a hyperconverged system doing a bunch of different things. You don't need dedicated storage hardware to run the entire full fidelity file system in all its features.
And it can connect back up into this dataset sitting in the cloud and provide that local caching and that local like performance, even though the data's sitting back up in the cloud. So the advantages here are, are pretty, uh, pretty extreme. You get extremely high durability, you get real time protection.
You're not having to schedule replication or go through any of that. Uh, no need to fail over or fail back. You have single source of truth no matter where you access the data from.
If I go talk to an s and b share or NFS share or S3 bucket from any of these three locations, I see the exact same data at all time. Uh, you, you obviously caching, uh, minimizes egress if not entirely eliminates it based on where the data's created from. Uh, and then you have, so It must be a point in time where the data's sitting in local cache and hasn't been replicated to the cloud.
And if you're saying that, you know, access to that data can, can occur from any place. Mm-hmm. So if some other location needs to access that data, it's gonna have to go to that cache to get the data.
'cause that's where it only resides. Is that sort of stuff that you're doing, you're Talking about like the moment something's written? Sure.
Yeah. I mean, there's, there's a point in time when the data is written to local storage, whatever that turns out to be. And it's not, it's not actually backed up yet.
Yeah. Or back Written. So the moment a file is created at any of these points, that metadata is distributed and strongly consistent, I can then see that file, open that file.
When I open that file, it will automatically go grab that either from the on-prem cashier that's yet to move it back up to cloud or from the cloud side or from another nook edge device. Wherever that data is, it'll go grab it and pull that out in real time. In fact, the demo that really kind of made me finally a believer that our engineering team did Ray, back when they were first developing this, was they drop a video file into a on-prem system, or no, it was a cloud system.
They drop a, a video file from a desktop into, uh, a cloud native chemo. The moment that file starts transferring, it shows up, OnPrem to the client, connect to the on-prem cluster, and they double click and it plays. And I'm like, wait, wait a minute.
What dark magic nonsense is happening here? What I thought was gonna happen was, okay, when the file gets done transferring metadata will sync. Okay, then maybe I have to wait for the cache to get down here, then I can open it.
But no, it's because it's at a block level underneath the file system, under the protocol layer. I can just stream the blocks on the fly as they're coming through. In fact, the the file system that's that's on OnPrem here doesn't even know the data's not there.
It just goes to the BRE structure and gets redirected back accordingly. So it really advances that, that ability to collaborate and move together. So you get high performance and then multi-site collaboration.
Quick Question. Yeah. Here, oh, hey, now you have my DevOps brain working.
Yeah. What I would do, um, let me know if this has happened too. Now I'm wondering like, so did, what about the reverse?
So like if you did on-prem first Yeah, sure. Do you start seeing it pretty instantly in the cloud? Yep.
Yep. We have, for example, we have customers who have a workflow where two come to mind. One is where they are, uh, writing in images, sports broadcasting.
So they're taking in images from photographers live video feed, it's writing to a local system in the van. Um, and then they have their editors all virtual in the cloud and they're pulling that data up. The moment it starts writing, they start working on it and it'll, by the way, dynamically change the stream of blocks based on what part of the file's being touched.
Okay. So it's not just like, you know, sequential from beginning of file to end. It'll, it'll flip.
In fact, in that video example that I gave, if you skip forward, you see a slight pause and then it'll start playing again, then it cuts because it just moved the bitstream to where that that part of the file's being requested. Mm-hmm. Okay.
And then another, And we also by the way have, uh, congestion protocol optimization across the wan, extremely high WAN utilization. So you don't get, like if you try to run SMB over a WAN connection, you get single digit utilization. Yeah.
It's too chatty. Wasn't meant for that. We're handling it all outside of the protocol layer.
You get really, really high. Almost the entire bandwidth utilized. Okay, cool.
And then, um, another one of my questions got answered, um, 'cause I was gonna ask about, uh, IAC so Terraform and stuff like that. Mm-hmm. So a question I have is like just how, like from a customer standpoint using IAC, like just how fine tuned can a customer like tune the dials to get it to do?
Like, you know, like is there like fine tuning or is it just kind of, I Mean, to be honest, most it handles completely on Its own. That's what I was wondering. Yeah, But look, we're an API first platform, so there's not a single thing that can't be done on the system through an API.
So we have customers that run us entirely as infrastructure, as code deployed automatically through, you know, chef or whatever automation they have in their CI ICD pipelines. Everything's done through that. There's no one standing up purpose-built hardware anywhere.
It's commodity. Even on-prem, they kind of run it like cloud native infrastructure on-prem, even in that environment as well, or in the cloud.