Protecting the Intelligence and Infrastructure Behind AI
In his presentation at AI Field Day 7, Sathya Sankaran, Head of Cloud Products at HYCU, emphasizes the importance of protecting the data and infrastructure that underpin AI systems. He highlights that while much of the AI conversation tends to focus on GPUs and models, the foundational data that fuels AI often lacks comprehensive protection. During AI implementation, vast and varied datasets are generated, modified, and analyzed—through data lakes, object storage, and lakehouses—posing significant challenges in maintaining consistency, accuracy, and recoverability. Sankaran underscores that much of this data resides in the cloud, making cloud the “home of AI,” but also introduces new threats due to fragmented services, inefficiencies, and blind spots in current protection measures.
HYCU aims to solve these challenges by offering broad and deep coverage across diverse cloud workloads, ensuring consistent and meaningful backup and recovery. Unlike traditional backup solutions that may not cater to AI-specific workflows or protect more than raw data, HYCU’s platform captures the entire ecosystem, including metadata, views, access policies, and AI-specific formats such as enriched JSON and vector databases. This level of comprehensive protection enables traceability and rollback capabilities for AI pipelines, which are critical when dealing with issues like schema drift, corrupted data, or poisoned datasets. HYCU’s approach involves aligning backups with stages like model training checkpoints, and doing so in a way that maintains consistency across fragmented and asynchronous data processes.
Adding to this, HYCU’s partnership with Dell and use of deduplication technologies such as DD Boost make backing up even large-scale AI data cost-effective and cloud-resilient. Their solution minimizes storage use and egress costs by identifying and transferring only changed data segments, often achieving up to 40:1 savings. This also supports cross-cloud backups, offering organizations flexibility and protection from vendor lock-in or catastrophic cloud failures. Ultimately, HYCU positions itself as an essential component in modern AI architecture by centralizing protection, enabling long-term recoverability, and reducing operational risk, all while keeping pace with the rapidly evolving landscape of AI workloads.
Recorded as part of AI Field Day 7 on October 29, 2025. Watch the entire presentation at https://techfieldday.com/appearance/hycu-presents-at-ai-field-day-7/ or visit https://www.HYCU.com or https://TechFieldDay.com/events/aifd7/ for more information.
Transcript
Um, just, uh, as an introduction, I run cloud products within Haiku. So any workload that touches any three hyperscalers, uh, I have portfolio coverage over, and, uh, I'm in year two. Um, uh, haiku as well come from the world of, uh, you know, protecting Kubernetes and, and a lot of the infrastructure elements that, uh, you know, power ai, uh, today as well.
Um, I know very often when people talk about ai, they, they, there is a lot of importance given to the, to the shovels, the GPUs, the models. Um, but not a lot of conversations happen on the data that actually drives the intelligence, uh, uh, for ai, right? So our goal today is to kind of explore, um, you know, how data protection needs to evolve, um, in the age of ai and how do we protect the data sets that are being generated both by AI and for ai, um, uh, these days, right?
And also the, all the infrastructure that needs that gets put in place, uh, to make AI possible for a lot of the organizations. And how today, cloud is the default delivery model for a lot of ai, uh, in, in many organizations. So let's, as we talked about, let's, let's start with, uh, cloud, I call it the home of AI these days, right?
This is where very often people start, uh, when they start their AI initiatives, uh, cloud providers have more access to, you know, GPUs done, you know, an average enterprise, uh, in the market. Uh, and, and you get the burst compute you need, uh, for a lot of the model work that you are doing within the organizations, right? The cloud has become the home of ai, but with cloud, uh, there is a whole bunch of problems that exist when it comes to data protection, right?
We call them, uh, the silent killers. We call them the lifestyle diseases, um, with this dataset, right? You have too many, uh, applications, you have too many as a service, uh, elements that are kind of put together to drive your pipeline, and that creates a lot of the console chaos.
You have to use multiple consoles to manage, you know, what's happening with your, with your pipeline, right? And, and that also creates a lot of a i and automation spread all. And the other part of it, there is a lot of data being generated, but there is not a lot of efficiency built into the cloud.
Um, and again, cloud is in compensated for driving a whole bunch of efficiency in terms of how you store that data. Um, right? So we are also seeing that there is this obesity when it comes to the amount of data that you're storing, you know, within, uh, uh, your infrastructure and within your tenants.
And despite all of this, there are still blind spots when it comes to AI and AI services. There is a large number of services being built in all of these cloud providers to make AI possible, but the data protection vendors haven't really kept up with, uh, all that dataset and even being able to serve that data back to the customer in an AI friendly fashion. You know, that's not really possible today with a lot of the solutions and, and what we are looking to solve, uh, with, with heco.
That's What you, Satya. Satya, yeah. Hi, guy Courier with Futurum here.
Um, can you talk a little bit more about this first one, fragmentation? Um, one of the big drivers, I think, towards data lakes has to do with avoiding this by just having a simple way to throw everything in a single bucket. And that alone, I feel like a sort of a typical data lake or data lakehouse vendor would say that solves this because that provides you with a single con uh, console.
Maybe it creates the challenge of all the chaos having to be reached out to, to pull everything in. But how would you, how would you respond to, to that? Yeah, look, uh, uh, it's a great question.
In, in a couple of two or three slides down the road, we actually walk you through what a typical stack look like. And With that actually, so, uh, please continue, Right? So in, in what you do see is the lakehouse, which we'll spend a lot of time today talking about, uh, are becoming the unified platform in many ways, right?
Uh, it is doing more and more unification. It's multimodal, it's versatile. You can do both, you know, optimize for both storage and analytics.
You know, all of that is true. Uh, but it is not true that today's data pipeline stores primarily in those workloads. It is, again, even if it is primarily you need to preserve, meaning when you're protecting AI data sets, there is a lot of data that sits around these data lake houses.
Yeah. That needs to be protected as well to protect your full infrastructure. Again, uh, great question, and, and thank you for that.
In a couple of slides, we'll actually show you what a sample AI data stack looks like and, and why that kind of coverage and, uh, uh, for that fragmented workload matters. Well, In, in, in my opinion, I'm actually following with my fellow delegate, uh, Keith Townsend here. Uh, his, his guidance on this, it also does create a problem of its own because of data gravity.
The more you collect, uh, data, however processed or managed in one place, uh, the stickier it can be there, even, you know, that, that really hasn't been solved. Yeah. The, the good thing is as we are in the world of, you know, providing insurance for that data, right?
And when data gravity shifts, it's our job and responsibility to kinda shift along with that and build our applications to make sure, you know, that data gets protected. And I think we are seeing that with the cloud where the object storage is perhaps now the landing spot, uh, uh, for a lot of the data sets and the lakehouse of the new systems of records. So between object storage and lakehouse, you're seeing more, more and more data sets kind of landing in the cloud.
And that means more and more applications and the processes and, and, and, and those things will also have to be built in the cloud, uh, because again, the data is gonna pull you where the data is, uh, right? And, and, uh, we certainly see that and, and why we think cloud is the new home of, I mean, it's not the new home. It is the home of AI because again, lots of customers data is, is being stored in those lakehouse and, and object storage in general, right?
And, and talking about cloud, one of the things that Haiku is really, really good at, uh, is this coverage. And, and we talk about insurance industry and data insurance industry, everybody knows coverage is currency for us, right? How many workloads can you protect?
Um, and I'm using, I mean, each of these cloud providers is now what, offering over a hundred to 200 services, uh, uh, for a variety of things that the, that a customer wants to do. How do you keep up with these growing set of diversification of your data sets? What used to be within four walls of your data center is now just spread everywhere, right?
And how do you protect your data sets in this context? And one of the things ICO is really good at is the coverage matrix. We have protect over 90 plus workloads today, and we already have the broadest coverage, uh, when it comes to cloud in terms of the number of workloads, uh, uh, we protect in the, in the cloud, right?
And so let's talk about how AI, uh, changes the game a bit, uh, in, in the cloud and, and how a lot of our trends kind of carry over in this, in this context. As you know, uh, AI is not just, I mean, delegate clearly knows that AI is not just about chat GPT and, and, and your interactions, uh, with, with the chat platform, right? And, and there is so much that is happening in the backend, uh, when it comes to, um, what an AI effort in, in entails in an organization.
You start with, you know, ingesting training data sets and actually creating training data sets. One of the things that we are learning is there is a huge increase in spike in storage, in, in, in storage around data lake houses, as well as object storage, because everybody's creating one additional copy of all their data sets to kind of feed AI to drive those insights, right? People didn't have a centralized repository of all their data sets in one place.
You know, now AI is kind of pushing you to create that set of training data sets, right? But raw dataset doesn't actually get the job done, or doesn't get, uh, get, uh, doesn't optimize. Uh, for those model training, you do have to enrich the data and make the data look, you know, friendly to AI so that it's easier for AI to ingest.
There's a lot of enrichment, uh, that happens. And then you go through training, it's multiple runs, uh, and at each point you're creating weights, you're grading logs, you, you create a lot of the, uh, tunability that happens in, in, in that case. And then you finally kind of serve that to your users.
And, and, and then there is a full feedback loop in terms of just making that, again, part of what you learn from, uh, in general. The key thing to note is, uh, this is ai, uh, implementation in a lot of organization is a relay race, right? The baton gets passed from one to another to another to actually drive the outcomes you need, uh, uh, in the organization.
And if losing the baton once, you know, in this relay race, you kind of break the traceability in AI or the trust in ai. And, and this is really where the challenge is when it comes to data protection. If you don't follow this relay race and ensure the batan gets passed around from one to the other, and you're fully capturing that, uh, particular race, you know, you're, you're not really in the race in that, in that context, right?
So, um, The quick question on, 'cause there's several, this is Keith Townsend, the advisory bench. There's several challenges around data protection as you hand off from one process to another. And the challenge is not just protecting the data, but the transformation.
So let's, you know, uh, focus on a rag process that ingests the data, however we're getting it, and then we go to enrich the data for ai, whether that's serving up via inference or we're serving it up for training. A lot of the mistakes happen at this level, and the visibility of being able to, uh, trace back the issue and revert back to state is a legitimate challenge. Are you guys helping with that problem or just the ability to recover the underlying dataset?
Yeah, look for each of those service. Like for example, if you are storing your vector databases in, uh, in a I DB as an example in Google Cloud, right? Um, what we do have is the ability to capture that state, kind of bring that back to the state.
One of your model runs, runs amuck, and, and it's, uh, uh, uh, it's, it's not the data set you we want. And, and a lot of filtering happens in, in, again, AI model training as well. You have a data set that is potentially poisoned, right?
In all of those cases, we do allow you to kind of bring back, uh, to a state, um, where your previous models ran from potentially, right? So we make that possible to be able to restore back to an earlier state. But a lot of the times it's also about restoring the data in a format that AI itself can understand, uh, as well.
So it's about being able to bring back the states, but also being able to produce that data back for future analysis to see if there was a problem with it, right? Yeah. So for case in point, like if I have, if I'm enriching the data with JSON so that it's better, uh, injected into my, uh, data pipeline for, uh, my vector database process, what happens two or three revisions down a line, I discovered that a edit that I made to the JSON structure actually broke something else.
And reverting back to that clean state is not simple. Yeah. Yeah.
Schema drift, pipeline corruption, you know, all of these are, are, are legit problems. Um, and, and my, my, uh, the, the good thing about iku is you can go back to an earlier bit, but you can also produce an offline copy of the data you already had so that you can understand what went wrong. So both those elements are, are supported and, and enabled through iku.
So sya, this is, uh, Ray Lui, Silverton Consulting. Do you plug into the checkpointing process during training? I mean, so I mean, data protection historically has been driven on a, a time basis or a change basis or something like that.
I'm just trying to figure out where you play in training checkpointing. Yeah, it's, it's not so much that we plug into checkpoints, but it is a natural point for you to take your backups, right? A checkpoint is usually integrating with your backup application to make sure that, hey, I have this model checkpoint, I've stored all the, you know, sampling weights and, and, and, and all of my metrics around my model implementation that's been stored in the registry.
And that is a perfect time to kick off a backup so that you have, um, you know, all the data up until the checkpoint. So well, so whether you are your models as a new, new run or whether you're taking a checkpoint to make sure that, you know, I have this state captured. So that's when they plug in with backup solution to make sure all these services that make the checkpoint protected as well.
So if I have, and this is, that's not necessarily a bad thing. It's, that's just if I have a custom data pipeline, just as long as you're storing the metadata that I need to do the recovery that I need, and this you're work working with a sophisticated group, this is different than enterprise IT backup. This is usually someone working specifically with data pipeline, and they're going to have specific requirements around how they recover and what, what, what they're recovering within the data.
So it'll be interesting on the kind of the brick level, what I'm able to recover, uh, in this environment versus what I'd be traditionally looking at in haiku's traditional type of customer. So just to jump in here, this is Dave Graham from Men Commons. So what moved to asynchronous checkpointing where you're in the entirety of your workflow is driven, you know, in a different direction where checkpoints aren't just fixed points in time anymore, they're something that is streamed.
How does this obviate some of what you were talking about? Because there is no return to zero anymore. There is a, an ongoing practice.
You're ongoing and you're flowing through these models and through these training epoch in a different sort of way. So you're return to zero, you're return to a known good becomes best, best state y Yeah. It becomes, becomes a little bit more ephemeral than it ever was before, right?
So I think it was Carl that might have brought it up before, but you know, you can act on a checkpoint file, you know, 80 ks plus whatever or whatever dump comes outta that. But how are you integrating above that stack then into the command and the command and control center of that training system or whatever ends being used, whether it be cuda or otherwise Metadata, uh, to control where you, That stuff like that is, is, is this becomes the hydra problem, right? It becomes, there's so many different aspects of things that you would want to control.
'cause a lot of those checkpoint files are just there. It's garbage anyway in the end, right? It's just there for, you know, cover your ass type moments, right?
The metadata ends up becoming the driving point. Yeah. You wanna make sure you maintain that.
So anyway, so yeah. So from a practical problem, I've run into this. So the, I transformed my data, but I need to recover not the entire state, their your, your com your comment of returning to zero.
I'm not trying to return to zero. I I want to save the work that I've done up to this point and, uh, return a portion of my data pipeline in its original or to a different version of the state. I don't know how to, I don't know how to solve the problem if you guys are solving the problem.
That is really interesting. Well, if you have an entire NDL 72 that's chunking away at this stuff, you're gonna return to a known good state. You're gonna now have to pull 72 GPUs back to a known good state gonna recalculate that epoch, right?
Right. That's a big lift if you're gonna try to do that. So consistency wise, again, and it's not, I, I think the idea here is noble.
Absolutely. I just, I'm more curious to like, how do you now embed yourself into a process that is literally changing week on week, uh, as we, as we discuss these technologies? Yeah.
And, and, and what's really, uh, cool to hear is, I mean, what you are talking about in terms of, hey, being able to know go, go back to last good state and, and so on. I mean, this is what backup windows have been doing for years for any data set, right? Any data set that we protect is always changing.
Any data set that we recover is because we need to go back to a known good spot, uh, in majority of the cases, whether it's because of disaster, whether it's because of cyber attacks, whatever the reason may be, we go back. In this case it's because of experimentation, right? So at the end of the day, it's, it we allow, we capture state and we allow you to go back the state.
And when, when it comes to capturing state, we'll also talk about how you capture that state in a more consistent fashion. And that's also things that, you know, we're, we're investing in and, and, and actually, you know, build some patent painting technology around how do you drive consistency among the data sets that you're actually protecting? That when you come back, it's not going to look at the dataset that just came back as, Hey, what is this that I don't recognize?
Right? You need to be able to put back and expect things to run, uh, normally, and there is a lot of work to be done. We're not saying that we've solved every problem there, but we have solved some problems that are, you know, very key.
And, and we'll walk you through those, uh, uh, consistency element as well, uh, when we, when we go through the solution in more detail, right? So again, the point is lots of living data sets. Yeah, yeah.
So lots of living data sets being produced, and again, those data sets need to be protected. Uh, you drop the baton once, you know, you kind of lose a lot of the traceability and, and trust, you know, that AI generates, you don't want to be producing results later saying, Hey, we don't know how, why my models behave the way they behave. Uh, you need to be able to again, trace back and say, what, what actually happened?
Uh, to an earlier question that came up. You know, let's take the example of, uh, uh, a customer building an AI pipeline in, uh, in, in Google Cloud, right? You start with a workbench, um, that act that has, you know, direct integration into GitHub and, and BigQuery and so on.
You deal with your training data sets, you create your feature, uh, data sets where your digital wins are or stored. You create prediction data sets with both what you send inside and, and what it comes out of it. And you generate a lot of model artifacts in, in, you know, verex AI pipelines, right?
So you, what you see is dataset is frauded across multiple services. People use the best of breed solutions for, you know, each of this dataset. Yes, you are gonna see BigQuery, which is the lakehouse platform that Google has.
Uh, right? You see this in multiple columns. And, and to an earlier comment, yes, lakehouse are becoming that unified, central platform, and we're really happy about it, but you're also seeing that it's not the only platform people use.
At the end of the day, BigQuery can't be your workbench. We got it can't be where your store, your source codes, and it is not the best platform for, you know, real time in insights or, or, or feature data in, in, in general. So there is a lot of cloud elements that customers use to be able to create an AI pipeline.
And what we are doing is to ensure protection across all of these services, uh, uh, with Haiku so that more and more of your AI infrastructure is covered, protected, and you have confidence in, in those capabilities, right? So, and that's why this graph where we talk about, uh, all of our, uh, uh, coverage matters because you want to be able to capture everything from VMs, which could be your workbench, all the way to vector embeds, right? So all of that data needs to be protected and why, you know, our coverage really matters, uh, in the, in this context, right?
And let's talk, uh, Lakehouse, right? Um, and, and again, glad that group is already tuned in and, and, and, and kind of see this change happening, uh, like we see it, uh, as well, Lakehouse, um, I like to call this, they're kind of bringing together AI BI and what I call ci, right? The, the continuous insight for an organization, right?
So the late houses are becoming the new system of record, and it is growing at a phenomenal pace in in, in, in customer organization. We've worked with customers, you know, who said they have near a hundred percent growth year over year on some of their lakehouse data sets. Um, and, and these are not a hundred percent growth when they are 70 gb, they're a hundred percent growth when they're at 700 terabytes, right?
They're talking about, uh, uh, massive growth even at scale. Um, and, and, and why you're seeing Lakehouse platforms like BigQuery and, and Databricks and Snowflake, each multi-billion dollar businesses, they're all around, you know, three to $5 billion businesses on their own, but each of them growing at nearly 50% year over year. Um, they're creating, uh, uh, a nearly a $2 billion business every year and kinda adding to their, uh, uh, balance sheet at this point.
And that's because again, customers are adopting lake houses, that single unified platform where they can store all this, uh, uh, dataset, right? And, and as we talked about, that creates gravity when more and more dataset lands in, in, in a specific platform, you know, you start building solutions that protect that particular platform, that deals well with that particular platform and, and integrates first party integration into some of these lakehouse architectures with a lot of solutions that, you know, are bit around these data sets, right? Sanya, can, Can I interrupt and ask you a question on this?
Yeah, please Do. So from your customer interactions, Scott Roon was al, um, you know, it seems like there's a spread of very mature customers who are really disciplined and have gravity in one place, but a very long tail and continuum of customers that don't and maybe will never need to. Where do you work best on that spectrum?
Because like, when you say lake house, it's like, well, I have a place at the lake, but I also have a place at the beach, and I've got a place in the mountains where there's streams, not lakes or oceans. And I think that's more a real representation of where many enterprises are today. Yeah.
And, and, and, and you're truth, the, the, the workload, the, the customer environments are often, you know, on both ends of the spectrum, right? You have a very long tail, uh, in, in, in, in, in this curve. You have a whole bunch of customers who have datasets that are kind of spread everywhere, and you've got a lot of customers who are also going through consolidation of all that dataset into single platform, so they could drive outcomes, their AI outcomes a lot faster, right?
And, and, and the good thing about our solution is that we were already built for the customers that had data everywhere, okay? Right. Uh, we had 90 plus workloads already supported, um, right?
Whether you keep your CRM data in Salesforce and GitHub data, uh, your source codes in, in, in GitHub and, uh, your databases in, in cloud sql, it didn't really matter. We provided support for all of those workloads individually through a platform of, uh, uh, through a platform. What we are seeing now is not just a group of customers that have data everywhere, but also a group of customers who are consolidating and creating these massive data sets that are feeding into ai.
And that's the journey we're seeing, and we wanna be there, uh, to, to support those customers who are doing that, uh, on this front as well, because these are workloads that have not historically been talked about in the context of data protection, right? Backup windows didn't talk about BigQuery, backup windows, didn't talk about Databricks, um, and still don't. Um, and, and, and we're seeing that shift and, and, and are adapting our solutions to kind of become the best in breed for those workloads as well.
Okay? Okay. Um, in, in, in, and the thing that I also wanna highlight is, you know, th this data set is yes, it's special, it can handle scale, uh, um, but there is also potential for, you know, data loss in this context.
Uh, we talked about a couple of scenarios already. I think Keith mentioned, Hey, I could have a corrupted JSON file completely screwing up my, my mar my models. And, and that's certainly possible.
Schema drift is, is a consistent problem where models break. Um, you can also have poison data sets. I mean, we, uh, the first thing I hear from cloud vendors when they talk about these, you know, BigQuery and Databricks and so on, they say, Hey, there is no, we don't, don't give you access to a compute.
Somebody can't just come in and run, uh, uh, uh, uh, or corrupt your dataset there. But what you do give is access to APIs. And with those APIs, you could poison datasets that your models are being trained on.
And that could be the face of AI in this model, but whatever is the issue, we do need to understand that any data, any service has chances of going wrong, you could lose them. And you wanna have the ability to go back and protect that. And, and, and that's one of the key elements, uh, to really think through.
Again, our job is to kind of follow where the data is going. We're seeing that the data is going to these lakehouse and object storage predominantly in the context of ai, and we wanna make sure that we have the best in class solution. So you could get back those data sets if you ever need for whatever reason, uh, uh, that is just Saying, you should update that last to include Amazon, because US East going down is kinda legendary at this point.
Might as well make, make, make a note of that. Yeah. And, and, and, and, and look, I, I think, uh, pointing that out is a little bit of like ambulance chasing as well.
Uh, but, but in this context and, and, and, and, and absolutely. Um, and, and we'll talk about why that problem also is permeated across all these enterprises is because, you know, AI is the cloud is the home of ai, but, you know, cloud is also puts a lot of people under house arrest, uh, right? And, and people don't really have a way to create that cross cloud resilience, uh, across these data sets.
So yes, it's your home, but it's, you're also under arrest to, to some extent. And, and, and that again, needs to change. And, uh, you'll see that we have some solutions that helps provide that cross-cloud resilience and, and shareability of dataset from one cloud provider to another, um, as well.
This is just to show that, um, data recreation, if you lost your dataset, is very expensive. Yes, it's possible, and it's possible in some cases, I'll explain the scenario where it's not possible. But generally when it's possible, if you add infinite time, infinite resources in infinite APIs, of course you can build a lot of that dataset back.
Very often what sits in Lakehouse is your second copy of a set. It's already sitting, uh, uh, someplace else. And you need to transform that and store that in a lakehouse before you feed it to ai.
So one of the reasons people think, don't think about backups immediately, but what we are also seeing, uh, is these lake houses are really good at streaming data, right? And, and customers are feeding, uh, whether it's for sentimental analysis based on Twitter messages, right? Or, or any social media messages that that may be, or ai, I mean, IOT data sets and sensor data sets that are coming in from multiple places.
Those data sets are getting streamed directly into these lakehouse. So it's your only copy of dataset there, and if you lose them, you know, the sensor is not gonna regenerate those, those, those information for you. And Twitter is not going to give you back all of those old tweets without charging you for, you know, a spike in API usage, right?
I think at this point, what they charge, uh, almost 500 KA year, uh, to search through and, and collect about 50,000 tweets. And, and that's the kind of plan you're gonna be paying for in, in, in, in Salesforce, uh, as well. I, so what we deliver, uh, to this ecosystem, we talked about why it's important.
We talked about what it would cost to bring back the data set if you didn't have a reliable mechanism. Um, and, and now I want to kind of talk about what we do, uh, uh, uh, for this dataset, right? And, and the key thing to highlight is we don't want to be just the FedEx that takes data from point A and store it in, in, in a different spot, right?
We're not just trying to move data blindly from point A point B. Um, one of the things that we want to do when you protect AI is not just protect the raw data behind, uh, that, but also protect the meaning, protect the intelligence that feeds into, uh, uh, uh, the, these pipelines. So when, for example, you protect, uh, a table, a raw data set that you're trying to protect, what Haiku does is that it has the mechanism to capture not just the raw data sets, but also the views, the models, the routines, the access policies, uh, the schema.
Uh, so we are protecting all of that dataset and not just your raw tables, uh, that are coming in. Again, a lot of the times you can get the raw table out and you get the raw table in, but if you don't know how that raw table is processed and, and then you don't have good visibility into that dataset, why I think we're extending beyond just protecting raw dataset, but also everything that kind of sits around it, the views, the, the models and the routines and so on. And there was a question on a consistency, uh, right in, in workloads like data lake houses, you know, take BigQuery as an example, right?
People store data sets that data. It could be a couple of petabytes of dataset, but they're being stored in smaller and smaller chunks. There is partitioning and people store with different data sets.
What that means is when you're taking a backup of a data set, um, and you kick off, uh, an export or a backup, you may get to dataset one at 11:00 AM in the morning, and you get to dataset 35 at 1230 in the afternoon, and you get to dataset 200 at 5:00 PM in, in the evening, right? So you are actually protecting these datasets are very different types. So one at 11, one at 1 30, 1 at, uh, uh, 5:00 PM then how do you drive consistency?
How do we make sure they're all saying the same thing? They're all referring to the same state. What we do is deliver automat in this context.
We line up these data sets as look like at a certain point in time. A lot of these solutions offer what is called time travel that allows you to go back to any point in time over the last seven days. We actually tap into that and leverage that to say, Hey, line up all of your data sets to make sure that it look like from a common reference point at say 11:00 AM in the morning, and then back up all of those dataset as it looked, um, at, at 11:00 AM.
So even though we backed it up and we moved those data sets at 5:00 PM they all look like how it was at the same time together. And how we drive consistency is something that we have a patent pending on. And, and, and, uh, uh, what we are driving with these data sets, uh, in, in this context.
Hey, Satya, this, and This is Ray ese, Silverton Consulting. Uh, don't, most of these lake houses have internal backup procedures or, uh, processes, you just turn on a flag that they automatically back up your data for you and you don't know necessarily need to do anything else. I mean, and all these guys is based on the cloud, which also has backup capabilities built in, uh, the underlying storage cap.
So I'm just trying to understand the need for, you know, a solution like Haiku to come on top of these things. Yeah. Uh, uh, Ray, what what you say is really what a lot of people believe, but not what is right, uh, right in the sense that yes, uh, if you take BigQuery as an example, uh, yes, it comes with time travel that you can configure, um, and you have a anywhere from two days to seven days where it could roll back to a point in time.
It's very fast. It's, it's, it's awesome. It's the technology that, you know, every time I work with, it puts a smile on your face, right?
Um, but your detection window in this context is seven days. If you find a problem within seven days, yes, you can go back to it almost instantaneously, and the solution makes, makes it possible. But tell me, in the world of data protection, how many workloads have a seven day data retention?
Right? Um, very few. You go to any company, ask them, what is your corporate policy for backups?
It's 30 days, 60 days, 90 days. And some of them will say even three years and five years, depending on whether they're a regulated and not, right? If you're subject to N-Y-D-F-S, you know, you're going to be stewarding the data for a lot longer, period.
If you're subject to a doa, you wanna make sure that that copy of dataset that has customer data is at least stored in another location that is completely independent, right? These are all regulations that are in place for good reasons. And when you have all of your dataset in one place, and you're leveraging co-located copies for data protection, when someone gets access to that IAM boundary, you basically can get rid of both your primary copy and your six three copy.
And your detection window for recovering from that event is, uh, is measured in days and not in, don't Get me wrong, I understand the need for independent copies of data outside of the, the co-located solutions that are available. Uh, but I thought even that was something that, uh, many cloud providers supply. Yeah.
So when a cloud provider, none of the cloud providers, especially for this type of workload, you know, I've spoken with a couple of these, uh, uh, product teams, they don't want to be in the backup business, right? But if you wanna dump a copy of it sometime once, you know, they will allow you to, and there are APIs to, to do that, but you are talking about a two petabyte data set with no way to generate an incremental, uh, uh, storage or a copy, right? And, and, we'll, I'll actually show that in example, in the backend, what the cloud vendors offer is a way to keep a copy of it someplace.
But they don't, they don't, like, for example, if you're immutability turned on on those buckets, you cannot actually back up to those buckets per day. You have an example where they're, again, only supporting raw data to be moved, not protecting your ML workloads, your access policies and so on, right? So what they offer is very primitive, uh, and often still locks you into the same platform.
Uh, right? You can only export my warehouse data to my object storage because I've got a petabyte network in between them. So I will only let the hop go into my own storage.
You, I will not let you export into one other, uh, cloud bucket. And those are all things that, you know, we can help address, uh, for a customer. I gotcha.
Thanks. So, and, and, and that kind of leads to the next section, which is to say the protection of this datasets, when you're talking two petabytes and 20 petabytes and 10 petabytes type dataset, you know, protection itself becomes cost prohibitive. It is no longer, uh, uh, a thing that kinda hides behind the budget.
If you are creating multiple copies of your datasets and they are petabyte and scale, you know, uh, backup sometimes becomes cost prohibitive and why customers kind of rung with it with, even though they understand shared responsibility model, oftentimes they realize, well, it's just cost prohibitive for me to do it, and I'm measured on my cloud spend, so I can't spend money on this, right? And that is a real thing we see in, in, in some of our customers. And, and what you see with the customer is, Hey, I want an offline copy and these are large data sets, so I wanna be able to store it kind of incrementally.
I don't want cyber resilience exposure. Uh, I don't wanna land up in New York Times or Wall Street Journal, uh, because I didn't even do a backup. Like, I didn't even do basic backup of my, my data set.
Um, it's not like you can't foresee cyber attacks anymore. I mean, it's the new reality. Everybody's gotta design a solution around the possibility of being a di at some point.
But there are limitations that exist in these platforms for these customers. You do have to make sure that you can actually egress out of where you are today if you want that cross cloud resilience. Okay?
So take our example and a real world example, you're generating a hundred terabytes a day. Being able to move that data as a full base copy every day, which is what these cloud providers offer you every day, and pushing it into an object storage. You know, you take 30 days, 60 days, 90 days, you could see this, a hundred terabytes quickly become petabytes and a big, uh, line item in your budget, right?
And they still can't back up to an immutable bucket. Recovery is slow. You're only getting a raw table and not everything that sits around those raw tables, those problems still remain, it's still freaking expensive, right?
So where we come into the picture, and it's also a partnership that we have in, in the ecosystem is Haiku works with, um, what people used to take for granted deduplication, right? In, in on-prem world. Uh, people took compression, deduplication, and all of these things as granted because it was just available.
You know, almost every organization had some form of data domain installed irrespective what backup application you had. And a lot of backup applications also offered, uh, you know, on-prem appliances that delivered deduplication. We kind of take that and we partnered with, uh, uh, Dell in this context to drive deduplication in the cloud.
So we work with data domain, virtual appliance. They are the number one storage target for backups. We work with them and we leverage their DD boost protocol to say, take these lakehouse datasets, take the a hundred terabytes of data that they're sending us, leverage DD boost protocol to actually only ship the changes that the system things it needs.
'cause a lot of these are upend only dataset. So if you're creating a new base copy every day, you don't really need to store all that much data. You only need to store what changed between yesterday and today.
And so that deduplication is something that we are offering, and we're actually seeing customers get up to 40 to one space savings, because again, you keep getting the same data again and again, uh, from cloud providers, but when we store it, we store only the fragments that have changed in formats that are very deduplication friendly, like parquet formats. And that allows us to deliver both mobility for this dataset, uh, and, and real protection for these data sets. Um, and if you also take into account that we can replicate this dataset from one cloud provider to another cloud provider post deduplication, you're not only going to save T to one in terms of your backup storage, you're also going to save t to one when it comes to egress charges.
So you can actually store some of your golden datasets, highly curated dataset from one cloud provider to another cloud provider, because again, you can have a cloud provider failure. Um, like someone mentioned earlier, there was a recent incident, there's also an incident where $150 billion, uh, uh, uh, company in, in Australia was literally wipe, uh, from cloud infrastructure. That account was deleted and their data was, was irretrievable.
Um, this is a not a small moment of shop. This is $150 billion, uh, uh, business in in Australia that just woke up one day and found their entire cloud account was lost. So how did they get back this business?
So Satya, was there a Question? Uh, yeah, from, this is Keith Towson, the advisor bench, from a practical perspective, one of the challenges that, uh, as I was hearing you, you're, you're answering a question I'm glad I didn't ask earlier, which is the data gravity problem that guy was talking about. So in this, this isn't workload management, I can't magically move, you know, my, my, uh, large data lake from GCP into AWS because I wanna use bad rock against it before this.
Now, if you're saying I can get 40 to one dupe, and now I can move my data out of, uh, aw, out of Google into, uh, AWS, this is the bad to use bedrock against it because, uh, I have some business process in which I wanna do that, that you're saying that Haiku helps me with that data gravity problem because of the DDU capability, Correct. But the, it, it does help you from data gravity problem, and it does help you from getting that cross cloud resilience. You don't have to feel like you're under host arrest.
But, but, but, and, But in those cases, if you wanna move the data from Google to a WSE duplicated, you have to have the data of the main clients in both locations. Correct. And data domain is essentially running on top of, uh, object storage in each platforms.
It is capable of running in all three hyperscalers. So it is replicating and reading reduplicated data that is then served back wherever you need them to be. Yeah, I'm sure that's still gonna be cheaper than, uh, moving had aby of data, data dumping is gonna be cheaper than doing.
Yeah. And, and, and especially considering that egress charges from one cloud to another cloud provider is five times the cost of storage. So you really think about in terms of dollar value, you're getting 40 to one in storage.
You're actually getting 200 to one, uh, in, in savings when it comes to egress. If you are trying to move from point A to point B, I'm not saying every customer is doing that today, but what this integration makes is it makes it possible for you to not store all your dataset in one cloud provider and hope and pray that they've got everything together. Um, right.
And, and this will be something that you drive resilience in another place. It also provides cross. You can use the best of tools available in either cloud platforms on the data sets that you would really care most about.
Right? Um, so, uh, that kind of wraps up at least my section before I hand it over to, to, uh, uh, so, right. The idea here is, is are real challenges in the cloud.
I mean, there were already problems in the cloud, but in the world of ai, it gets exacerbated. And, and what we have built with our solution is to really centralize the management of all of these services that people deploy for AI and deliver protection to one place, um, and actually provide coverage for a lot of workloads that act as blind spots in, in, in the industry today. Provide backup coverage for DBAs workload, provide backup coverage for data lakes and Lakehouse, and do it in a fashion where it's efficient and lean so that you can actually deliver that kind of resilience.
We, we've talked about. So hopefully, again, a lot of AI work going on. We're super excited to become the, the home of your AI resilience and, and really looking forward to working with, uh, our customers, seeing where the industry is going and adopt our backup in, in cloud technologies to, um, what a, where AI is taking us.
And that said, we, we talked about VMs to vector embeddings, like we need to be able to protect everything from VMs. This is where we all started from to now thinking about where vector embeddings and, and so they are. Once we'll walk through, uh, some of our key capabilities around vector databases and what we are doing there as well, and as well as the SaaS workloads and, and how, uh, we are, we are adapting to the age of SaaS as well.
So I do have a question for you. Uh, this is, uh, Fred Van Hern hyphens consulting. So how fast does the DUP deduplication work?
So for example, if you have a new batch of data, at what point or how fast can you figure out if the data can be duplicated or not? Yeah, there are two elements to, uh, the way we implement our data data protection. Um, one is when we transport the data itself, um, we, we diversify where the deduplication happens, right?
So when we transport the data itself, we match checksums to make sure that we don't send the same data again and again. Um, so a lot of network acceleration, uh, mechanisms that you've probably seen in the past, very similar. We use checksumming to make sure we don't even send the data that you, that needs, that we know is already there, right?
On top of that data domain does its own, um, uh, uh, process to make sure that, you know, hey, if I find unique shots, then I keep just those unique shots and not keep any more of those data sets. It has got the intelligence to compress and, and duplicate even further, but it only does that on the 5% of data that we think is really unique, um, and not on the a hundred percent of data that we need to send and then make it to the work, right? So the combination of the DD boost logic, which diversifies where deduplication happens, but also a combination of what a data domain does really, really well, which is really shrinking it down to, uh, just the bits you need.