Protecting the heartbeat of ML/AI use cases with HYCU
HYCU’s presentation at Cloud Field Day focused on the critical need for data protection within the rapidly expanding AI/ML landscape. The increasing adoption of AI mandates across organizations necessitates robust protection for the underlying data lakes and lake houses that fuel these systems, as well as the repositories AI creates. The presentation highlighted the broad coverage HYCU provides for this data stack, emphasizing its innovative solutions for Google BigQuery, a major data framework used in production environments.
A key aspect of the presentation centered on the various reasons why protecting AI data is essential beyond simply recreating results. Speakers discussed the importance of cyber resilience, the ability to revert to specific points in time to address performance issues or model drift (re-vectoring), and the crucial role of data protection in meeting legal and compliance requirements, such as demonstrating the absence of PII or IP infringement in training data. Furthermore, the complexity of reconstructing datasets spread across diverse sources (on-premises and cloud) was underscored as a significant challenge requiring a comprehensive data protection strategy.
The presentation showcased HYCU’s capabilities in addressing these challenges, specifically demonstrating its solutions for BigQuery. HYCU’s platform provides automated discovery and protection for a wide range of Google Cloud services and boasts a patent-pending technology enabling atomic backups. This innovation addresses the critical issue of data synchronization across multiple tables and datasets within a data lake house, ensuring consistency during backups and recovery. The discussion also highlighted the increasing reliance on data lake houses as central repositories for AI-related data, emphasizing the importance of robust protection for these often singular copies of crucial datasets.
Presented by David Noy, Vice President of Product Management, Dell DPS and Sathya Sankaran, Head of Cloud Products, HYCU. Recorded live in Santa Clara, California on February 20, 2025 as part of Cloud Field Day 22. Watch the entire presentation at https://techfieldday.com/appearance/fortinet-presents-at-cloud-field-day-22/, https://techfieldday.com/event/cfd22/ or visit https://www.hycu.com/ for more information.
Transcript
All of our customers are going on embarking on AI and ML journey. Pretty much every customer you talk to should be doing, if not, uh, for folks who do not know me, this is Aya, AYA syndrome. I run products for Haiku.
I want introduce the section and I wanna pass it on to the speakers. David, now, and Satya, every customer, there is a stat, you guys must have read about it. 80% of the company enterprises are supposed to actually have or expected to have an AI based solution before end of the year.
Why is that actually important? If they're gonna come up with a production service, it means they gotta keep the data safe. That's a requirement.
It's not a nice to have, it's a must up because now you're running a production in production service. That's where it comes down to it. And people talk about ai, most people think of gen ai.
They just think of, oh, the chat GPTs of the world. There's a lot more to it than that. You guys know that it's not, people don't always think of.
That's what I want to invite our friend, David, now to first give us, set the context on what does he see as some of the needs for customers protecting the data. And then Satya will walk through some of the things we can do. David, Thanks.
Thank you. Alright. Uh, this one.
Thanks. Um, yeah. So look, uh, David, no again, VP of products for Dell for the data protection portfolio.
Dell is the industry leader in providing AI servers to customers. Um, and you know, what Aya just mentioned is the importance of protecting ai. Look, when we talk about AI and data protection, there's all kinds of cute things.
You can discuss chat bots that help you figure out what commands to run. Um, people try to do analytics on your backup data. I'm curious how that's gonna come out when you have multiple versions of a document and which answer you're gonna get.
But that aside, um, at the end of the day, yeah, when we were thinking about the AI factory at Dell, uh, you know, you've heard this announced Dell has the AI factories that we sell. Um, we asked the question of, well, why do people protect AI data? Because if you ask, do you protect AI data, one of the common answers you get back is, well, you don't need to protect AI data 'cause you can just recreate the answers.
Well, it's kind of true, but kind of not true. So you put yourself in the mind of a large enterprise. Think of a top 10 bank.
Think of a, uh, any regulated industry. And there's a couple of use cases that pop up. Number one is cyber resilience.
So if I take the, the training data or I take the results of inference and I poison them, um, in some way, shape or form, I can actually modify your models. I can modify your, uh, training data. I can modify your results.
So look, cyber attacks happen. They're getting more sophisticated. The ability to go and recover back to a specific point in time is critical.
That's, that's motherhood and apple pie stuff. But let's talk about some other interesting reasons. Workload state.
So let's say that you start to get performance, um, variations and, and drift in the way that your models are behaving. And you wanna get back to a specific point in time and you wanna re vector potentially while re vectoring requires you to go back to a specific point in time. And if that means that you have to know what the state was, including all the configuration parameters that were used in the time that vectors, uh, vector database was created, and the models were basically trained.
This is another way to go back, another interesting one. Put yourself in the minds of a large enterprise legal protection. How do I prove that the training data that I use to train a model didn't have person identifiable information or some sort of IP infringement because I'm, I'm, I'm using someone else's intellectual property in my training models.
The way to prove it is to back it up. If you wanna do that cost effectively over time, data protection solutions like the Dell data domain plus backup vendors like haiku make a lot of sense. Actually, I'll also gonna start to believe that, um, compliance is gonna start to kick in.
So the data that you use, oftentimes there's compliance requirements to keep data from multiple years, seven years. How do you do that cost effectively? And not just the data that came that was used for training data that came out of inference, not just the in, you know, the inference results, but also the prompts that were put into a system and the results that came back from prompts.
So in some large financial institutions, every text message, every message that goes across the wire in any way, shape or form, is actually kept and retained for compliance reasons. So you can go back and say, Hey, when I had this conversation, here's what we discussed. That's an important reason as well.
And then dataset reconstruction. So one of the interesting things, and again, we're at a cloud field days, if I'm going and I'm pulling training data, I'm potentially pulling it from multiple disparate sources. Might be O 365, might be, um, you know, some on-prem NASA environments might, whatever you name it, multiple workloads, vector databases are being produced.
You have databases like MongoDB, Postgres, elastic searches used in, in keeping track of where things are. The, the whole in the processing of the training job. All of these different data sources, whether they're training data sources or config sources or, or, or databases that are used in the process of creating a model, potentially need to be reconstructed in the case of catastrophic data loss.
And so how do you go back and make sure that you have a consistent point in time across multiple disparate data sources where they weren't all residing on one piece of infrastructure? And that's the answer. So, um, I just wanted to give you some of the context of why it's important.
Now, let's go into, uh, a real life example. So this example is a large, uh, retail grocery chain. And in this particular case, we're talking about data sets that existed in object storage, on-premise, object storage, off-premise, um, large database as a service provided by one of the cloud service providers.
We're talking about multiple tens of petabytes of data across different cloud providers as well as on-prem. And so, again, what the value of a Haiku Dell relationship here is that Haiku provides the knowledge of those cloud database as a service vendors and knows how to read those object storage, whether it's on-prem or in the cloud. And basically take what is essentially a combination of, we saw here, uh, procurement, manufacturing, pharmacy, inventory, all kinds of different, uh, documents, images, um, config files, you name it.
And basically be able to pull that together to use for, in this particular case, it's used for an AI use case, but you saw the reasons why you would wanna do that on the prior page. And then we protect that and store that. And again, Matt massively reduce it.
So 15 petabytes on-prem, we saw multiple petabytes going o uh, off-prem as well. And, um, at the end of the day, we wanted to make sure that that's really resilient. This is a customer who's been a long time Dell customer, so they understand the value of the data and main product, but they need to be able to get access to all of those workloads that go into that ai, uh, trading environment.
Okay. Makes sense. Alright, Thank you David.
Thank you. Actually, I do wanna just, what's interesting is looking at like the, I literally just to check, I scanned the NIST framework in the RMF playbook for ai. And it's wild that in one side there's literally two references to the words data protection.
There's nothing in about backup in there, but you highlighted something that's significant, like the re vectorizing of, of a database is, it's brutal. And a lot of people are discovering this the hard way. They accidentally leak bad data into the training and then they gotta go back too far, or it's extremely expensive to go in.
So they're literally weighing the risk of maybe it'll just kind of disappear inside. So I like to see this, and I think that's a really strong story that we may not have seen yet because people haven't hit that wall, but they're gonna hit that wall and there's nothing out there that's doing that. I just want to add one quick thing to that, which is the NIST framework, uh, references it, but you're right.
Doesn't enforce it. The DORA framework in Europe absolutely enforces it. So it's very interesting to see kind of what's happening in Europe with data protection and how this is all gonna play out there as well.
Yeah, yeah. I think AI has been moving so fast and, and we've been looking as a tech industry. We've been looking at like, how do we build bigger GPUs and faster GPUs and put more bus speed into the, into the, you know, onto the motherboard.
And, uh, I'll, I'll, I'll give you a little insider baseball, just share with something with you. Like when, at last Delt Tech's world, when we went up and announced AI factory, they built a huge slide and they put servers and they put storage. They didn't put data protection.
And I went and looked at it and I said, well, that, that's cool. But you have to understand the people who are doing this, these are gonna be large enterprises with these kind of concerns. Like at the end of the day, they will get regulated by Dora.
At the end of the day, they will have compliance requirements. And so we have to make sure that we think about things holistically. We, we have the portfolio to do it.
And so once we start to have those conversations, boom, the light bulbs come. And, and so people are really thinking about this stuff. Uh, again, the goal goal here is to kind of walk you through, um, you know, what we see, uh, with, with AI workloads, um, kinda identify what a modern AI stack looks like.
Where is the data actually getting generated and whether we are following it and what kind of trends we see, um, with this, uh, ecosystem. Um, and, and we'll also do a quick demo if, if time permits, uh, as well, right? So a typical AI stack, um, I think all of them try to do five things, uh, uh, really in terms of stages, right?
Uh, they ingest, they ingest a lot of data, and that drives a set of trends, uh, in this ecosystem. Um, and, and they ingest from files, they ingest from objects, databases, uh, you know, you've got data lakes and, and also realtime streams. Uh, and then you enrich.
This is where a lot of the vector DB vector, uh, embeddings come into play. You, you take raw data, process it, transform it, uh, you create features out of it, uh, so that you have an optimal path to training your, uh, data sets. And finally, there is training.
And, and the training itself is a multi-day effort. Uh, it's not one that happens overnight. Uh, you throw data at it, you run your model, and then yeah, you've got something to work with.
It doesn't happen that way. Uh, people train people, tune people, experiment. Um, and, and, uh, as, uh, the example you already used, there may be times where, uh, you run against a benchmark, uh, you, you check explainability and you later find out, uh, even if all of those KPIs are met, you've actually exposed some sensitive information, uh, to that module.
So you can have to go back and do a whole lot of this, uh, all over again. So training is a multi-day effort. Uh, in fact, some of the largest language models, they all kind of talk about 90 day to 120 day, uh, training window, uh, even though you're throwing, you know, uh, millions of, uh, GPUs or a hundred thousand plus GPUs in, in, in some of these cases, uh, at your, uh, workload, right?
And then once you've built that model, uh, you serve and you deploy that module somewhere, and that comes through our, uh, artifacts, you know, container images, um, and you deploy them and you've got an observ architecture around it. And finally, you use that model to serve predictions. That's really what, uh, all of this AI ML is, is doing, right?
You then give new data sets and ask questions of it, and then the model gives you answers. And all of this data is, is important data that, uh, an organization is looking to protect. So let's take an example of a Google customer, right?
And I chose Google because they've probably got the most, uh, you know, uh, modern data stack, uh, for ai, right? That's, they're probably the most matured among the three cloud providers when it comes to, uh, the AI ecosystem. And, and where does the customer put all of this data?
Uh, the training data, uh, comes from all over the place. Um, they can choose to, uh, you know, go through a unified platform like BigQuery. Um, feature and vector data typically tends to be, uh, sitting in databases like all IDB or, or fire store BigQuery.
Uh, you've got prediction input data that also, again, sits on databases typically. And the output data is either a stream, or, again, if you're a web application, uh, uh, sits in a database or can actually sit in, uh, GCS as well. You've got model artifacts where there's a registry where you can store the model artifacts.
There's metadata about the models that you can store. And there's also workbench, which is effectively a GCE instance that, uh, uh, is synchronized with, you know, BigQuery that's, uh, got the ability to store code in, in GitHub and has, uh, direct connectivity to GCS as well, right? So there is a lot of data being generated at every stage of this AI life cycle.
And, and what we do, uh, with Haiku is ensure that all of these services are protected. Um, right? So all these services are listed here.
Uh, we can actually auto discover and protect all of these services, uh, through, um, uh, our system. And again, uh, it's not just on Google. We actually support, uh, the largest number of services in both AWS as well as, uh, uh, Google Cloud, uh, today.
And, and, uh, the other trend I wanna highlight here, uh, is the emergence of Data Lake House. Uh, in the example that David walked through the major grocery retail chain, um, people are seeing value of a unifying platform where you can throw a lot of structured data and unstructured data and multimodal data and, and, and be able to point your ai uh, uh, training engine on top of that dataset, right? So you need some place to store everything.
And, and that's turning out to be, uh, data lake, uh, Lakehouse, right? And, and what is a Lakehouse? Uh, we historically add something called Data lakes, where, which is optimized for storage efficiency and the flexibility of being able to store lots of different types of data, uh, and it can store at scale, right?
But our data warehouse, uh, which, you know, you've heard of, uh, BigQuery, snowflake, Databricks, uh, you know, Redshift, all of these, uh, uh, uh, solutions, their, their warehouse, they basically provide a level of governance. They provide database like capabilities on top of, you know, effectively data lakes, and that you can actually query and analyze data, uh, uh, very quickly. What Lakehouse do is actually build a system that can handle both store at scale, but also analyze at scale.
And, and that's what, uh, these lakehouse do. And BigQuery is certainly one of them. Databricks actually pioneered, uh, this concept.
And, and BigQuery is one of the largest lakehouse out there, uh, in the market. And again, in the example that, uh, uh, the major grocery retailer, uh, they actually store 85% of the data that they're looking to protect in the stack in BigQuery, right? Because they're storing all of the training data there.
They're using BigQuery for feature data, they're using BigQuery for protection data. Um, so it is becoming a single unified platform, uh, that is storing a lot of data, uh, related to, uh, uh, AI and AI stack. Uh, it is a multi-billion dollar ecosystem.
Uh, Google doesn't disclose revenue, uh, for each of their services, but it is estimated to be between three to $4 billion a RR, uh, uh, ecosystem for, uh, Google Cloud and, and, uh, a significant portion of their, uh, data portfolio, uh, today. So it's not just one retailer using it, it is the, it is becoming the centerpiece of the data strategy for, uh, Google Cloud, right? And, and what I wanna also highlight is, yes, this is emerging.
People are throwing a lot of data at it, but your exposure hasn't quite changed. Uh, right? You still carry the same kind of exposure you carried when you were storing this data in multiple places, right?
You still have the possibility of a deleted table. You still have the possibility of a Terraform query. There's a schema change, and you just end up going and deleting a whole bunch of tables and, and, and data sets.
And these things do happen. Um, and, and there are some inbuilt capabilities with some of these, uh, uh, solutions where you can kind of go back to a seven day window and, and change some of these things. But what if it's greater than seven day window?
You do need backup, right? Um, and the other myth around surrounding this type of workload is people see lakehouse and warehouses as a second copy, uh, uh, typically, and, and why that happens is because usually you're ingesting from another source, right? Again, lakehouse is, are challenging the norm.
Uh, you're not just storing a copy of a database you already have. You're not just storing a copy of a file that you already have in Office 365 that is coming into BigQuery before it's being fed, uh, into, uh, an AI ML model. Uh, we are actually seeing a lot of realtime streaming, uh, data sets being generated as well.
All your IO OT data sets, um, all of your sentiment analysis on social media, right? All of that is realtime stream. And that data is actually skipping any other data store, and they're landing straight into these lakehouse.
So for some of these data sets, your only copy of data is in these lakehouse. And if you don't protect it, um, you don't actually have a way of getting those data back, right? So, and also if you, even if you do have another copy of dataset, the cost of recreating that dataset, uh, in whether it's a big query or any of these lakehouse, is extremely high.
You have to worry about ingestion costs, you have to worry about processing costs, you have to worry about ETL costs. You have to worry about egress and co-location costs because not all of your data is sitting in one location. So recreating this data set, uh, at will is both expensive, uh, and, and also takes a whole lot of time, uh, again, explaining why they should be thinking about backup of these workloads, right?
So this is where PU comes in, is where the PU advantage, uh, shows up for our customers, right? Um, and again, when you support 80 plus workload, uh, it is very easy to become a platform that is tailored for the least common denominator, right? Uh, do something very simple, just dump everything and get it back.
Um, but that's not what Haiku platform is about. And I'll, I'll show that, uh, when I, when I do a demo as well, the platform is flexible up to understand the dataset that we are protecting, uh, and ensure we deliver consistency, we deliver immutability, we deliver cross, uh, regional project protection and, and also support, uh, uh, high level of granularity, both in terms of what we back up and as well as what we can restore, uh, within this, uh, uh, dataset, right? I need to warn you that we are back time for the presentation.
So please do, uh, bring it to a reasonably quick close. I don't think it'll be time. Awesome.
Will do. And, and the other thing I wanna highlight is BigQuery is, is ecosystem. It's not a workload.
Uh, there are over a hundred connectors to BigQuery that can read from all kinds of connectors. Uh, and again, having a broadest platform that supports, uh, you know, a 80 plus workloads allows us to actually not just protect the Lakehouse, but also a whole lot of connecting ecosystem players that are out there, uh, in the market as well. The, the, the final takeaway from the Lakehouse section is that Haiku has a patent pending technology, um, around atomic backups.
That is something that has never been, uh, built in the industry before. Um, and so Satya, if you could just very briefly, in 30 seconds or less, let's just provide our, our friends here with an overview of what this patent means to the industry as we start to drive towards a world in which we are really, really providing end-to-end protection of ai. I think that ultimately is what we're trying to get to.
And, you know, there is so much jargon and talk about AI these days that sometimes we feel like we need to provide a lot of context because people aren't talking about the infrastructure. They're not, as David really pointed out nicely previously, you know, at Dell even, they're building a whole AI factory. And then David said, wait a second, we gotta protect this.
These are enterprise customers. And so this innovation is something that we believe is gonna change the way that people think about that protection of ai. So, Satya, just really quickly, 30 seconds or less, and then we're gonna wrap things up so you guys can all get onto your lunch.
Awesome. And, and so this is the customer spotlight that actually brought to us the problem, right? What they have is a 700 terabyte, uh, data lakehouse that's ingesting data from lots of different sources, and they're all segregated in different forms and, and different data sets, right?
And they have dependencies across these data sets. So they came and told us that whenever we run backups, by the time we get to the fifth table, the first table is no longer in sync with the fifth table that I'm backing up. I've got a 700 terabyte, uh, footprint, right?
So today with cloud native capabilities, they can only export one table at a time, and they're all kind of outta sync whenever there is dependency. So what we build is we leverage native time travel capabilities within BigQuery to line up all of these data sets to a single reference point. And from there, we back up so that you get an atomic backup where all your data sets are as of a certain point in time when you run your backups.
And again, that's, uh, uh, very, very unique to Haiku and, and, uh, we, we have a patent pending, uh, uh, on that, uh, particular solution. So.