UT07x02: Building an AI Training Data Pipeline with VAST Data – Utilizing Tech
Model training seriously stresses data infrastructure, but preparing that data to be used is a much more difficult challenge. This episode of Utilizing Tech features Subramanian Kartik of VAST Data discussing the broad data pipeline with Jeniece Wnorowski of Solidigm and Stephen Foskett. The first step in building an AI model is collecting, organizing, tagging, and transforming data. Yet this data is spread around the organization in databases, data lakes, and unstructured repositories. The challenge of building a data pipeline is familiar to most businesses, since a similar process is required in analytics, business intelligence, observability, and simulation, but generative AI applications have an insatiable appetite for data. These applications also demand extreme levels of storage performance, and only flash SSDs can meet this demand. A side benefit is the improvements in power consumption and cooling versus hard disk drives, and this is especially true as massive SSDs come to market. Ultimately the success of generative AI will drive greater collection and processing of data on the inferencing side, perhaps at the edge, and this will drive AI data infrastructure further.
Transcript
Model training seriously stresses data infrastructure, but preparing that data to be used is a much more difficult challenge. This episode of utilizing tech features, Kartik from Vast Data, discussing the Broad Data Pipeline with Janice Roski from Soddy and myself. The first step in building an AI data model is collecting, organizing, tagging, and transforming data.
And once the model is trained, you've got to present that data to be used. That's what we're talking about on this episode of Utilizing Tech. Welcome to Utilizing Tech, the podcast about emerging technology from Tech Field Day, part of the future Improve.
This season is presented by soy and focuses on AI data infrastructure. I'm your host, Steven Foskett, organizer of the Tech Field Day event series. And joining me today as co-host is Janice from soy.
Welcome to the show. Thank you for having us, Steven. So, as you know, we've been, uh, planning this, uh, whole season out to have lots of great guests from soy partners, from industry partners, from people who really know quite a lot about the data challenges of ai.
But I think that sometimes people get a little bit wrapped around the wheel thinking about feeding GPUs. They think about training data, they think about high performance. They think that that's like the only thing that matters.
Now we see that when it comes to the GPUs, they think that that's the only thing that matters for AI infrastructure. They forget storage, they forget compute, they forget inferencing, they forget everything else. It's all about the mg.
But in, in storage too, you know, you can't train models unless you've got data, right? Lots and lots of data, and everyone knows that, right? That is the headline every day, all day when it comes to AI loads and loads of data.
And to your earlier point, right? It's all about the GPU, the compute performance. But we are really excited today to have one of our very best partners here, uh, on the show.
Uh, and he's going to talk a lot about why data is so important all the way through, not just looking at how do you feed the GPU or the compute, but the overall data pipeline. And it doesn't just stop at training, right? You have to look at the whole thing, the whole kit and caboodle all the way through to inferencing and then some.
So, uh, we're delighted to have, uh, Kartik, Superman and join us from Vast Data. Welcome, Kartik. Thank you, Janice.
Uh, hi folks. Just by way of introducing myself, I'm Kartik. I work for Vast Data.
I've been working with these guys for four years. Um, prior to that I was at EMC and Dell, and before that I used to be a particle physicist. Um, strange change of careers over here.
Last two years, I've been obsessed with anything connected with generative ai, as all you guys know. So, super happy to be here and to be able to contribute whatever I can. So we've talked a lot about the value of, um, of data, uh, to generative ai, um, especially when it comes to model training, but also transfer learning and fine tuning, and of course actual inferencing and generation.
Um, but there's a lot more to it than that. I mean, it, it's one of those things I think where people don't realize, you know, when you're gonna go to that training, you know, what data are you gonna use and where does that data come from? That's a much, much bigger process than just turning it on and saying, okay, go train.
Right? Yeah, you are a hundred percent right. The people tend to think of generative AI and AI in general for that matter, as something where you consume data for the GPUs subsystems that you have, and they take a first look at it and say, Hey, let's look at, say, chat GPT.
You know, when, when, like G PT three was first trained, that's 175 billion parameter model. It was trained on a little over 500 gigabytes of data. It was like, oh, well, I don't really need that much data.
And, but the, the training ran for months and cost like a whole bunch of money. So it was like, oh, okay. GPUs are important.
No question. GPOs are important. Uh, the, they need data and they have a lot of performance requirements due to checkpointing and other things like that while the training is going on.
But they forget that there's actually a whole bunch of heavy lifting that happened before the 500 plus gigabytes of tokens were created that was distilled from petabytes and petabytes of data all on the internet to build these foundation models, which we all know and love, like llama and, you know, Palm and, and others like that. Uh, what they don't realize is that that data actually fits as part of a pipeline. And this is a one takeaway I'd le like to leave for our customer or for anybody listening, is that pipeline encompasses all the way from the raw data to inference.
And every one of these is data intensive as we go along, not just a little bit under the, uh, GPUs. Now, this situation becomes even more complicated when you look at how enterprise is going to ultimately use it. 'cause keep in mind, these big models are trained by people who are professional model trainers.
Enterprise now wants to incorporate their internal wisdom, their big corpora of data into the input stream for training these models. This is a data wrangling problem, which is enormous. And here you literally have customers who are tens if not hundreds of petabytes of data who are now trying to figure out, how do I get a grip on this and how do I actually make these fit in the whole AI pipeline, uh, start to end.
So, so Kartik on that note, I mean, everyone's kind of looking at the overall data pipeline, and I think vast in particular, right? Um, we've worked with you guys for over 10 years, uh, and you've always continued to innovate and do things a little bit different. So can you tell us how Vast might be looking at, you know, AI with the data pipeline and, and how it's different from, you know, maybe other types of, you know, platforms in the in industry, kind of, you know, looking at the overall pipeline?
That's a great question. Janice VAs is a data platform company, not just a storage company. Of course, at the heart of what we do is the most scalable multi-protocol storage subsystem on earth.
Uh, we are extremely high performing. We have all the necessary certifications from Nvidia, base pad, super pad, now most recently NCP also. However, what distinguishes us is not only that we are high performing and we scale well, but we also expose our data through other modalities than just file and object protocols.
We also expose ourself as a table. So now we are open to analyzing data using other toolings, such as Spark or Trino or Dremio on top with native tabular structures within us, which is crucial for the kind of data crunching you need to actually take the raw data, which is currently sitting in large data lakes on Hadoop or Iceberg or Minia or, or object stores or something like that. We wanna be able to corral that and to be able to give the transformation platform to convert these into the things that can be actually input within, into a model.
Now, we do this with an exceptional degree of security and governance and controls across the whole thing. And probably most importantly, because we are able to consolidate the entire data pipeline, the data connector with the data pipeline on a single platform, we eliminate copying of data or moving of data, and we are able to reduce the data footprint much more efficiently than anybody else. So we think just on just the basis of the unification of the different pieces of the pipeline, as well as extending this to a global namespace, that alone makes us the perfect fit for the AI world that's emerging right now.
Yeah, I would imagine that, um, the many people listening could, um, kind of get their hand their heads around this. I mean, think about the data that you have, think about how it's organized or probably disorganized, and then think about how w how much work would it take me to locate, identify, tag, organize, consolidate, you know, basically get all that data ready to even start doing any kind of, uh, an AI model training or retraining or fine tuning. Think about all of those tasks.
Think about all of that data, and throughout it all, of course, it's not just gonna be one person, it's gonna be an entire team. It's gonna be a disparate team from multiple parts of the organization, maybe even multiple organizations. It's gonna have many different data types, many different data sources, and all of that has to come together in a nice organized fashion and then be presented to the model.
That's incredibly difficult. If you've ever sat there and, and watched that little, that little bar grow on your laptop as you're copying a file, well multiply that by a thousand x When you're talking about enterprise data, it's incredibly difficult to move this. And, and so it's really not overstating it to say that what you just described Kartik by having all of this data on a unified platform with a unified namespace.
Well, that's pretty transformative, wouldn't you say? Talk to us a little bit about how that really works in customers, um, environments when they're preparing to prepare a model. Yeah, so I wanna, that that's, that's, that's a fantastic, uh, way to describe the problem.
I think you nailed it there, Steven. Um, in the enterprise, there's multiple levels of problems to solve. First is all the data which was created over the last 30 years is potentially going to be something that we are gonna want to train models with.
This data is in data silos all over the place. Some of them are in data lakes, some of them are in data warehouses, some of them are on the cloud and Snowflake or Databricks or something like that. Uh, some are on-prem, some are off-prem.
Secondly, very few people actually understand what data they have or what the semantic meaning is of that data to the business processes, which are core to them. There is no ontological model, which typically exists in these large enterprises, and people are scrambling to build one. So you'll hear a lot of talk about things like knowledge graphs and, you know, data fabrics and data measures and things like that.
These are all basically efforts to get a grip in simple terms on where is the data, what is the data and what use is It to me then comes the actual task of calling that data and making it available to be able to train larger models like this, which will then give rise to business use cases to move forward. So the philosophy most people are doing are, are adapting that right now is to say, I need a fundamental transformation of enterprise architecture to get what we call AI ready. Now, the silos of information that stuff lies on is often infrastructure that was built 10 years ago.
It's sitting on 10 gigabit networks, and there are petabytes and petabytes of data. I, I mean, how am I ever gonna get, get a handle on this? How am I ever gonna move it, et cetera.
8 trillion in this transformation, is to start to aggregate the data, to understand the meaning of the data, and then to prepare it and decide how I'm going to vector it towards a variety of models, which are then going to be able to transform how my business operates. At the end of the day, training is important, like we said, but frankly, money doesn't get made in training. It's a money sink.
It's all in the inference. You gotta get it in the hands of the end users, and that's what needs to happen. And all these needs to be done in a secure, highly governed fashion as transparency, copyright violation, intellectual property issues, all that start to become more and more important.
So there's a lot of work that the traditional enterprise has to do to do this, uh, but it all starts by understanding the data and starting to consolidate it in modern infrastructure. And this is where Ban comes in. They got the best mousetrap over there to provide the underlying storage, which is high performance, and, uh, can really add advantage.
Yeah. Yeah. So a couple of questions before I have a couple.
Well, one, I'll just wanna, um, you talked a lot Kartik about enterprise, and I just wanted to get your thoughts a little bit more about, uh, HPC. You know, we've worked with you over the years. Uh, you know, high performance computing is, you know, obviously turning more toward what, what is now ai, but you don't really go backwards, right?
You don't, uh, look at AI and go, go to HPC. So just get, love to get your point on how, um, how do your partners in the HPC space tackle this AI data pipeline? And is it different from enterprise customers?
So AI is just a subset of HPC. I should probably be a little more careful than that. People will take umbrage to me calling what we are doing right now.
Ai, AI is actually very broad subset. You know, what we are talking about is specifically is generative ai, which is neur network based. And that's a, and especially large models and how, how they operate here, uh, there's absolutely no reason that anyone should believe that large language models are the only kind of competition that's gonna go on.
It is shocking that 95% of analytics that goes on in most enterprises is actually good old fashioned structured data. And it's standard machine learning algorithms like linear regression, logistic regression, basin forest, that's the kind of stuff that gets done. Now, when I translate this to the lens of an HPC operator, and many of them are customers of ours, we are widely deployed across most of the national labs, most of the large HVC centers, et cetera, they are seeing that the traditional classic HPC cluster, one that had a lot of compute nodes and high performance parallel file systems under it, is giving way to heavily GPU accelerated codes as well, right?
Not just GPUs, could be GPUs, could be fegs, some kind of co-processor acceleration is getting more and more ubiquitous in here. So they too are transforming. One of the big d big things they're doing is they're moving away from traditional hard drive based technologies to solid state technologies to be able to do this.
Because the IO patterns for AI are tend to be more random read dominated, and at this point it becomes an IOPS game and describes, unfortunately, a constraint because of the mechanical construction they have. And Nan can do a heck of a lot better job at being able to deliver this. So they're all looking at high performance, all flash namespace to be able to deliver the kind of performance they need to handle this unusual mix of workloads, traditional HPC simulation, NPI jab, high throughput, large block, you know, type, you know, sequential, read, sequential right type workload contrasting with these heavy random IO intensive workloads.
At the same time, the combination of these perfect platform is a completely solid, safe platform. That transformation is well underway. The only hindrance to this is economics.
They say, oh, I like flash, but oh my God, uh, it's too expensive. I can't afford it. This is where soddy comes in and gives us very affordable, dense NAND with very high capacity in some of the secret sauce that we've developed in conjunction with soy, has allowed us to increase the endurance of those systems to a point where it's structurally become affordable.
And then the data services to shrink this data, again, bend the cost curve to a point where it becomes something which is, you know, which enterprise can use, which is very, very tenable for what they want to build at this point. But also for high performance, uh, computing shops. They also are moving in exactly the same direction.
Pipelines are the same. At the end of the day. It's data guys, people, it goes through successive, uh, processing steps and successive refinement.
And ultimately what emerges is either great science or great business, either ways we wanna be right in the middle of it. This is what we do In a way, Kartik, there's an analogy here to the GPU question. Mm-hmm.
One of the things that we talked about all last season on utilizing tech was the fact that, you know, when it comes to training, you have to keep those GPUs fed or you're not making maximum use of that incredible hardware investment. Hmm. It seems to me that the same question is true of solid state.
You're buying solid state storage, not for capacity, but for a combination of capacity and performance. And so it's a good idea to make sure that the way you're using it is able to leverage that performance. And that means that you have an intelligent storage engine, you have an intelligent storage approach that's able to make the me best use of this incredible resource that's literally orders of magnitude faster than what you'd get on spending media.
Mm-hmm. And I think that today we've gotten to the point where, and, and maybe you can tell me if you agree with this, we've gotten to the point where primary storage really is flash, full stop. It's flash, let's face it, And, and discs have a place.
Mm-hmm. But it's not in primary storage, it's in secondary applications. It's in data protection, it's in archives, it's in, it's in the, the, the, the areas that don't need any kind of performance.
Is, is that what you're seeing in the market? Yeah. To, to a large extent, yes.
There's, there is another that yes, people like flash, uh, for multiple reasons. One is it is high performance. Okay?
Secondly, the power form factor is much more manageable at scale as, as drive size become bigger. Uh, but probably as important is the, uh, you know, what I would describe as the, uh, dependability or of, of the actual performance itself. So it's not just raw performance.
It's seems counterintuitive, but training actually is a GPU bound problem. You actually do very little io the input datasets and the output models are pretty small. That's not where the, you get hammered for iil, where you get hammered for IO is actually doing checkpointing.
So absolutely, you, you need to have systems that are very high performance, and that's true. But what's crucial about all flash namespace is as many, many of our customers have found out, is you no longer have to worry about is my right data at the right place at the right time. So, as you know, well, Steven, we eshoo the idea of tiering just for that reason.
We believe that AI and what is emerging now, we are going to go into multimodal models. Data volumes are in the petabytes. I just talked to a customer who's gonna buy a couple of thousand Blackwell GPUs with 60 petabytes of storage.
That's a lot. That's a hell of a lot. I can guarantee you that there's some very, very dense content in there that they intend to analyze.
Key thing is the nature of AI says that you cannot predict what you're going to need when, when the GPU subsystems need it. They need to go and find it. And at that point, that data can't be stuck on a slow object store somewhere else, orano another disc based tier somewhere else.
Because that's a, that's a buzz scale. It just, you know, all of a sudden you go, whom my latency went from one millisecond to 200 milliseconds or 800 milliseconds? You know, what am I gonna do?
Uh, so the uniformity and the, uh, predictability of the workload I think is as important. Now, the other point I would make is that the reason why, I think the other reason why I think flash is ultimately going to reign in this space is so solid. Im is introducing bigger and bigger spindles.
You're gonna, we are reaching limits as to how much you can do with mechanical drives. You know, right now people are shipping 20 terabyte drives, maybe 24 terabyte drives, you know, uh, the, the long awaited promise of hammer is still not materialized. Uh, in the meantime, soine is like charging along 60 terabyte drives, 120 terabyte drives the density from a capacity, from a floor space and power is going to be significantly lower in these kind of environments than what we had in disc space environments.
I think just that Delta alone is going to basically eliminate drives. Mm-hmm. They're just not power efficient enough, not space efficient enough, and not performant enough.
They're gonna get squeezed out on the low end with tape and on the high end with all, all flash based systems. And this is something I've seen in conference after conference that I've been in that that's really what's gonna happen in the end. Yeah.
Thank you for saying that Karta. 'cause that was gonna be one of my questions for you was, you know, what's your opinion on the space and the power consumption? Obviously this is a really hot topic.
Uh, can you comment on any of your partners or customers that you're working with that have, you know, really been able to reap the benefits of both feeding that GPU, keeping the performance well all, you know, all the, well, bringing down the overall, um, you know, cooling and, and heat? Yeah, so we've been humbled and privileged for being the de facto standard for variety of people in the AI space. On one side, we have several superpower deployments, which are going on with Nvidia, as you know, that's, uh, their flagship offering from a single tenant perspective.
Uh, but we are also the standard for a large number of the, a new tier of cloud service providers. I call 'em AI Cloud Service, pro ai CSPs. These defer from the tier one, uh, cloud providers such as AWS and Azure and Google in the sense that they're not general purpose.
They will build ground up with the intense power footprint, the RDMA networks to be able to do, uh, communication between the GPUs and high performance storage specifically to be able to tackle the challenges of generative AI itself. So in this space, we have many, core Weave is a great customer of ours. They deploying over a hundred thousand GPUs globally.
Uh, we routinely see production jobs running huge, huge training over there. But equally interesting to us is some of their customers came to them and said, I'm having data wrangling issues. Can you help me here?
Can you help me actually make sense? Uh, do the pre-processing as well, because you start with a large amount of data and you're gonna condense it down to a small amount. Many, many cases.
That's very CPU led workload as opposed to g pled workload. Again, to your older question about HPCI, I don't think these are different. These are actually blending together to form one common infrastructure where both CPU as well as GPU and other forms of co-processor will coexist in the same data pipeline.
So that's also gonna be very much be there. Lambdas and other huge customer of ours, they're also extremely, uh, in, in, in, in the same boat. They also have a lot of GPUs and their intent is to offer GPUs a service.
So here's here, again, security and governance becomes super critical, right? Because if you're in a multi-tenant environment, guess what? You're gonna have to keep complete logical separation of data comp, you know, and you're gonna have to be able to encrypt the data, encrypt the traffic connected with this, uh, to be able to reduce the data within that, that kind of domain to have to build, you know, FedRAMP capable infrastructures, IL five, IL six level protection to build zero trust architecture.
All of those are things that we excel at, and that's really where we go, uh, much more than anything else. So these people consume data at scale, and they're gonna start consuming data at a much larger scale. So far, we were looking at tech Space lms, there's a whole class of large scale vision problems.
There's a whole class of large scale multimodal problems and the huge sucking sound of all the enterprise data getting filtered into these generative models, which is going to drive data volumes through the roof. Along with that, it will also drive the stability, the, you know, dependability, the predictable performance, the scale, the security, the governance, all of those will start to come into the forefront. And so we're, like I said, really fortunate working with these customers has taught us a lot.
Mm-hmm. We now understand what the real requirements are for how to actually go to market with this. And this helps us innovate continuously with ourselves, with our partners.
You know, we've just reached an no EM agreement with Super Micro. You may have heard that we've already had a deep relationship with HPE, with GreenLake for file. They have different hardware stack, but it doesn't matter to us.
We are a, we are a software offering, but we want that kind of diversity to be able to support all kinds of environments. But the core problem is that we wanna solve is we want to solve the customer's data platform problem is how do you build something which can take all data, analyze it in any modality you like, be it as through a database engine or through PyTorch, and then ultimately take it through to inference and fundamentally change how they operate their business. Yeah, that's, that's a really good point that you make there with, uh, with the inferencing question as well, because one of the things I think we're gonna see is, uh, proliferation of data on the, well, rubber meets the road side of things, you know, on the edge Type of Thing.
Yeah. Your model's trained, everything's ready to go. Now let's turn the data fire hose on on that end.
Yeah. And that's just gonna cause even more data to be collected, to be processed, to be stored, to be acted on, and it's gonna drive up the demands for, uh, at the edge, uh, in the cloud, in the data center everywhere. And, um, and I think that that's natural.
I think all of us can see that because, you know, chat GPT is lovely, but it's text, you know? Yeah. Okay, now let's, let's throw, uh, images at that.
Let's show documents at that. Let's show throw, uh, video. Mm-hmm.
Uh, streaming video, multiple cameras now, now what kind of data requirements do we have? It just grows and grows and grows. And, and I think that that's, you know, it's good because of course these are new applications that are hopefully gonna be productive and, and profitable for people.
But it's also a challenge because we're gonna need to be able to handle that kind of data, which again, you know, you need to have intelligent infrastructure and you need to have high performance, and you need to worry about environmental impact and, and all these things. Yeah. They couldn't agree with you More.
Inference has its own peculiar set of technical and business challenges, which are, which are interestingly enough not present in training. It, it, it on the technical side, you know, training, I can hog A GPU, it's all, mind, mind. Mine inference is not that way.
In a multi-tenant edge inference use case, my edge GPUs or whatever, accelerator's doing inference, uh, need to have access to multiple models. I may need to load and unload these models depending on who is exactly doing the inference. Worse, if I'm doing retrieval augmented generation, or what they call rag, and people have vector databases, which codify the internal things that the company does, then those again would start to move into the edge.
How do I do this rapidly? How do I do this in a governed way? The EU AI Act was passed, as you know, a couple of months ago.
It mandates that any of the data that is used for query value, especially for high risk business, needs to be preserved for eternity. All the queries, all the keys have to be preserved for eternity. And these are like, these are very difficult technical problems as well as governance problems to solve.
And we, we believe we have a superior mousetrap in that sense to be able to do this. Mm-hmm. Well, thank you so much, Kartik, for joining us and talking about this.
I think you've really opened up the eye my eyes and, and, and those of our audience as well to the greater question, because again, it it's very easy to, to my point at the beginning, it's very easy to focus on how do I feed training? How do I keep these GPUs active and forget that the data pipeline starts way before that and continues way after that. And the volume of data is just absolutely incredible.
So thank you so much for joining us here, uh, today. If people are interested in this, where can they, they find you, where can they continue the conversation? com.
There's a massive wealth of information of customers we worked with of how our technology works. It's great. com/white paper.
Uh, me, I'm easy to reach. com. It's easy.
K-A-R-T-I-K. And uh, I'll be more than happy to respond to you guys, uh, but anyone from Vast Contact, anyone from Vast locally, and they'll be able to help you as well. And they'll be able to guide you to me and to some of my colleagues.
That's the easiest way to do it. So thank you for listening to, uh, utilizing tech focused on AI data infrastructure. The Utilizing Tech podcast series is available in your favorite podcast application as well as on YouTube.
If you enjoyed this discussion, please do give us a rating and a nice review in your podcast application of choice. It's always great to hear from you. This podcast was brought to you by soy as well as Tech Field Day, home of IT experts from across the enterprise now part of the Futurum group.
com or find us on X, Twitter and Mastodon at utilizing Tech. Thanks for listening, and we will see you next week.