Why Storage Matters to AI in 2025 with Solidigm
Solidigm focused on the evolving role of storage in AI, specifically highlighting its significance in the AI data pipeline through 2025. The presentation emphasizes that the AI workflow involves a series of distinct tasks, each with unique demands on hardware, and that it’s often distributed among different organizations. The presentation identifies six primary stages in this pipeline: ingestion, data preparation, model training, fine-tuning, inference, and archiving. Solidigm’s perspective underscores the increasing data intensity of these steps, especially during data ingestion, inference, and archiving.
Solidigm explores how storage choices impact these stages, distinguishing between direct-attached storage within GPU servers and network-attached storage. Direct-attached storage is optimized for performance-intensive tasks, while network-attached storage is used for larger datasets and offers capacity and cost-effectiveness. Ace Stryker highlights the rising importance of network storage, driven by advances in network bandwidth and the growing size and complexity of AI models and larger datasets for RAG. The key takeaway is that great storage facilitates larger models, longer interactions, and improved outputs on the inference side.
Finally, the presentation showcases a collaboration with Metrum AI, presenting a real-world demo that addresses the challenges of data-intensive inference by offloading model weights and RAG data to SSDs. This allows running 70 billion parameter models on less powerful hardware, which would have been impossible without storage-based offloading, saving on costs and hardware requirements. The demo emphasizes the potential of SSDs to enhance performance in AI applications by reducing GPU memory usage. The collaboration offers insights into the benefits of leveraging storage in AI, as the partnership also showed that offloading to SSD could provide the same or better performance as DRAM in this application of retrieval augmented generation.
Presented by Ace Stryker, Director of Market Development, Solidigm. Recorded live in Santa Clara, California, on April 23, 2025, as part of AI Infrastructure Field Day. Watch the entire presentation at https://techfieldday.com/appearance/solidigm-presents-at-ai-infrastructure-field-day-2/or https://techfieldday.com/event/aiifd2/ for more information.
Transcript
My name is Ace Stryker. I am director of Market Development for Soine. As Scott mentioned, what I wanna do now is spend some time with you on what we call the AI data pipeline.
And if you were with us for the last AI infrastructure field day, which was what, October, November, somewhere in there. Mm-hmm. Uh, you heard from us on this topic.
Um, but we have learned a significant amount since then and refined our view of this. And I wanna share some of the new learnings with you. Okay?
So when we talk about the data pipeline, what we're talking about is the entire workflow from start to finish of doing AI work, right? It turns out it's not just one job, it's a series of tasks, each of which, uh, have discreet goals, have discreet workload characteristics and discreet hardware requirements associated with them, right? And so we're laying out what we view from the solid, IM kind of data centric point of view as the six main steps.
I think we had five when we came to you last year. We have now promoted fine tuning to a top level step, and we'll talk about why. Uh, but one of the, one of the new nuances here that we're trying to capture in the way that we talk about the data pipeline is that it's not typically one organization that's doing this work from start to finish.
It does happen. Google does train a model called Gemini, and then make a public chat bot available based on that model and run inference from their servers and controls the whole thing to end. But more and more often, we're moving into a world where this work is actually disaggregated, and there's at least a couple different kinds of entities involved in the work of doing AI and creating value from ai.
So what we're seeing, again, in broad strokes, because this is a generalized view, and the answer in AI is always, it depends, uh, is that the first three steps here where you ingest a bunch of raw data, you clean it up and prepare it, and then you use it to train A capable model tend to be done by one kind of organization. And the output of that process is called the foundation. So a relatively small number of companies doing this work on a global scale that have the resources to build these really complex foundation models, right?
These are the hyperscalers. So think about Google meta anthropic, open ai, deep seek, Alibaba, right? You're, you're probably familiar with most of the names on that list, but it's a, it's a list that numbers into the dozens or perhaps hundreds, but it's not a group of organizations.
Whereas the second half of the process, what we call enterprise solution deployment, is something that's done by thousands or even millions of organizations globally. And what they do is they come in and step forward, they shop around and they say, which of the foundation models that was created as an output of the first half of the process is the best fit for me? And then they proceed to fine tune it on proprietary data and then actually deploy it, run inference, and archive the results.
And so this is what we mean by the data pipeline. And I've got a couple builds on this slide to layer some additional stuff into it, uh, that's highly relevant to us as a, a data focused company. So the first here is just, uh, just examples.
Not comprehensive, but everywhere you see this little disc icon, the same one up there is an example of an incremental data set that's being generated in the course of doing ai. And it happens all over the place. There is not just a training data set here, a model here, right?
Uh, as you grow through, you're creating all kinds of data. Uh, as an example, you know, the training data set that's cleaned up coming outta. Step number two is incremental to the raw data set.
You started with during training, you're gonna end up with, um, uh, sort of intermittent safe states of the model called checkpoints. All that data must be stored somewhere, right? If you're doing inference, that's really where we see data magnitude exploding.
And we'll talk more about that in a bit. But you've got inputs and outputs many times over many, many more times. Uh, is inference happening than training on a given model if you're using rag or retrieval augmented generation, which we'll also talk about that can multiply data magnitude.
And so when we say data is everywhere, it's not just a a, a catchy phrase. It's a, it's an observation across many customer discussions in our own path finding in the lab. It turns out you can, again, in a very general sense, assign data magnitudes to each of these steps.
No doubt there are folks in this room looking at this and thinking, I know of examples where that's not true on that step and that step, right? Uh, and we're not asserting that these are absolute ranges or that they apply in every case, but across, again, many customer conversations, lots of research, lots of digging into different use cases. What we've seen is that there are steps in the data pipeline that are more data intensive than others, right?
And it tends to be upfront where you're ingesting a bunch of raw data to train a model. And it tends to be at the end where you've deployed and you're doing inference, potentially with a rag component, and then you're archiving everything for retraining later, or audit purposes, or whatever the reason might be. But this is all cumulative, right?
As we said before, there's no AI without data. There's no data without infrastructure. So as you look at how this all gets generated, moving from left to right, you've gotta think about what's the most cost effective way to store that.
Uh, what are my kind of minimum performance requirements by stage? And that's what this next kinda layer on the build is intended to address. So, again, very high level, but when you look across the solid I product portfolio and the many cool products that Scott, Scott just changed his shirt and threw me off there, he's, uh, sorry, uh, I didn't mean to spoil a surprise there.
Um, but, uh, broadly speaking, within an AI cluster, you can allocate, uh, the products that solid eye makes into one place or another. When we talk about direct attached storage, that's a term that refers to the storage slots inside of a GPU server, right? So you take a DGX box from nvidia, it may have eight U2 slots, right?
And you put a certain amount of storage in there that is directly attached to the GPUs by PCIE lanes, super low latency, relatively constrained on capacity. 'cause there's just not a ton of slots in those boxes. And in those places, it tends to be, uh, a very performance focused game where the goal is really keeping GPUs maximally utilized really high iops from the storage drives, means high utilization from the GPUs, which you've spent a lot of money on.
Uh, and so that's where we see the focus among customers. Whereas on the blue side, what we refer to as network storage, these are now boxes that are storage focused. Sometimes they're called, uh, storage servers.
Sometimes they're called JBoss, just a box of drives or J bs, just a box of flash. There are various flavors of this thing, but the idea is they, uh, have a bunch more slots, maybe 24, maybe 32 slots, uh, for a lot of storage. Uh, and they are connected to the direct attach boxes over the network using ethernet or InfiniBand or what have you.
Uh, and so we can layer those different, um, tiers within an AI cluster onto the stages of the pipeline. Again, very broadly speaking in the following way, you typically start with a lot of data, uh, in network storage. You read it into the GPU servers during data prep and you manipulate it.
That's primarily a CPU intensive activity. And then once you have that cleaned up, ready to train data set, the training, uh, fine tuning is occurring obviously within the GPU servers. Uh, inference can, uh, historically has been sort of a direct attach focused, uh, area of activity from a storage perspective.
But for reasons we'll talk about with STEAM here in a minute, uh, uh, we are seeing more and more demand for network attached storage involvement and inference workloads as well for pulling in rag data, uh, for pulling in, um, you know, multiple models and complex use cases where you gotta swap 'em in and out. And then finally, archive is a process of sending everything back to that colder tier of storage that Scott talked about. Question about the, yeah.
Advancement of network storage from that perspective, do you think that's a result of advancements in networking technology that can get you that storage connectivity faster? I think there are bandwidth, two reasons. Yeah.
One absolutely is the improvements in network bandwidth. We're now approaching a world where, um, the, the, the bandwidth available to move data across the network matches the bandwidth in an individual drive, right? And so you're no longer bottlenecked in that way if you look at kind of the future specs for PCIE versus, you know, future generations of these networking protocols.
But the other thing, so that's kind of on the supply side, right? Like networks are becoming more capable and faster on the demand side, we're seeing bigger models, we're seeing more complex and longer interactions with models. We're seeing bigger and bigger rag data sets that enterprises wanna plug into pre-trained models, which we'll talk about a lot, uh, with steam here in a minute.
Uh, and so yes, the, the technology's becoming available, but it's, it's happening alongside, uh, uh, interest from the users and consumers of these models, uh, to pull a whole bunch more data into the inference side of the pipeline than ever before. Gotcha. Good question.
Okay. And the last thing to understand, uh, at a high level here on the data pipeline is predominant IO types, which I've just added here in the middle. So this is really for the storage geeks.
We don't typically expect our customers to worry a lot about this stuff. You know, this is where we, we have conversations or can kind of give some advice, but it turns out that depending on the stage of the pipeline, you're asking the storage sub subsystem to do very different things, right? Training, for example, tends to be a random read, intensive kind of activity interspersed with sequential rights as you checkpoint along the way.
Uh, whereas, you know, something like, uh, archiving data at the very end, you're, you're sort of reading straight from your direct attach and writing straight to your network attach, right? And so understanding those parameters and mapping them onto, uh, characteristics and, and capabilities of individual storage drives is important to make sure that you're making kind of optimal storage subsystem choices. Does that make sense?
Any questions on this one? Okay, I Wanna Alright, slide. Sorry.
It's a really good slide, particularly with the build. Oh, good. This is an example of one that makes it very clear from what you're doing.
You also have core DC near edge and far edge. I like near edge and far edge. That's, that makes sense to me.
Good. I'm glad to hear that. Thank you for the feedback.
Uh, okay. So the next couple slides before I bring steam up, I want to talk specifically about why, um, great storage matters in training. And then the next slide is about inference, but we've kind of talked about, okay, faster storage in general means faster outcomes for ai, higher capacity storage in general means you can feed more data into your AI pipeline, right?
But it's important to understand that there's a level of nuance beyond that, particularly in how data moves between memory and storage. Because let's be honest, like in the real world, AI developers are trying not to touch storage whenever possible, right? If they could do everything in memory, they absolutely would.
It's not possible economically, uh, for reasons that Scott highlighted with the data pyramid, right? Um, so how does this stuff actually get moved around? I'll show you an example here, um, of a, of a workflow.
And again, this is very general, but what we tend to see is raw data being written to network attached storage to begin, uh, the process of training a model. And that stuff then gets read out of network attached storage across the network into the GPU servers, where it's now located in direct attached storage, and it goes through this process called ETL or extract, transform load. This is the cleaning up of the data.
This is deduplicating and vectorizing the raw data to make it into training ready data sets. That's a CPU intensive activity. That prep data then can go in a couple of places.
One is it's written back to disk, uh, back to storage inside of the GPU server. And the other, of course is it starts to go into a machine learning algorithm to actually develop the model. Right?
And outta there a series of things happens. One of the data out outputs of the machine learning algorithm is the model itself. That real Quick.
Yeah, please. Sorry. Um, uh, is this a batch process or a continuous process?
Uh, I mean are each, are these stages running all in parallel? Yeah, absolutely. Yeah.
Yeah. To make best use of your AI infrastructure, you know, the smart kind of operators are gonna stage this stuff so that, you know, it's not like sort of zero to 100% here, and then this starts, okay. And then that starts, right?
It's an, it's an ongoing process absolutely. Coming out of machine learning, uh, you have the model itself, which is written initially to, to memory, and then to disc. And then you have the checkpoints.
And I'm, what I'm highlighting here is a couple of different ways checkpointing can happen. We don't need to really deep dive on this unless there's particular interest, but traditionally, checkpointing has been something where your, you know, these training epics can take a long time to, to, to train. Uh, a foundation model can take weeks or even months in some cases of just GPUs churning 24 7.
You know, a single hardware failure software issue can scrap all your progress unless you're saving the state of the model as you go along. Uh, that's the checkpoint. And what it, what's happened historically is they've been written directly to disc.
The challenge with that is when you're checkpointing you're not training, the GPUs are waiting for the checkpoints to complete. And so it became super important to checkpoint efficiently with high, uh, sequential right speed in this case, in order to get back to the work of actually training the model as quickly as possible. There are now approaches, and that's what this dotted line is.
It's called, uh, asymmetrical checkpointing, where the, the checkpoint data is actually offloaded, uh, um, first system memory and then written directly across the network to your network storage. And your training does not have to wait in the same way. Right?
So they've sort of pipelined a way that performance penalty in a lot of cases. What we've seen is really, it's only the hyperscalers who are kind of have that architecture built out that are capable of, uh, mitigating that performance penalty. And everyone else is still kind of dealing with, with synchronous checkpointing and inter uping their training.
But I just wanna give you a sense, there's a couple of different approaches going on here. And then finally at the end, everything moves into archive data at the end, right? So at the, the, the big name of the game here, oh, and I do have this, if folks have heard about GP direct storage, if that's a familiar term, just as a, uh, quick primer, that's what this arrow is.
It's an NVIDIA technology that enables a direct pipeline from your training data set, uh, on NAND or, or your storage directly into your, uh, GPU memory and it skips the CPU hop and reduces, uh, CPU, uh, utilization associated with processing iOS. Yeah. So just to check, has, um, SSDs the, the issue with sequential right?
Being the least, uh, or the most wear outable of solid state devices, at least back when they first came on board mm-hmm. Is that still true or are you have vendors like dym? Have you been able to overcome it?
I guess this is more of a generic question. Okay. Yes.
I mean, it would seem to me that especially writing that raw data set would just pummeled mm-hmm. Call state device and especially, uh, the, the checkpointing possibly Right. Archive, who cares?
It goes to, to dis Right? And it kind of sits there, right? Just curious, is there, is it the QLC aspect that actually makes sequential Right.
Okay. At this point, No, QQLC actually tends to have lower right endurance than TLC. Um, so it's, it's not a solution to the problem.
Okay. But we are seeing, um, uh, you know, advances in right. Endurance, uh, scaling in particular with capacity.
Uh, so I think Scott will talk about our 1 22 terabyte product in a little bit, which okay. Literally cannot be worn out in the five year warranty period. You can write to it 24 7 for five straight years and not hit the rated endurance.
Wow. Um, and so there are cases where it's becoming, uh, less of a concern, but on the lower capacity drives where you're having to write to the same blocks over and over. Right.
Um, gotcha. It remains a concern. Sure.
It's still Okay. Is there anything you would add to that, Scott? I feel like you're, you're closer to this than I am.
Yeah. So, um, no, I would agree that, um, there's a, there's a concern there and part of the, the concept of Right shaping is a big piece of that. Okay?
So, as to his point, we'll show you in a minute, because of the way the, the format of this show works, you get it in a minute as far as 1 22, but the ability to ensure the way the rights are to a drive does impact it. Okay. But as he mentioned, when we were talking 4, 8 16 terabyte drives, you could blow them up, no problem.
You get into 32 64, you start to do sequential rights. Our ability to do the background cleanup, the garbage collection that we call it Okay. Becomes less chaotic and less aggressive, and therefore artificially lowers the, what we call right amplification factor.
So the idea is every time you write, there's a factor associated to that 1, 2, 3, 4 higher the number, the faster you wear it out, right? We've got the capability to bring that right amplification factor down. Okay.
And we're also working with the industry on new techniques for data writing to drives mostly at the hyperscale level. Gotcha. Which is where a lot of the training takes place, where that Right.
Amplification becomes a one. Got it. So, and then it means every right is one, right?
You've got a hundred thousand or 10,000 cycles per cell, right? You've got 122 terabytes of cells, which is Sure. A few more zeros, right?
So, yeah. Yeah. Well, thanks for answering.
Yeah, Absolutely. Right on. Thank you, sir.
Thank you. All right. I'd like to keep us moving along here 'cause I wanna make sure we get steam up here quickly.
But the, the other side of the why storage matters for AI is here, this is not in flowchart form, but wanted to kind of highlight for you, uh, when the question arises of why does it matter if the storage choices you're making on the inference side, it turns out, uh, there's some really exciting work going on. And some of that path finding is what we're gonna talk about with Metro Maya here in a minute. There's obviously mainstream reasons why storage matters and inference, like if you're ingesting, uh, uh, rag data or, uh, your inputs for your model off of, uh, Nan somewhere, it needs to be fast enough to keep up with the demands, uh, of the cluster.
And then writing it back, uh, at the end, the archive step, uh, speed matters there, right? So good source decisions are important. Um, there's other, the, these two in the middle I'm gonna very briefly mention, and then we're gonna show you the demo where we actually do these things and talk about it.
But a couple of the, the areas where great storage shows up as a, as a solution for real world problems. And inference one has to do with rag database scaling. We'll do a quick primer on rag in a minute, but what we're learning is enterprises love it.
Last, uh, figure I saw was 80 something percent of generative AI deployments in enterprise have a rag component. So they're, they're consulting other external data sources that the model wasn't originally trained on. Uh, and, uh, the demand to go bigger and bigger on those external data sources, uh, has created an opportunity for storage to fill a gap and solve some problems.
Uh, and then model weight offload, I think I'll hold, and I'll lette ste ste talk about that one because that's a really, uh, cool one. But essentially you can, you can actually read some of the model weights themselves from storage as opposed to keeping them in memory. And then one thing we heard about a lot from, uh, Nvidia at, uh, GTC just recently was the key value cash, which you can think of as sort of the model's short-term memory.
If you get on, um, you know, Gemini and you ask it to create you a, a picture and you don't quite like it and you ask for some changes and some more changes, some more changes, the, the model has to remember that whole history of the interaction to have any idea what you're talking about or have the ability to iterate. That's the KV cash. And on more complex models and longer interactions, that can get quite big, quite fast.
Uh, and so there's an opportunity for storage to play a role there, uh, in, uh, uh, as the location of the KV cash, which has historically been in memory. So when we talk about the value of storage and inference, this is kind of the, the focus area. The takeaway here is it's really about you can do things with great storage on the inference side to use larger models than you'd otherwise be able to have longer interactions and better outputs.
We don't typically talk about storage unlocking, you know, greater, uh, inference speeds, right? 'cause again, that's a, that tends to be a memory game, but there are very material ways in which great storage, uh, shows up and, and improves outcomes on the inference side as well. Any questions on that?
Cool. Okay. Um, hello everybody.
I'm a Stryker. I'm the director of market development at Soy. And what I want to talk to you about now is some really cool groundbreaking work that we are doing with an organization called Metro ai, uh, in the field of data intensive inferencing.
And I wanna start with a picture of a giant robot in the desert. This is an art installation. It's called Meen nine.
I hope I'm pronouncing that right. The artist is Tyler Fuqua. Uh, last week was spring break, uh, for my family.
So I took my kids to Las Vegas and, uh, family friendly stuff only. Mm-hmm. Of course.
Uh, but there's, uh, uh, this guy is in Vegas now. There's an area, uh, called Area 15 in Las Vegas. Has anyone ever been there?
I see some nods of recognition. Yeah. If you haven't been, highly recommend it.
Omega Mart at Area 15 is one of the most bizarre tourist attractions I've ever been to in my life. It's very, very memorable. So That transpose 51, I guess.
Yeah. Yeah. They take a lot of your kind of assumptions and, and, uh, you know, preexisting knowledge and turn it on its head when you walk through the doors there.
But, so this guy's a big art installation. He's also a puzzle. He's covered in plaques that have encrypted messages on them.
Huh. Wow. And the game is, you can walk around and there's a key provided on one of the plaques, and you can try to decrypt the messages and get the story of the robot.
Okay? So I took my two boys ages 14 and 12, and I said, I'm really gonna impress them here. We're gonna use AI to solve these, uh, these messages, right?
Uh, and we'll do it so quickly and they'll think dad's really cool and tech savvy. So I took a picture of the key. These are screenshots from my phone, and I gave it to Gemini and I said, Hey, use this key to decipher the messages that I'm gonna give you and Gemini in fine form, very chipper and enthusiastic said, you bet.
Let's do it. Show me the first message. And so this is a, a shot of the first message that I found on the robot.
It's got about five rows of different symbols. And I said, this is the one I'd like you to solve first. And it came back with a very friendly, very confident, and completely incorrect response.
Said, oh, this is, uh, this is saying H-E-L-E-O, which decipher to, hello. Neither of those things are true. None of these symbols actually exist, uh, in that plaque.
And if they did, they certainly don't correspond to the letters that Gemini is so confidently asserting that they do. The actual message was, your quest begins with particles who's close, which I suppose is a hello sort of a message. It's about the begin, right?
But, um, my kids, did you do this? So your kids won't use Gemini to Classes? I wanted them to.
Really good strategy really is I wanted them to get excited about dad's work for once in their lives. But obviously I shot myself in the foot 'cause it clearly didn't work. And Maybe even understand what you do a Little bit.
Yeah. Or understand where it can go wrong. Anyway, message, um, so did not work right despite the fact that Gemini gave me all the confidence in the world, that it understood the request perfectly, that it had sufficient data to answer the request.
And yes, in fact, here's your solution, right? This is what we call a hallucination. This is the questions.
This is a hallucination. When an AI model doesn't just tell you, I have no answer, or I don't know how to answer the question, it actually says, here's the answer and it's wrong. Which is way worse than saying, I have no answer, right?
This is my segue into rag. What is RAG and why is it useful? Why do we use it?
RAG stands for retrieval augmented generation. So this is a, a wrinkle that's used in generative AI models without rag. This is what happens.
A model gets trained on a static data set, and then you feed in the query. What does this message say? Right?
If the model has sufficient training data to give you an accurate response, then the response is correct. Green light, all systems go. If it does not have sufficient data to give you an accurate response, one of two things can happen.
Either it will say, I can't help you, which is a bummer. Or it will say, you bet, here you go and be way off base. Hello.
Which is much more of a bummer in the long term, right? What does rag do? Rag creates a step in between the query coming from your fingers on the phone and the model seeing it at all.
So when you enter, uh, a user query, if you're using a sort of a rag enabled, uh, inference setup, uh, your query will first go to an external data set, uh, where that data will be searched for anything that's relevant to your query. And then it will grab that information and your query is sent with that data from the rag data set together into the model, right? And the model still has its initial knowledge that it was trained on two months ago or whenever the model training was done, but it now has more up to date, more specific stuff from this dynamic data set.
Highly increasing the likelihood of a correct or an accurate response, right? So that's the beauty of rag at a very high level. These data sets, what are they?
They could be anything. It could be your corporate SharePoint, it could be your Outlook inbox, it could be Wikipedia, it could be news feeds, right? There's, there's any number of things you could connect, uh, your model two to get better answers.
And so that's where we're gonna go with STEAM here in just a second and talk about, um, some cool things that storage unlocks. Because like I mentioned, the problem is that enterprises are so excited about the possibilities of RAG that they want this orange circle to get bigger and bigger and bigger and bigger. And that gets very expensive very quickly, right?
That's the problem in pursuit of more valuable insights. We wanna use more of this cool new technology. Uh, but memory's awfully expensive, right?
So are there opportunities to do some things in storage instead with that Ste Graham, come on up. Thanks, ACE. Uh, Ste Graham, CEO of Metro ai.
And um, as, uh, ACE alluded to, we've been wanting to have a little bit of fun kind of testing the traditional paradigms of how we deploy large language models in full stack rag applications. And we'll show you a quick demo about that. We, we put together, it's gonna be a live demo too, which is always fun.
Um, what we're gonna demo is this particular stack that we're showing here. Um, one of the things that I'm super excited about on this, on this particular demo, and Ace alluded to it, is for, for most of us that work on application development in ai, and what we spend a lot of time doing is, is building AI agents as well as, you know, evaluating the, the latest state-of-the-art models and performance. We're gonna run a a 70 billion parameter model in this demo, but we're gonna do it in a single L 40 SGPU.
And if those of you that pay kind of attention to the memory constraints on that, you would go, oh, normally that'd be, if we're gonna run it Precision 16, which, you know, some people have really strong opinions about precision. Some AI researchers say, give me, you know, FP 16 or BF 16 or gimme death. Um, because they like high precision, because you get high accuracy and high precision, but you can quantize models too.
But in this context, we're gonna go take a high precision model and then we're gonna offload that model because we don't have the memory footprint in the GPU, uh, which is a common problem as we look at more and more state of the art models, you know, pick your favorite state-of-the-art, you know, model right now on the leaderboard. And, you know, many of 'em are massive and getting more big and getting bigger as well. 3 again, won't run natively on that single L 40 SGP.
So we're gonna try to offload it into SSDs, which you might ask like, why would you do that? Well, you know, if you have existing infrastructure and you're waiting for your next order of GPUs, or you're building out a new data center and you don't have A-A-G-P-U native infrastructure with the latest, you know, the latest, greatest GPUs with a great memory footprint and you want to use something now and deploy today, or make use of your prior capital investments, it kind of makes sense, right? Mm-hmm.
Um, so I also think it's super fun. Um, the hero the story though, I think, and, and the other really cool thing I think we're gonna show is, is we're also, you know, what we typically do when we run a RAG applications is we'll do the vector indexing. Um, and we'll do that in memory and it's kind of, there's a, there's a hierarchy navigable, you know, small world algo that does that, that we traditionally do.
And most leading vector databases, like in this case we're running BU DB does just that, but we'll actually offload that to SSDs too, which is kind of cool. Okay. So now we're gonna take a model that's, you know, theoretically not possible to run, you know, in that, in that GPU we're gonna offload the model parameters to the ssd, which is kind of cool.
And then we're gonna offload, you know, all that, uh, work as far as vector indexing as well. So kind of cool. Um, the other thing that's cool about this solution that we built, um, we're running a vision language model.
So we, you know, text data kind of small, so we don't want to, you know, text is not that interesting. It's too small. So what we're gonna do is like upload a video for the purposes of today.
It's, you know, in, in the attention span of, of the average human in the year 2025, which I hear is roughly 20 seconds. Um, we'll probably we'll do just a relatively small video, but we built this solution to process videos. 'cause you know, text data just isn't big enough.
Um, and so we've, we've done that too. What we're gonna do here, um, on the pipeline is we're actually gonna use a vision language model. In this case we're gonna run FI three.
I'm gonna run that on a different GPU. Um, so I'm gonna run that a different GPU and the FI three model we're gonna do, we're gonna do kind of transcription. We're gonna describe what's happening in those videos as well.
So that's what we'll do. That'll be fun. Um, and then we're using a, a, you know, an embedding model as well.
We gotta create those vector embeddings, another cri critical important, uh, piece. And we'll run that on another GPU. This server that we're running actually has four GPUs, but just for the purposes of this demo, we'll we'll use three out of four.
Um, so that's kind of the, the fun stuff in this demo. And, uh, I'm just gonna switch machines. Um, we we're actually gonna do is like, this is a kind of a traditional public safety and security use case, or it, those of us that have worked in this space in the past, we used to call it digital security and surveillance.
Um, and um, like, so the, the ask here, is it, you know, we're, we're processing a video. Um, this is more of a smart 3D application. So we're actually looking at like a street corner.
We're looking at traffic flow patterns and obviously humans don't want to sit around and monitor traffic flow patterns. There's a lot of traffic and there's, uh, not many humans. And so the vision language model and the, um, the language model will summarize that kind of traffic flow report, okay.
Is the particular use case and application that they'll be showing today. Okay. Um, as well.
And so, you know, just for fun, I've got like, this is kind of our ui. We've got the, the hardware set up here using Epic. Of course we're using a solid IPCI, gen five is fantastic for this particular use case.
Um, got, you know, a bunch of traditional DRAM memory as well. Um, the, the technique I didn't talk about, but, um, most, a lot a lot of us use, uh, deep speed, uh, for training. But there's this really cool feature where it has the capability to offload models to SSDs too.
So kind of hidden gem in deep speed as well. That's what we're using vision language model here. And you know, what we're really looking for is like when we run like the vector indexing workloads, is we're looking to see if, hey, can SSDs perform kind of roughly equivalent of that traditional, um, in-memory technique that's used.
And we're gonna use dis a NN to do that. So Disk n is a great project team, did a fantastic job, and they've done a ton of optimization, um, in the algorithms to essentially kind of catch up the SSD to the, um, in-memory performance you would traditionally, uh, use in, in this particular use case. So I've, I've already kind of preset this demo for time purposes.
We kind of do do a lot of fun stuff. Um, I preset this demo. I actually ran already LLM with GPU offload.
This demo actually doesn't let you run the LLM without GPU offload 'cause the system theoretically couldn't run it. Um, so I'm not, not allowed to do that. And if we get some more time, I might run it again.
Um, I just throw it through in a small video clip. And this is like a pedestrian crossing type intersection video clip. Um, and it's just easy in this ui.
It's like you can update labeled videos really easy as well. Um, and I'm just gonna kind of click begin analysis. And so this will take a little bit of time because right now we're gonna process the videos and we're gonna take, usually when you process videos, you take, you know, you clip images out of the video and then we're gonna store that in our Vector DB on, on the SSD.
Um, and then summarize in the context. This is, this is actually where our vision language model will come into place. This is where the PHI three based vision language model, they're describing what's happening in these video clips, in these images.
Um, so it's gonna do that. And then we're gonna generate, you know, our beddings. This is where we're gonna try to do it.
We're gonna generate our embeddings, um, with disc A and n with SSD offload as well. So that, that piece is gonna happen as we're generating embeddings. This is gonna take a little bit longer 'cause this is a little bit more of the labor intensive part as we pre-process all this stuff into the vector db.
And then a little bit later we'll start like querying insights. And so the telemetry on this system we're using like Prometheus for telemetry. It's gonna be a little bit, um, lagging the real time.
But you see as we go through this, as we go through this process, you'll see kind of the, the reads and writes to the SSDs will be, will be called over time as well. And it's going to kind of take a few seconds. So let me think about any other fun cool things.
The other performance metrics that we're gonna track on this, and this isn't gonna come out until we generate the report. A lot of people ask me like, when do you actually use the 70 billion parameter model? So the 70 billion parameter model is gonna call back into our vector DB and the embeddings and it's gonna generate the report.
And so this would be like a smart city pedestrian traffic flow report, our drivers driving correctly as they go through this intersection. And that's where the LLM comes into play. That's gonna show up a little bit later.
And then we're gonna measure not only the, you know, throughput and tokens per second, which, because we gotta, we're fitting a really big model on that one L 40 SGPU and we're, we're offloading a ton of parameters on the SSD. It's not gonna be blazing fast, but again, you theoretically couldn't do this without this technique. And then we'll also show like the, the GPU, uh, memory usage as well, which will be kind of fun.
'cause typically a 70 billion parameter model, you're gonna see 1 1 40 GB for that as required and not available. Uh, yeah. Questions while we, Yeah.
So question while this is running. Yeah. Are you ingesting any part of the audio stream or running any audio models in parallel?
No, we're not. But uh, so there's some, uh, legal requirements in different states mm-hmm. Where you legally can't record audio in public settings.
So you do have kind of that, that requirement as well. Um, I'm not quite, quite up to date on the particular states that, that, that, um, enable that, but we are in the state of California, which, unless policy has changed since I moved outta California, you're not able to, uh, take audio in public settings without consent. So yeah, we, we don't typically do that.
Um, although, you know, occasionally for surveillance use cases, uh, that will be a requirement in an RFP or RFQ and then you have to, Well, could it be in an enterprise setting? That's what I Yeah, of course. You know, as you guys all know, when we, when we become W2 employees in enterprises, we do, you're consenting we do consent away a lot of privacy for that, that great W2 paycheck that we receive.
So, um, yeah, and I think we, you know, I think everybody should be aware that there's a lot of really cool monitoring software, um, as far as employee monitoring that I don't participate in Go, go ahead. So if I rephrase in the enterprise setting, we could ingest audio and run an audio model. Yeah, I think, I think in theory, but to, based on the code of conduct and the agreements that you can technically not ethically or Yeah, Yeah, technically yeah, Technically of course.
Sure. Yes. We're an enterprise based in Texas.
You're good to go. Yeah. And what I, what I would do like, and it would take me, you know, I could throw it on the extra GP that were run it for the, uh, purpose of the scenario.
I'd go grab a whisper model, um, you know, open ai gr made, made a great whisper model, which they do open source and is one of the industry leading, you know, models, um, for, for audio. And I'd go grab the whisper model, I'd throw it onto this demo and we could have it right here, right now. It'd be no problem to do something like that, uh, if thinks is appropriate.
And you can see the, you'll see like the different GPUs 'cause like GPU one, as, as we're running that PHI three model was the first model we were running where we're taking those, those image clips from the video and then we're transcribing them. Um, and then we'll, we'll start using the other, the other GPUs for the embeddings piece and the report generation piece. And I am trying to buy myself some time.
'cause this is a bit of, a bit of a longer, longer demo in in most, uh, scenarios. I'm hoping we'll get some outputs here shortly. So if I can Just put a Yeah, a fine point on it.
I'm a lot less technical than steam, right? So it helps me to kinda hear the highlights a couple times. But what we're doing here is really Two things.
We're offloading model weights to storage, right? What does that buy you? That buys you the ability to run complex models on hardware?
You couldn't run it on before. So without this switch, this demo doesn't exist. It's not possible to run the 70 billion parameter model that we're running in step four on the hardware that we're running it on.
It's a binary not possibly before. Now it is. Right.
And Just to like, to look at numerically, I think like of the 70 billion parameters, like actively, we're only running like 300 million actively in, you know, in the GPU and then we, we transition back between the GPU and the SSD for the other panels. Again, like something quite unique. Um, right, Totally.
The second part is, is the other purple box, right? And so that's where we're actually taking some of that rag data we talked about previously. The big orange circle, if it grows infinitely large, it very quickly becomes unmanageable in memory, right?
And so what we're actually doing is we're relocating some of that data from memory where it's typically been in a lot of rag use cases to SSD and we're, we're accessing it directly from there in real time is what you're seeing. So these are, these are indepe. What what we'll do, and I don't wanna steal your thunder here, but we have like a GitHub repo that we intend to make public, and you'll all have the ability to kick the tires on this if you want.
But you can flip these switches independently and see kind of before and after results. We've done a, a pre-run here, uh, where we've kept it in memory and as I point to that, it zeroes out as he, so The cool cool thing just happened, um, is in memory, this is shocking if you saw that it was 299 queries per second mm-hmm. By pre-run.
And then we offload to SSD and we got 320 queries per second. Okay. So that's really, really unique and it's counterintuitive, right?
Yeah. Yeah. And so, you know, three things that you're thinking about there is one, the team at DIS and did a fantastic job on the algorithms.
Two, you know, the, the way that you know, HNNS works, um, is we like, or HNSW SW is small world by the way. Um, and the small world is that DRAM footprint. So there are some memory brown constraints when we use end memory.
And then the, the other component too, um, that you think about is like these great PCIE Gen five drives, um, you know, from, so from soddy are, are really great. So the evolution and performance of SSDs enable this as well because, uh, as we test and we'll share some results on gen on gen, but we get pretty dramatic performances in gen on gen, but super cool. So now we got better performance than in memory.
We loaded a model that was not possible to load. So I have two really cool things, um, and then like this generates this report. I can just download this port, which I pre downloaded, um, as well.
And so this is where you see like here's the, the PHI vision language model just describing, hey, this white vans driving through this intersection, um, were they following the right of way rules in this particular scenario? So this is all AI generated and the, the, the 70 billion parameter model is the one that writes up the report, but the vector embeddings were, and the vector database was based on that, that original five vision language model that, that put that together as well. So this is all kind of AI generated report.
Why, why are we talking? So yeah. Very cool.
Now we're gonna switch back and we'll look at kind like, let's test this ve like this vector indexing with SSDs to the limit and look at what performance we actually see versus a real time short demo. Just brief question. Yeah.
In terms of the video input and the embeddings that were generated, the HNSW index, was it ginormous by comparison? I was just curious if, you know, how, how much data are we really talking about for the video itself plus the HNSW index roughly, if you know, Um, yeah. Well the video is really small for the purpose of this demo.
It's like 50 M 50 MB. Okay. And then, you know, you can see like the, like HNSW we use as much of the memory footprint as they have in dram, which, um, I didn't, I don't think I have telemetry on their utilization of that, but it was like 380 four-ish GDS of DRAM memory footprint for that as well.
So Yeah, seven to one. Wow. Um, and then so the kind of, again, like just hitting this count counterintuitive point that we were, we were kind of all experiencing is like, wow, you know, with this a NN with this offloading to SSDs, like, and we actually start testing, this is a, a cohere database set.
Um, we start testing with Vector DB bench, which is a great tool to test different vector databases on performance. In this case we're using viss, viss is great, but pick your ve your favorite Vector DB as well and you can see like in memory performance and then you can see, okay, higher queries per second at this 1 million dataset size as we go up, it tends to even out. So in memory we'll kind of perform, um, um, a little bit, you know, a little bit lower as you, you go over here we get this great performance on 10 mil, um, as well.
So the, the other question I think is like, you somehow look at like, okay, so what's the overall DRAM footprint impact? And maybe this is kind of to our prior question is, you know, performance of DRAM usage actually dramatically lowers at that system level as well. So there's some, there's some great cost savings as you would expect if you're using less DRAM and using SSDs for offloading as, yep.
What has to happen in the file system or object storage system to enable this to go on? Is there any software coding? I think the, I mean if we just go back to the kind of the typical RAG framework, I think the, you know, the one thing that I would allude to is miskin missing here, like overtly is, you know, it's just kind of, plus here you gotta add your disc a and n algo capability, um, you know, two year vector database to make this happen.
But it's not gonna be a huge lift. Um, the team at discount end's done a really great job. They've, you know, and they've got a great, you know, release schedule as well.
So, um, just enabling capabilities. So, um, cost savings you can see is there. The other cool thing, you know, the real question is, you know, okay, how are we doing, you know, on recall performance?
Are we as accurate? Are we making some accuracy trade off that I'm not telling you about? And so you're getting about the same level recall as well in both case scenarios.
I, I, I think maybe Kimberly had more to that question. Oh, sorry. 'cause I'm thinking that if it's a, if it's a Linux system, I don't know if it's a Linux or a Windows, it's a Linux, but it'll, it'll probably just use a raw disc write, read, and write because it's NVME.
Okay. Okay. I don't think there's a file system on it.
Yeah. Okay. Thanks.
Yeah, but recall performance, the, the one thing I will say though, the trade off you will make is the, the up the upfront, there'll be a little additional indexing time that happens as well. So all the metrics look great. There's a little bit of upfront indexing, uh, time and trade off as well.
Um, the other thing that we looked at, I think, you know, thanks to the, the team at Sola time, we, we, we got the latest generation drives. We looked at last generation, um, and, you know, got some great gen on gen performance pops, um, you know, 30% as well in this particular scenario, which is, which was really fun. So appreciate, appreciate that opportunity.
So you talked a lot about the great work others did to make this possible, but Metro was the driving force behind the, putting all these pieces together and we've really enjoyed our relationship with them so far. Um, so just to wrap this up and then I'll hand it back to Scott. Uh, we do have a white paper that we'll be making available to you that kind of goes through the methodology, goes through, uh, some of the key findings that's seen, highlighted.
You'll be able to peruse that at your leisure. And then, as I mentioned before, uh, the repo as well, that has all the, the, the parts, the guidance on how to put it all together. If you're interested in, in playing around with the solution yourself, that will all be available to you as well.
So is there any part of the solution you're holding back or will it all be available on the, on the repo and is it It's all open source and Yeah. The intent is just to, uh, make the case that storage can surprise some people in the, in the inference space. Right.
Particularly in data intensive inference and, and to encourage folks to go try it out. Cool.