MLCommons MLPerf Storage
Presented by David Kanter, Executive Director, MLCommons live in San Jose, California on January 29, 2025 as part of AI Field Day 6. Watch the entire presentation at https://techfieldday.com/appearance/ml-commons-presents-at-ai-field-day-6/ or visit https://TechFieldDay.com/event/aifd6/ or https://MLCommons.org for more information.
Transcript
Next, we're gonna talk about ML PERF storage. Uh, and, uh, I'm David Canter. I'm, uh, one of the founders and, and the head of ML Perf, and I'm with Curtis Anderson.
I've, uh, one of the co-chairs of the ML perf Storage working group. Yeah. And so this, this is a community effort.
Um, and so I, I'm, I'm gonna just give one brief introduction, hand it over to Curtis to talk about the work of the community. But, you know, really the motivation for ML PERF storage, I, I is up in this slide here. Right?
And so this is data from meta talking about the growth in the use of storage for AI training. And what you can see is, uh, over eight quarters, uh, they essentially doubled the number, the amount of storage they needed, the amount of data they needed to store. But the, the actual ingress bandwidth, or sorry, I guess egress from storage ingress into compute, has, uh, grown even faster than that, right?
And so storage is truly critical for our big data models. So why do we wanna measure storage? Curtis, why don't you take it away?
Thank you. Um, so you need to train a, to make use of ai. You need, clearly, you need accelerators to do the training.
You need networks to, to communicate the data from storage into the, the accelerators. There's a bunch of different pieces that go into that. So full system performance says I need to know what storage is capable of from the point of view of the benchmark we're trying to produce, uh, relative comparison for, uh, the, the people who want to acquire storage for use in a, in an AI environment.
But we're also trying to help the people who build the infrastructure, like PyTorch, I TensorFlow, those kinds of frameworks to optimize for, uh, storage systems. Um, as you can imagine from the, uh, uh, the majors that are training the foundation models, there's an enormous amount of data and enormous number of accelerators. And that means an enormous amount of bandwidth and capacity to man for the storage system to manage.
Um, so it's in about an informed decision for the purchasers. It's about the researchers who are trying to build better storage systems. And it's about the, the product vendors who are building storage systems who want to optimize for the, uh, the AI workloads.
Hey, Chris, I mean, most of the storage for, uh, AI training is typically, uh, direct access storage or something like that. It's not necessarily network storage. Actually, the vast majority is network storage, because the, the fundamental thing that that causes fits for the storage people is the data access pattern has to be random.
You can't feed this the data, you can't show all of your cat pictures and then all of your dog pictures because the neural network won't learn that way. You have to intersperse them randomly. And so you go through one epic feeding the data through, and then the next epic has to be a different pattern.
And so that, that defeats all of the normal caching techniques that storage soft storage systems tend to use. And so, and in general, the amount of data that most people have, uh, is much too large for any on-prem, or, sorry, in node storage solution. So generally, most storage is, uh, actually external outside the nodes.
I, I guess I agree with respect to capacity, but the actual, the, the, the data coming into the GPU, you think that's typically off the network. Yeah. Mm-hmm.
And that's, that's one of the challenges, right? The, uh, the storage system providers are working hard to figure out how to make that work better than you might. Oh, well, why would that work?
Well, well, but yet that's one of the requirements. And so, I mean, like GPU Direct and things of that nature, is that Sure. Yeah.
Yeah. 'cause we're not talking about like traditional network attached storage necessarily. I mean, I think we're talking about all sorts of exciting new solutions for network storage.
So yeah, we've got GPU direct, uh, obviously we've got a lot of parallel NFS going on in there. We've got a lot of object storage, uh, scalable, you know, scale out, uh, object storage solutions being used. Um, I've even heard about, uh, things like S3 over RDMA and yes.
You know, obviously Infin Band is in there, so, right. Yeah. It's not like traditional network storage, but like, I mean, the thing that, that kills me when I see graphs like this is, you know, we, we've never designed storage systems, right?
I mean, we never designed these things to be able to have twice as much IO as they were designed for. Yes. You know, per per unit volume.
We've never designed them to be able to do, to handle this kind of randomization or this kind of parallel access. It's just not what we were making storage for generally. But I think like folks like you were actually making storage like that, like with Panasas and so on, where, you know, storage for h HPC is more applicable to what's happening here with ML than traditional corporate nas, right?
The, I don't wanna put words into your mouth. Yeah, no, no, no. That's, That's all fair.
That's a good summary. Yeah. But I totally put words in your mouth.
So, um, storage is a, a slow moving business in terms of architectures, right? The, the, the quality requirements are so high. If you lose data, you're toast.
Right? And so the quality requirements are high enough that people don't take large risks unless they need to. And so that's what happening with AI is the, the industry is being forced to take larger risks, do more innovation, increase the speed of innovation in order to meet the needs, right?
And so that's, that's one of the interesting, exciting things about it is, uh, storage wouldn't change unless there was a driver that forced us to change, right? Because of the quality that we have to, to do. That's actually a really interesting point, because things have changed, like what we wanna do with the data and how we're going to process the data.
So the storage, the, the access to storage, stor access to data within the storage systems has to change. Yes. With everything else.
We've just come to that point, Right? Yeah. The, the, uh, um, I mean, tapes didn't die hard.
Drives will never die. Uh, stands will never die. The the, they'll shrink, uh, and they'll, they'll be fit for a particular set of use cases.
And so now AI is taking over. Uh, it's on everybody, every storage vendor or supplier's, uh, uh, top of the list saying, okay, how do I play in this exploding market? Right?
And so what do I need to do to my product to best serve that the needs of that market? Of course, this is for training though, but there's also, yeah. Different D and other, you know, rag and other elements of storage as well.
They're gonna have completely different access patterns. Absolutely. That's completely different.
So You're not addressing that with ml, with, with storage performance yet, yet? We are in the process. Okay.
So, um, let me, uh, oh, okay. 0 of the benchmark. And those each apply a different workload to the storage from each other.
Some of those run under PyTorch, some run under TensorFlow. And the TensorFlow versus PyTorch are also very different work, uh, uh, workloads applied to storage. And so the training, Training versions, right?
Those are training versions. Yes. Now we're in the process of working through a, a rag pipeline.
So, um, a large language model with a vector database underneath and the database, it hasn't yet another different workload. The challenge with inference is, in general, there's not much data involved in inference except for possibly in the rag pipeline. It depends, right?
What your point of context is, there is a lot of data from some traditional client application standards. Yes. But not much data compared to training.
Not much compared to training, but, um, and it's, it's a, a single ton, a one-off, right? You, yeah. You say, okay, here's an X-ray analyzer, here's The data, right?
Yeah. And, and so the, uh, um, Index too Rag it's index and Rag is Yes. There's a lot of sort of now in context rag.
Yes. Like get all of the chunks compared to rag, which that changes stored from traditional rag, even to cag, what they're calling conec. Mm-hmm.
Anyway, so yeah. So I'd love to hear all, but that I think would be incredibly interesting. It It is interesting.
0 with a benchmark. Okay. Um, the, uh, the, our, our best guess at the beginning was that, uh, uh, normal inferencing wasn't interesting from a storage perspective, wasn't challenging to the, to the solutions out there.
Uh, vector database we're getting farther into it. And, uh, in, in a chat kind of, uh, context, the amount of compute time is very large compared to the amount of IO time to do the rag lookup. Right.
Um, but things are changing constantly, right? And so we're, we're tracking we're Rag C and all that. Yeah.
Right. So in timelines, uh, we saw the prior list of like the benchmarks that started and when they came before, and the ones that were deprecated and became the dash lines that Ray didn't like. With within this world, uh, we've heard about storage arrays that have been created, and solutions and storage architectures have been created.
They're probably chasing just enterprise workloads. Go chase the biggest market, you know, optimize for this, right. Synthetic benchmarks that made you as a vendor look better or worse compared to you.
Absolutely. Where would you say where maturity wise, I know you're saying we're just getting started, but maturity wise, uh, are you starting to see specific vendor, uh, uh, adoption of these in their marketing collateral or stories they're telling to the market about how they're performing using this benchmark? Yes.
Okay. 5 release, which is sort of the, the traditional proof of concept thing that the ML comment does. 0.
No, that was the first one. 5 to say, okay, yeah, this, I, we think we can make a real benchmark outta this. Okay.
5, we had four submitters and 20, uh, submissions, or maybe it was 30, and now, uh, then we went to 14 submitters and 105 or 130 submissions, right? So the, the ramp we think is very, uh, steep. I'm getting a lot more interest, uh, in, uh, organizations wanting to join the working group in order to, uh, to post this, right?
Um, AI is, is an enormous market, and it's the new hot markets where all the investment is going. And so, um, the way I describe it, uh, 10 years ago, every CIO said, what's my cloud story? Now, every CIO and every other C-level executive saying, what's my AI story?
It's a much bigger wave, right? And so all the storage renders will respond, right? So, yeah, I'm expecting there the, uh, um, that they, that the growth of the use of the benchmark will continue, right?
And we're also intending to, to help the, the product organizations optimize for this workload. But that was one of the things that we had to do on our zero to five is, well, what is the, the storage workload that comes from AI training? And it's very different from most other things.
It's not a database or a streaming workload or anything else. And there was a, there was a trend years ago that if you were a storage engineer, your career is over. You need to Oh, yeah.
You need to become a cloud engineer because cloud Oh, yes. So are, are you, do you anticipate a similar, like if you're a storage, if you're still a storage engineer and you didn't quite make it to being a cloud engineer, just forget that. Go straight to AI engineer, you know?
Yes. Maybe it's outta your wheelhouse, I don't know. But what's your thinking on these kinds of, like, you know, your career, your title is dead, you know?
Oh, No, I, it's moved. I, I started storage 40 years ago, and I, my kids go into computer science and I says, no, no, don't do storage. Don't storage.
It's over. And now it's not again, which is very cool. Cry.
It's awful. So, so let me describe a little bit about the, the structure of the benchmark, and then we'll go into what the workload looks like. So in the, the normal environment, you've got your data set on storage, you've got, uh, system memory, uh, you've got the accelerator, um, uh, on the far right.
And then there's some, uh, online pre-processing. So let me define what that is from a storage perspective, um, I've got a certain amount of, I'm gonna use cats and dog images just because it's easy, right? I got a bunch of pictures of cats, but more data is generally better.
So I can take that picture of, pull a cat picture in, and then I can rotate it 15 degrees. Still a cat now, but the, but the neural network sees this as a different image. I can change the color palette a little bit, right?
So that's data pre-processing. Uh, it multiplies the effective amount of data, and I get a higher quality resolution of, of, uh, predictive ability on, on cats. So that's the, the structure that we're trying to model, what we're doing, because we're only interested in the storage piece, is we've got the storage, um, it comes into system memory.
We don't need the accelerator. We just put a sleep in there saying, well, the computation time for this type of data by this accelerator on this workload is 327 milliseconds per batch. And so we just sleep because, and we are just pulling the data into memory, and we're ignoring all the other things that that's sufficient to stress the storage out.
Why burn down the rainforest when you don't need to do, you Don't abs. Absolutely. Well, and, and as you know, the, uh, uh, accelerators are expensive, and so the storage vendors don't really wanna buy 150 of those if they don't have to.
Right. Have, have any of the vendors of those, uh, uh, cards that do that, um, have they tried to influence, like No, no. Really we need to bring in actual, it needs to be real traffic, not sympathetic traffic.
Any, anything you've heard in that space? No. The, the, the suppliers of, of accelerators.
I, I say accelerators because I'm trying not to, you know, I, I really want the, the non NVIDIA people to also participate, right? Yeah. Uh, we have some Nvidia Nvidia people in the working Group and, and there's all kinds of custom silicon that's going on Yes.
In other places too, so, okay. Very good. So, uh, significant variables, and then I'll go through, uh, there's a couple more I need to add to this.
The framework, high torch versus TensorFlow, high torch. Generally speaking, if you got images, for example, cats and dogs, uh, you're gonna have, uh, a direct retrie full of files. Every file is a single image.
Um, with TensorFlow, it's much more like a tar file for people to know storage, right? So you put a whole bunch of images, like a thousand images into a single file from the storage perspective to out open to, to get a single data point, a single image, I have to do some directory lookups and a file open, and a whole bunch of traversal metadata operations. TensorFlow, I just read the next image out of the already open file.
So, vastly different workloads from the point of view of the, the metadata and data operations of a storage system, storage network, how we're connected. Um, the, we've had a, a, a several submitters, uh, realize a little late that lossless networks were required. So you get a TCP packet dropout, and then all of a sudden you perform storage performance goods to the floor because you've, you're waiting for retransmission.
And so there's, there's nuances that are interesting about the storage, networking, the solution, uh, hardware and software, um, uh, how it works internally. Is it a an HPC class parallel file system? Uh, is it, uh, you know, lots of different things.
0 that had had, uh, an SSD vendor Yeah. Submit SSD versions of submissions for ML birth storage. So, yep.
Yeah, I got three minutes left. So, um, gotta keep moving. It's, it's valid, right?
We are, we are trying to, to bridge that dynamic or range from single SSD vendor all the way up to a thousand node HPC submissions. 'cause we have Argon National Labs submitted, you know, Intense Pel Neutron source Yes. Type of training, um, distributed training, the, a single host, we started with that.
Then we may move to what's called data parallel training. So, uh, data parallel, the way I describe it as a storage guy is, uh, every accelerator has the complete model, and you're gonna pass a shard of the data across that model and the periodically exchange weights model parallelism as opposed to data parallelism. Um, every GPU or every accelerator has a shard of the, the model, and you're gonna pass all the data past every one of them, and then you exchange some weights.
So we're doing data parallel caching. So here are the, uh, uh, workloads. Uh, there are different categories like image segmentation, image classification, scientific cosmology.
We're looking at a rag pipeline. We're looking at, uh, checkpointing for, for large language model training, you have to take checkpoints because the training will last for, you know, a month or two, and you can't afford to take a, uh, an outage in the middle and lose all of that. Um, so I wanted to go through, oh, well, storage 13th organizations on how results here, the organizations that submitted.
So I'll leave that up for a second. Well, the, some structure about the, the, uh, um, the benchmark proper. So the thing we wanted to say is, or we were trying to demonstrate is the storage is not gonna slow down the compute.
So that means that, uh, you, the, there's a request from the, uh, the framework say, I need these data points. I need to pull them into system memory before the current computation is finished in the GPU. Otherwise, the GPU goes idle.
Those things are expensive. Nobody wants those to go idle. So we have a metric of accelerator utilization, which is AU percent.
Um, is that we have to keep the, the accelerator 95% busy. So if you are, you have excess latency on some IO operations that counts against you, too many of those, and you don't have a valid submission any longer, right? It's keeping the thing busy.
Is, is the sort of the core metric. Can you keep up with the demand? Uh, we also have two categories, open and closed.
Closed is, uh, apples to apples comparison. It's virtually identical workload against from all submitters. Uh, no innovation, no changes.
Here's the baseline. Can you do it open? On the other hand is for the, the, the submitter to say, uh, here's my close number.
But now, if you are willing to make these changes in your environment, look what I can do for you. Right? So they're not comparable against each other, but it's a place for the innovation to happen.
Right? This, uh, training has a similar structure, a little bit different with open and closed. Oh, here we go.
Open and closed. 0. Um, uh, only one organization chose to, uh, to do open mostly because the, uh, uh, almost all of these were new.
5. And so they were all, Are you working with any rag vendors right now? And if you are, can you mention who they are?
We haven't yet decided, um, that the, the current proposal is that we pick a representative rag solution and a vector database, and then that's the standard workload that everyone applies in a closed. So they vary, right? You've got your Postgres, you've got your doctor.
They do. And so that's, you've got Your PAT files, so it's, anyway, that's why I'd asked the question, right. Currently we're, we're proposing to, to have one.
Okay. Right. Because this is a crawl, walk, run structure, uh, key insights.
And so we're down to the last two slides, 48 seconds, I can do it. Explosion a micro, uh, driving a wave. Innovation said that before, uh, stability, uh, lots of different architectures out there are coming out because people are saying, oh, uh, the, I had this odd little architecture that I, that wasn't getting a lot of market traction.
But look, I can do a really good job on this benchmark. I, I think I have a shot at, at large growth opportunity. Um, yes.
Listening results, it's important to know that the, uh, um, for storage vendors, this is new territory. And so they're doing a lot of work, uh, to understand what the workload means, what is what, uh, where is my architecture stressed by this workload. And, uh, just read a single host scale of a different frameworks.
Yeah. So there's a lot of challenges. It's That last Bullet I think that stands out.
You know, it's, it is, you have, you know, you go to your old spec websites and you work for, or you, or you're purchasing that next box from that next storage vendor. And it seems like you're, you're, you're, you're going with some arbitrary thing that someone says they've optimized for, right. In this one.
Um, because it is community. Mm-hmm. Because it is transparent and traceable.
Right. Um, do you, do you think people are gonna make better IT purchasing decisions because of the existence of this? Yes.
Because the, the, the data scientists that rightly they are focused on their neural network architecture, they don't care about storage and networking, all the rest of that, that because the innovation is coming with the neural network architecture. Yeah. Right.
That's the value. And so they say, ah, whatever. Fix me up with something that's gonna work for me.
That's the measure of what's gonna work for them. Right. And so that's our goal at least.
And we think it's working. Yeah. So, you know, the thing I find interesting, I'm more of a on the data engineering side Okay.
And database administration side, and, you know, all the stuff that you have to do to be, not the GPUs, but the developers. Ah, I just find it fascinating, uh, just popped into my head that, you know, even 10 years ago we were freaking out if A CPU was overt tasking. Mm-hmm.
And now we're worried about GPUs that have to be fed all the time. Yes. It's just interesting to see how things have turned inside out.
Yes. For people that are in the industry, in the trenches trying to figure this stuff out. Do you see that a standard, like your standard, you know, for auditing for accuracy would be useful below the CTO or C level for people in the data engineering and, uh, database administration space and even application development space?
Absolutely. The uh, um, uh, the storage system architecture is one piece of the puzzle, right? The, uh, right.
And so in order to, uh, if you don't measure it, you're not gonna improve it, right? The, uh, right. We we're measuring the architectural performance.
Here's a, a, a set of gear and, uh, you know, how, what performance you're getting out of it in terms of, uh, batches per minute, right? Or whatever, right? The, the, the AI metrics, right?
And so that's driving innovation there, but the vector databases are also brand new. And so they're going to go through their own shakeout in terms of who's got the better architecture and the better, uh, performance, Right. And even things like, uh, different vector indexes.
Absolutely. Performance, approximate, exact searches, all kinds of things like that. Right.
Good point. Cool. Alright.
If I could just make a comment. Go ahead. Um, so first of all, this fascinating.
I had heard about your work before. Um, I was strangely, particularly grabbed by your discussion about PyTorch versus TensorFlow. I think there is such a massive imbalance in the way people think about ai.
It's like all, all GPUs, that's everything. And maybe this model and this 3 billion, a trillion, things like this when there's so many really critical, important considerations, right? Like storage, like file systems, like things like this.
Yes. Um, and, and, and I say that in a sense of, you know, many of us talk to other people as we, and we, we have to nudge and educate and, and just make the discussion much broader. Oh yes.
Than, than I really think it is with ai. Uh, because these considerations are, are just paramount and Oh yeah. Completely clog up all this wonderful process'.
Supposed to be totally, I mean, David Candice said it earlier. That's a, it's a whole system problem, but it's a whole system problem to, to achieve the, uh, the results that people want to keep. The accelerators busy.
You need networking and storage and orchestration and on and on and on in order to make the thing work, right? To, to get the best efficiency. I mean, the, at the price of Nvidia accelerators and GPUs, right?
That's a resource. You really don't want to be idle.