MLCommons and MLPerf – An Introduction
MLCommons is a non-profit industry consortium dedicated to improving AI for everyone by focusing on accuracy, safety, speed, and power efficiency. The organization boasts over 125 members across six continents and leverages community participation to achieve its goals. A key project is MLPerf, an open industry standard benchmark suite for measuring the performance and efficiency of AI systems, providing a common framework for comparison and progress tracking. This transparency fosters collaboration among researchers, vendors, and customers, driving innovation and preventing inflated claims.
The presentation highlights the crucial relationship between big data, big models, and big compute in achieving AI breakthroughs. A key chart illustrates how AI model performance significantly improves with increased data, but eventually plateaus. This necessitates larger models and more powerful computing resources, leading to an insatiable demand for compute power. MLPerf benchmarks help navigate this landscape by providing a standardized method of measuring performance across various factors including hardware, algorithms, software optimization, and scale, ensuring that improvements are verifiable and reproducible.
MLPerf offers a range of benchmarks covering diverse AI applications, including training, inference (data center, edge, mobile, tiny, and automotive), storage, and client systems. The benchmarks are designed to be representative of real-world use cases and are regularly updated to reflect technological advancements and evolving industry practices. While acknowledging the limitations of any benchmark, the presenter emphasizes MLPerf’s commitment to transparency and accountability through open-source results, peer review, and audits, ensuring that reported results are not merely flukes but can be validated and replicated. This approach promotes a collaborative, data-driven approach to developing more efficient and impactful AI solutions.
Presented by David Kanter, Executive Director, MLCommons live in San Jose, California on January 29, 2025 as part of AI Field Day 6. Watch the entire presentation at https://techfieldday.com/appearance/ml-commons-presents-at-ai-field-day-6/ or visit https://TechFieldDay.com/event/aifd6/ or https://MLCommons.org for more information.
Transcript
So first of all, good morning. Thank you for coming to join us. Uh, thank you for listening to me speak and giving me an opportunity to sort of share, uh, the mission of ML Commons and, and what we aim to do with the ML perf, uh, family of Benchmark.
So, this is an overview of my presentation. We will first start with, uh, the organization and benchmarks, and then do a deep dive into L Perf client and storage. Uh, ML Commons is a nonprofit.
We're an industry consortium. We are focused on making AI better for everyone. And, uh, as you can see on the right, or, uh, what that means is we want to make AI more accurate, safer, faster, and, and more power efficient, right?
Every day we are seeing in the news, uh, you know, this week in particular, as, as was discussed earlier this morning, uh, right, the excitement of, uh, you know, getting a new AI model that has, uh, faster capabilities or better capabilities and how, you know, this all promises to transform society in, in remarkable ways, whether it's, uh, you know, the self-driving cars that are buzzing around San Francisco, uh, that I get to see every day or better AI for medicine or, uh, you know, better translation. Um, ML Commons is an organization, uh, that is, we're powered by our members and by the broader community. We have over a hundred members spread all across the globe, uh, over 125.
In fact, we're on six out of the seven continents. Um, if you know of any IT companies, any folks active in AI and Antarctica, you know, uh, seven OUTTA seven is always better than six, outta seven. So, uh, please reach out, you know, maybe someone at McMurdo Station, uh, needs to, uh, get some systems for, uh, you know, one of our premier projects is ML perf.
Uh, we've got, you know, tens of thousands of performance results. Uh, ML PERF is the open industry standard for measuring the performance and efficiency of ai, and that's what we'll be talking about today. So, I guess to start and motivate, uh, many of you in the audience may be, uh, if I can be forgiven, professional skeptics, right?
And so I see some nods. That's good. That's good.
Yes. Right? So, so why do we want, and why are we building benchmarks as a way to make AI better?
Like, what is that gonna do for us? And, you know, I like to go back to the, uh, uh, aphorism of Peter Drucker up here, that if you can measure things, you get improved, right? And, and we've seen this when we're, we're writing online, right?
You can ab test your headlines and see what is attention grabbing. You know, does adding images work or does it not? What should we be doing?
Uh, do we need a new screen? Has screeny the first, you know, maybe needs to be a bit better, right? So once we start measuring these things, we get to improve them.
And, and, you know, this is really the image that to me captures the, the real goal of benchmarking, which is we're in this as a community, as, as customers, as researchers, as buyers, as sellers, all working to make, uh, AI more capable, more efficient, uh, and we go farther together. If you've ever gone canoeing or kayaking or in crew, when you get the whole team in harmony, you just fly across the water at incredible speed. And that's our goal.
How do we bring the community together and align on what does it mean to be better? Right? That's the point of a benchmark.
Uh, right? It gives us a true north to guide towards. Um, and we want them to be, in order for them to drive this trans this progress, right?
To help us get forward, right? We want we transparency, and we want to be able to, uh, again, move faster. And, and a benchmark, no benchmark is perfect, right?
Uh, again, for, for everyone in the audience, right? We're, we're trying to get to North, but you know, if we're five, 10 degrees off, that's okay, because we're all going together. We'll, course correct and we'll go more rapidly.
So the reason behind this, uh, this is what I like to call sort of my, my TED talk slides. Um, there was a great paper from Baidu, you know, may last decade, uh, that really talks about why we want big data for ai, and this drives into our needs for compute. And so really what the, what they did is they looked across many different mo modalities of ai.
And what they found is that when you don't have enough data, AI isn't actually that useful. Just talk to an expert, make a good guess. But what this chart shows is on the on, on the x axis, you've got the dataset size on the y axis, you have error, lower error is better, right?
And so you start out in this small data region where the error is high 'cause you don't have enough data, but when you start getting enough data, the error starts dropping a lot, the accuracy goes up tremendously. And somewhere along that curve is, you know, going from lane warning to lane following, to staying in the lane to start following the car ahead of you. Uh, it goes from, you know, a sort of jokingly good translation that you never really want to show to, to everyone else, to, you know, I'm on the streets of, uh, uh, Singapore and I just, you know, can use Google Translate, right?
And those are all gonna be points along those curve. And so as we drive up the amount of data, we get new capabilities that are, you know, truly amazing. You know, David, What's the proof for irreducible error?
Uh, well, the proof for reproducible error, um, I, I, I guess in theory, maybe there is no irreducible error, but I'm gonna claim that no one can see the future. Like no person, no ai. And that's in some ways what that's saying, right?
Which is, it doesn't matter how much data we feed into a system, right? There's always some things you can't predict. I mean, there's obviously a limiting factor here as, as you get more and more data, there's less and less improvement.
But that doesn't mean that Well, right? At some point it's gonna saturate. And that's really what this is saying is, you know, but, but we're clearly in this power law region, right?
Like everyone, you've, many of you may have heard of the scaling laws that, uh, OpenAI is popularized about how we build bigger and bigger networks with more and more data that's in the power law region. So the, the consequence of this is that in order to learn from all that data, we need bigger models. So this is some data from Rise Lab, again, time on the X axis, and this size of the networks is on the Y.
And what you can see, and this data is a little, a few years old, but what you can see is it's just up into the right, right? We're seeing transformer based models growing at 250 X every two years, uh, right? And the transformer is the, you know, the fundamental building block that we see behind things like chat, GPT, uh, deep seek, uh, et cetera.
And so what that means is our compute demands are growing even faster, right? The compute is in some ways, both the product of data and model size. You make the model bigger, you now need to learn from more data.
And this is even growing even faster. Again, you know, not 250 x, it's 750 x. So the, the need for compute is, uh, tremendous and insatiable is really the point I'm making.
And, and so you put all this together and, and sort of, here's what I like to call the fundamental theorem of machine learning is big data plus a big model, plus big compute is innovation, right? When we put this all together, right? We get things like better cancer detection cars that are safer on the road safer than me, safer than, than all of us.
And I, you know, I I, I look forward to that future. Ultimately, machine learning is a full system problem. So, uh, and we're gonna get to talk about some components of that system today and particular, you know, certain compute and storage.
But when you look at training these models and, and increasingly doing inference, uh, there's so many factors that come into play. And so that's both good and bad, right? It's, it makes building a benchmark hard, but it also means that we have so many different levers to pull to improve performance, right?
There's the data itself. If you, sometimes it turns out if you sort the data cleverly, you can boost performance by 20 30%. There's the silicon technology we have, you know, folks like Intel, TSMC are hard at work improving that.
And that gives us more performance, more transistors, better architectures, uh, better algorithms, uh, uh, you know, whether it's more efficient training or maybe a better fast four EA transform, better code generation, better compilers, better drivers, those contribute as well and larger scale, right? Uh, you know, once a long time ago, it was very common to see people training on, on, you know, a workstation, a small system, maybe a few processors. And now we see the most, uh, cutting edge models are using, uh, hundreds, thousands, tens of thousands of accelerators.
And so all David, Yes. Um, I'd just like to throw out here for our viewers, what they're probably thinking the elephant in the room of the last week or two, which is of course, deep sea. Yeah.
Which is challenging, and an awful lot of these assumptions about how much compute you really need to do to, uh, you know, get really good models. So, so I'm not asking for any sort of response from you now, but, um, I'm sure a lot of people are thinking about this as, as you're going through these venture, do we really know what we thought we knew? Right?
Well, and I mean, I think that, and I think it actually really highlights the point I'm making here, that, you know, to really accurately understand performance and what's going on, you need to look at all of these factors, right? It's, yes, we do need massive compute. And even if you look at like what Deepsea did it, it involved a tremendous amount of compute, but part of what they were showing is being clever and getting it right can really reduce the needs of compute, right?
That was one of the takeaways. And, and part of that is going back to that if you pick the right algorithms, the right tuning, the right architecture, the right optimizations, you can use your compute a lot more efficiently. Um, So when you talk about measuring it, just, you know, maybe just really quickly when we talk about measuring, uh, um, the paper by, uh, Emily Bender at all on, can language models be too big?
Mm-hmm. Um, really recommend also measuring the impact to the environment and to people and the rest of it. So are we gonna talk about that at all?
Um, so peripherally? Yeah. And so, I mean, the thing I would say is, so I actually, I helped, uh, start create lead ML perf power, and that is the industry standard way of measuring the perform the power efficiency, right?
And so, you know, I think the point is, uh, as we look at deploying any sort of new technology, we wanna look at the broader impact it has, right? On the environment and other places. And being efficient with our AI is critically important, right?
And so, you know, we, we are building the tools to help everyone measure what is the energy efficiency of training and inference for ai. And that's critical. And, and, you know, I I, I couldn't agree more that that is important.
I spent many years, you know, working on that. Very cool. Um, so the ML per benchmark suite is designed to help us measure performance, to help drive the industry forward, to help us understand, you know, what is the fastest solution?
Not so we can beat up vendors, although we do a bit of that, right? Or that I should say we don't, but people do. But to help really align buyers and sellers and understand, right, are we getting the right thing for what we want to do, right?
It's a tool, it's about information. So we want, so that means some of the goals of our benchmarks are we want the performance that we measure and the energy efficiency that we measure to be reproducible, right? You know, if you can't reproduce, what if I tell you that I, uh, manage to jump 50 feet in the air, but it was in the room the other day?
You might not believe me. In part, I, I, you know, it seems unlikely I can jump 50 feet in the air, but you know, if, if I tell you, Hey, here's how I did it. I said, ah, you know, big trampoline, and then you try it and you do it.
Ah, makes sense. Seems reasonable. And so we want, we need the replicability to establish trust and to really build useful tools, right?
You wanna be able to know that what the vendor's claiming actually can be done, or what the researcher has achieved is, is not just a fluke. And Ultimately, even if you're, Even if it's not always done, knowing that it could be done, hopefully keeps people from making ridiculous claims about their models. Yeah.
And, and that's sort of, uh, an endemic problem to the AI space because we're talking about such huge systems, so much hardware, so much data, uh, such an incredible importance of software and tuning and optimization that we need them to know that somebody's gonna hold them accountable if they're making ridiculous claims about the performance of their system. Yeah. And That's, that's was one, one of the things that I was looking at when I was looking at their slide is, uh, accountability and auditability of the results are a couple of things that, uh, I would like to see come out of this.
Yeah. That the, uh, the idea that you, you could actually come up with a source material. You can't just say you jumped 50 feet in the room, some type Yeah.
Somewhere. But you can actually say, here's a demonstration of me, me jumping 50 feet. That's right.
And the auditability of, I, I got this result, but can I actually go back and find the source material that gave me this result? Yeah. So, and, and that's great points.
I mean, so we, that is built into our benchmarks. There's peer review. So you know, everyone who does a benchmark, they're all in a room looking at each other's results.
And, uh, you know, It's not just, it's not just peer review, it's, it's the ability to actually have an audit trail of Yeah, you have a result. Where did this result come from And what are the artifacts used to produce it? All of our benchmarks, when you submit logs, the submission, not just the logs, not just the results, but like the, the things that you need to reproduce it all end up in GitHub.
And so, actually getting back to your point, one of the things that's great about that is it means that your recipe, your best known method is now public. And so your customers can see it and they can say, aha, that's how I tune things. Let me go do that.
Or, you know, maybe I can't quite tune things that way because, you know, my systems aren't set up right. Maybe I've got two data centers that are talking across a large distance, and that's not what they had, but I can see how to adapt it, right? And so that transparency, that visibility, that traceability, and in fact, having audits, which we do on some of our benchmarks where it makes sense is, is vital for that.
And so we want our benchmarks to be representative. Can I question this question about the audits? And, and I know more about like performance, the, or task oriented benchmarks, likewe bench and things like that.
Yeah. Um, as opposed to some general performance benchmarks. Yeah.
I find in the task oriented ones, particularly like SW bench and stuff like that, the logs don't really tell how they train to train or teach to train. You know, there's sort of the bias that the models, there's this whole discussion about the benchmarks are floored. I agree with Steven that they're, they're valuable and they're what we have, but a lot of the models are training with the benchmark data in mind.
And when I look at the logs of the task oriented ones, I don't see beyond, I don't see the depth of how they were sort of, I can't see how they taught to train to be able to achieve the benchmark results that they did. I don't know if that question makes sense or not, but, So, so yeah, I'm, I'm, I'm, uh, so the, the question, no, The auditability is suspect even with logs and the GitHub. I, I think that's actually an intrinsic challenge of measuring, right?
As you say, task perf task accuracy is what I would say. Yeah. And that's a really, really tough thing to do, right?
That's actually not what we're doing with ML perf, what we do is we establish, generally we establish a, a a a required level of accuracy. And if you pass that, got it, then you have a valid submission. And, and then you, I was just wondering if those, those characteristics, because you do have what model guard and, and like, are you gonna talk about any of that?
That's, uh, that's a different part of, okay. Alright. What's going on?
So as we build these benchmarks, because of what we mentioned, right? The, the having the reproducible artifacts being a best known method, we want all of our benchmarks to be representative of real use cases. Things that people are commercially, uh, uh, doing that are impactful, right?
So this means, you know, we end up actually having to retire benchmarks pretty often. There's a very rapid cadence. Would you could, could You give some color?
So, so we deprecate or we retire a particular benchmark. Yeah. What is, what are, what are some of the key characteristics that would say this benchmark is really no longer fit for purpose?
Is that, is that a function of the industry? The, the, the landscape or where the technology is moved? Uh, I mean, all of those factors come into play.
I'll, I'll give you an example from some of the earliest days of ML perf. When we started out, we had, uh, two machine translation benchmarks. Same task.
One was transformer based and one was using RNs. And when we started, RNs were really what people used in production, but Transformers were pretty new and they had promising results, and there was a sense that they were much more likely to scale into the future. And so we said, okay, we're gonna have both of them.
But fairly swiftly we ended up retiring the RNN in part because Transformers proved to be so much better that everyone in the industry just switched to that, um, for that particular task. Now, there are other places where RNs are still used, uh, today. They're, you know, oftentimes used on time series data mm-hmm.
Uh, on speech stuff. But, uh, ultimately it's, we want to be representative of what people are doing in production, although we also need to lean in and lead a little bit. Right.
You know, you don't, there's a spectrum of production uses and, you know, you don't want to be too far back. You don't want to be too far forward, right? And so it's dialing that in, talking to customers, talking to heavy users, talking to industry, and, and then looking at, uh, yeah.
Is it still a good benchmark? So, so to simplify it into like a 1, 2, 3. So is the, is the goal to be more associated with the pioneers associated with the settlers or pi?
Uh, or what would be the laggards? Wardly, honestly, I I did go Simon on you there. So, Um, I, I think for ML perf for a lot of what we, so it, it's a little bit different for training and inference, right?
Inference is fundamentally about production. Yeah. And so I think for training, you have to lead a little bit more than inference, but I would say you kind of want to be waited forward.
'cause part of this is we're illuminating for the whole community. Like what's important and what's going to be important, right? I would say, you know, ultimately, actually even some of the retired benchmarks are still useful, but maybe more for some of those lagging cases.
So I think the interesting thing is, I think over time, any given benchmark will, will pass from your, you know, stages. And, you know, we want to pick the right point. And so it's, when we introduce a benchmark, it should be a little bit leading.
Uh, or, you know, for example, when we introduced G PT three as a benchmark, g PT three was already out, but we followed very, very quickly. And again, we have to do everything open, transparently, et cetera. So, um, and, and what this gives us is, you know, we get to, to help improve the state of the art, right?
By making fair and useful measurements. Uh, right. And so the, the, the, the use here is of course, as it goes from, you know, out on the cutting edge performance starts improving, improving, improving, and eventually it might saturate, right?
But, uh, you know, helping to drive that flywheel is, is, is critical for us. Um, so I wanna talk about the overall ML perf family. Uh, so we started with, uh, uh, ML PERF training.
That was the one that, that got us started in, uh, 2018. Uh, we moved into inference, uh, both data center and Edge Mobile for smartphones, tiny for I OT devices, uh, storage, which we'll get to talk about, uh, uh, client for, you know, sort of windows and, and, and, and Mac and desktop laptop machines. And then, uh, automotive.
So there's been, you know, tremendous growth in expansion of breadth, uh, you know, really as ML has begun to touch everything. So David, there are a couple other benchmarks that's not on this list, uh, illuminate and, and, uh, algo perf, yeah. What's the distinction between ML perf benchmarks and non ML perf benchmarks?
Yeah, so I, I, that's a great question. So I think ML perf, the way to look at that is that's really focused on performance and efficiency. Uh, AI illuminate is, uh, focused on measuring a very specific quality that is sort of the risk and responsibility.
Um, you know, essentially is your model sort of doing as you intended, if it, is it, is it a chat bot? And if so, is it wandering off and recommending what car to buy? Or is it really sticking to the task at hand?
Things like that. And algo perf is focused on the algorithmic performance. Uh, so not, so it's focusing on just one aspect of training, not the whole end-to-end thing.
It's looking at are there better algorithms, for example, than SGD. And so, you know, one of the great things that came out of algo PERF was a renewed interest in something in a technique called shampoo. Um, and which is a different, uh, high level optimizer for training.
Uh, but ML PERF is maybe a little bit more hardware centric and, and very much performance and efficiency. And, and over time as we've grown, we've really gotten to, uh, improve the, uh, not just the breadth of the ML perf suite, but the depth, right. We've, uh, have, as you asked about earlier, we added power measurement for inference and Tiny first, then we added power measurement for training and cloud systems.
So, you know, you could actually, uh, if vendors were interested, you could measure the power efficiency of training in the cloud and then compare it to an on-prem system. Uh, How do you manage power? Do you actually do it algorithmically or are you doing a, it, It, uh, so for ml perf inference, uh, in smaller systems, the preferred methodology is essentially a power meter.
So, you know, you're actually measuring the, the power that is consumed for some of the larger at scale systems like, uh, uh, you might find in ML perf training where you've got dozens or, or, or, or hundreds of nodes, you are, uh, measuring power of the compute. And then, you know, there's certain portions that are either, uh, um, estimated or measured. But in general, we tend to favor full system measurement.
Okay. And are you including cooling Full system? So if you've got, if your system is on the wall, you know, if you plug in the system, you've got fans running, you choose to run them on high, that would be part of the Well, but you're, you're not taking into account where you actually dissipate the heat that you're generating to At the full data center scale.
Yes. Uh, so I Mean, you need environmental cooling or you need some sort of cooling method to, to dissipate. Yeah.
So it's not going to measure if you're in a data center. It may not leisure. I think you've answered the question.
Okay. The, the, the outermost cooling loop. Okay.
Yeah. Um, so I wanna talk about what we've managed to accomplish, uh, uh, in and with the broader community. And so this is my, my favorite chart to share.
And so, uh, again, on the X axis, we've got time on the y axis is performance on a log scale. And, uh, there's this solid blue line along the bottom of the screen, and that's Moore's loft, um, more, we all know that that's right. The increase in transistor density per unit area of silicon, maybe a little bit more complicated than that, but, but it's about transistors.
And if you assume that every transistor linearly translates into more performance, that blue line is what you would get. And each one of those colored lines above it is a benchmark from the ML perf training suite. Mm-hmm.
So the dotted lines are for ones that are discontinued. Um, the, the solid lines are ones that are still with us. And so what you see is that if you compare the best performance on, on any ml perf benchmark, right?
This is not a, Hey, let's go find a weak baseline. This is, let's get the best submission in every single round on every single benchmark and compare them over time. What you see is that we're beating Moore's Law, we're ahead of the curve by, you know, sometimes up to 10 x.
And the Biggest challenge, David, with this chart is that some of those benchmarks run on five GPUs, some run on 2000 GPUs. So I mean, over time that's, that's right. The amount of hardware that's been been associated with these runs has increased considerably.
Yes. That's, and that's actually designed into the benchmark, right? Because as a customer, right, if you look at what folks are doing, they are using larger and larger systems because it gives them the ability to solve bigger and bigger problems, right?
And so, and I said this upfront, which is scale is an intrinsic aspect of machine learning. Now you're right that, that, that many of these systems are larger, uh, and that is absolutely a capability that customers want customers need. That does drive innovation.
And so it, you know, from my standpoint, it's absolutely fair, fair to measure that. Hmm. Um, right.
And, and what we see is that, you know, in aggregate, you know, the performance that we've seen improve in some cases is 50 x. And you're right, some of that is through scale, through moving from, you know, maybe one node, uh, uh, uh, up to dozens, uh, you know, some of our largest submissions are, are thousands of accelerators. But it takes truly dedicated hard work, algorithmic innovation, all sorts of optimization to get a workload to scale across multiple systems.
So this is a strong scaling benchmark. We don't increase the size of the data set. So when you are able to use more machines, like that's actually very significant.
There are a lot of benchmarks where they just expand the size of the data set and it will scale to larger and larger systems. That's not true of ML perf training. So when you see, hey, we're able to use more machines that actually is truly valuable to, to customers and represents, you know, pushing forward in.
The other thing are obviously the efficiencies of the software have changed over time. Mm-hmm. 8 or whatever has also impacted how much these sorts of benchmarks can run and try.
Yeah, Yeah. Absolutely. And that's something we wanna measure, right?
Because again, it comes back to is this useful on the real problems that we're looking to solve, right? And, and, uh, one of the things that is exciting is when you see someone come in with a new numerical format and say, Hey, we got this to work. And sometimes what'll happen is, and, and it's been very exciting because sometimes they'll come in and they'll be say, okay, we're using this new numerical format only in some places, and then six months, a year later, they say, actually, we're able to now use it everywhere because we made our op software better.
Right? And so, you know, again, if it's a customer valued, if it's a intrinsically useful thing, right? We want to be able to measure that, right?