MemVerge Memory Machine AI Transparent Checkpointing
Presented by Bernie Wu, VP of Strategic Partnerships, MemVerge. Recorded live in San Jose, California on January 29, 2025 as part of AI Field Day 6. Watch the entire presentation at https://TechFieldDay.com/appearance/memverge-presents-at-ai-field-day-6/ or visit https://TechFieldDay.com/event/aifd6/ or https://memverge.com/memory-machine-ai/ for more information.
Transcript
I'm Bernie Will. I'm the VP of Strategic Partnerships for member. And, uh, I wanna pick up kind of where Steve and Charles, uh, left off.
They were talking about, uh, preemption and checkpointing. And I want to deep dive into that a little bit more for you guys, because that is one of our, uh, core technologies in, in this, in this, uh, in m ai, uh, architecture. And so the, the goal of this, we, we refer to our checkpointing as Transparent checkpointing.
And what we mean by that is the ability to, uh, either pause or preempt and relocate a running, uh, GPU workload, uh, without requiring any changes to the application or even having the application necessarily be aware of it. Uh, there are other, uh, types of platform schedulers that, uh, uh, they don't do this. So this is a pioneering effort here to use this kind of transparent checkpoint technology.
In this kind of use case. Other schedulers will send maybe perhaps a preemption notice, but then the application either has to be checkpoint aware already, or, or it has to, uh, be refactored, uh, or, or just face a cold restart situation. So what our goal is to, is to handle any kind of application, whether it it has checkpointing built in or not, and, and be able to restart it somewhere else hot, restart it somewhere else, or hibernate it and then restart it.
Uh, and this allows us to, uh, remove, uh, what we think is a quite a bit of the friction in doing any kind of resource scheduling and resource management and scheduling, uh, between the platform layer, uh, operators and the, the, and the users. Uh, and besides this basic scheduling stuff, there are other use cases we see that, that can expand the, uh, the benefits to platform engineering. So, for example, we, we mentioned, uh, graceful node, uh, maintenance or shutdowns.
We detect, for example, A GPU overheating. We can preemptively, uh, uh, evacuate that node. Uh, we're using this transparent techno, uh, transparent checkpoint technology.
Uh, another example is this elastic workload bin packing and, uh, optimization. So not only can we, uh, maybe, uh, you know, pack workloads more densely, uh, into, uh, different, uh, nodes, but also we can coordinate with Kubernetes infrastructure for horizontal scaling. You know, scaling up and scaling down, uh, especially long running jobs.
So things like inferencing are gonna be long running services, and at some point they may be using lots of memory. At some point they be using less of memory. Uh, so this kind of check pointing can also be used to more gracefully migrate things into, uh, in, into a scale upscale down, uh, situations.
And then also, uh, idle, uh, resource, uh, and, uh, spot instance reclamation. A lot of people are gonna nine cloud, they're gonna want to use low cost spot instances. Uh, and, and as Steve mentioned earlier, uh, we have the ability with our schedule to borrow, uh, unallocated resources from other departments and things like that.
But in addition to that, if somebody has already allocated it, but it's sitting these resources, but they're sitting idle, we could also use this kind of technology to detect, uh, uh, there's something idle. And then, and then, uh, Bernie. Yes.
It wasn't clear to me that that, that with the checkpointing solution you had, that you could actually change the size of the, let's call it the GPU partition that you're using. Let's say I checkpoint and over here, and I wanna move it over here, but I'm gonna use a smaller portion of the GPU. Yeah.
Or a larger Portion of, yeah, yeah. No, these are some of these capabilities of elastically, uh, increasing and decreasing currently are on the CPU U side of this checkpoint is this checkpoint technology is both a CPU and GPU technology. Uh, over time it'll be more elastic on the GPU side.
So that, that part I'll explain in a minute. Yeah. Uh, so, uh, another point that was brought up earlier in this panel was, uh, that, that there are other types of checkpointing technology.
So I wanted to contrast what we're doing at the, uh, infra level versus this, these, uh, what we call, uh, AI ML framework type of, uh, checkpointing technologies, which are really, uh, designed for the data scientists and the, uh, ML ops engineer. So, as you probably know, PyTorch, I, TensorFlow have built in, uh, checkpointing capabilities. And those are used mainly in the training side to save model, uh, parameters and optimize our states at specific, uh, epics or batch intervals.
And, uh, they can be used to restore the state of a, of a, of a training, uh, workload in case of some sort of, uh, failure of the infrastructure. But one of the primary use cases of these is actually for, uh, rolling back overtrained or, uh, overfitted models, and then also to facilitate hyper parameter tuning. Uh, so, uh, the reason I'm bringing this up is because a lot of people get confused about our kind of checkpointing, which is going on at the platform engineering level, versus this higher level type of checkpointing going on at the more at the data scientists ML ops layer.
Uh, and also these kinds of checkpoints are, this kind of checkpoint is also used for data for model interchange because the, uh, uh, the, there's enough metadata such that you can take a, a checkpoint on this kind of infrastructure and then move it over to a different infrastructure or even a different C-P-U-G-P-U combination and, and, and fire it up there. So that, that kind of, again, that kind of framework type of checkpointing is used for interchange. Uh, and, uh, one, one of the, uh, peculiarities of this approach, however, though, is that the, these checkpoints are framework specific, uh, not platform specific, not, not, not optimal necessarily for platform engineers.
And they're also user specific because they have to store their, uh, checkpoints in specific files and directories. And so in a larger scale environment with multiple tenants, you can imagine this can get pretty chaotic. Uh, and, uh, in addition to these frameworks, there are some pipeline, uh, workflow pipeline, uh, uh, uh, uh, tools out there that do have what I call checkpointing between the, the pipe segments, uh, are out there.
Uh, and if you look at checkpointing in general with this kind of approach, because it generally has to be done on a periodic basic basis for every epic or, or, or bashes, uh, the overhead's pretty significant. It can be up to 25% of the total runtime is spent doing checkpointing. Bernie, is, is the purpose of this type of checkpointing, is it What you're talking about here?
Is that optimization, or is that really Disaster recovery? Uh, most of this is for, for the model training itself. So you don't wanna overs, so you have these different training epics, right?
You don't wanna overshoot overfit the model, so you do wanna roll it back. You also wanna tweak all the hyper parameters. So the, the primary goal has been for, uh, for the model itself optimization, right?
But, but also as these bonds get bigger and bigger, uh, people start worrying about fault tolerance. And this is one way to, to a certain degree, to, uh, address fault tolerance as well. Mm-hmm.
But it's really, again, it's, uh, uh, these, these types of things are used by, at the ML ops, uh, uh, data scientist level, but somebody running a much larger, uh, uh, cloud native infrastructure underneath that, uh, he, he can't really address problems down in his level. Uh, and that's where, uh, going back to our kind of checkpointing is where we've been spending a lot of time working on that innovation. That's what we call transparent, uh, checkpointing.
That's really the, the target persona. There would be the site reliability engineers and the platform engineers. And this is a, a technology that does memory snapshotting.
So we're actually capturing the state of the machine, all the, all the caches and sockets, uh, and, and, uh, uh, as well as the, uh, uh, uh, uh, ephemeral files and things like that, and saving all that to a, uh, file system. Uh, and initially we, we developed this, uh, uh, for CPUs, uh, CPU based workloads, uh, and, and, and, uh, uh, compute instances. And, uh, we, we've been working on this for several years, but we've had actually products in production based on this for the last two years.
We have another product line called Memory Machine Cloud, where we, uh, we have been using this to, uh, checkpoint CPU instances, EC2 instances out in the cloud. And, uh, we've demonstrated this working on cloud bare metal container and VM environments. And so, uh, we've been hoping this been working on optimization of its, uh, uh, overhead.
So there's techniques that we've developed to, uh, reduce the, uh, interruption time. Uh, so one of the problems is that when you're doing this checkpoint, you have to pause. We, we, we also have to pause the application and, and dump all the memory and, and the files and everything else.
And that takes time. So we've been, we in some cases have been able to reduce it by over a hundred x, the, the interruption period we call before we can resume production. Likewise, we were doing work to reduce the amount of storage consumed by these checkpoints.
So the techniques we're using, there are things like, uh, and some of this is borrowed from the storage industry. There's like a notion of an incremental memory snapshot. There's, uh, data compression technologies.
There's a asynchronous checkpointing technology where we go to memory first and then D stage, uh, asynchronously to, to, uh, to a file system. Uh, so all these different techniques in combination can be used to drive down the overhead of checkpointing to make it viable for, uh, production. And also, we've been working on integrating this checkpointing engine with various types of, uh, batch schedulers, and now more recently on Kubernetes.
So, uh, uh, and then the, the, uh, uh, the pivot toward GPUs began, uh, with, uh, collaboration that started in the first half of 2023, where we started working with Nvidia on doing this process level checkpointing, uh, extending it to GPUs. So a lot of that work, you know, quite honestly, had to be done by the, the Nvidia side. So they have been making modifications of their Cuda driver to support, uh, checkpointing.
4 driver. And the same time, we, uh, we, uh, presented a poster at the, uh, GTC event last year where we, uh, uh, talk about, uh, achieving, uh, KAS Kubernetes and public cloud operational efficiency, using for the first time a checkpoint, a transparent checkpoint and restore that works at the GPU level. Uh, and the way this is done is right now, currently a, uh, a two stage process.
Right now, the, the, the, the, uh, X axis is the, is time. Uh, and so there's two things that are checkpoint and restore. And the y axis is there's GPUs, G-P-U-C-P-U in storage.
I know that's a little bit hard to read. Yeah. But, but basically it's a, it's a two stage process.
You can see, first of all, we check checkpoint the GPU and dumped that memory into the CPU. And then we checkpoint the CPU, including the GPU memory that was just dumped in there. And then combined with the files and femoral files and o other, uh, objects, uh, uh, storage objects, dump all that into a checkpoint image and store it there.
And then we, uh, rehydrate on the other side, just doing the reverse process. And where does it gets, so it gets stored locally? Uh, it could be, yeah, whatever you call locally.
Yeah. We, we would, Kubernetes, we would store on some sort of, uh, of, uh, persistent volume claim. Yeah.
A storage. Yeah. We put it on Storage.
It's super storage. Yeah. So this is not a, this is not, in contrast, this is not some sort of live migration.
This is a checkpoint store and recover. So we can use this to hibernate, uh, low priority workflows for an indefinite period of time. Somebody goes home from, for work in the evening, we wanna just hibernate his stuff, and then the next morning bring it back and resuscitate It's three.
And it reminded me of SRDF, y'all see? Yep. And then the way we implement this now in Kubernetes, this is a little one more level of detail, is that we've built an operator, uh, and, and the, and this operator allows us to take advantage of all the automation benefits of being on the Kubernetes platform.
And there's two basically ways that it can be deployed. One is that as, as a, uh, custom resource, uh, device, uh, CRD, uh, and, uh, and so we can deploy as a demon set across a, uh, an, an a namespace, or we can also deploy as a, uh, uh, an annotation to a specific application. So we can annotate or, uh, be labeled to specific applications.
And then we can be deployed as a, uh, what's known as a Kubernetes sidecar. So there's two ways we can deploy. And then again, as to your question, we store all of our checkpoints into this persistent volume, uh, shown on the, on the right.
And the, again, the, the beauty of this is that it really helps reduce what we call the friction of moving hibernating or moving or hibernating, uh, stateful, uh, long running, uh, uh, workloads, which are, a lot of AI workloads are like that. Or also long init initiation time workloads. So a lot of, a lot of workloads take a long time just to load up the model ways or just to get initiated.
So this is also a useful, Right, just using Kubernetes. If Kubernetes recognizes, for example, an A MD MI cart, would you be able to do the same thing with, uh, A MD? Yeah, we, we plan to support a MD in the As along.
Yeah. Yeah, absolutely. Yeah.
So we can use a similar mechanism for, for a MD. Uh, so what's ahead, uh, right now? Uh, there are gonna be further performance enhancements and, and optimizations.
As you saw. The initial, uh, work that we've done so far with Nvidia is a two stage process. We can parallelize that so that the, that the, the memory dump and the from the GPU and CPU are occurring in parallel.
Uh, there's also, there's also been a lot of talk about fractional and, uh, and, and also, uh, uh, multi GPU, uh, node supports some of this work as is gonna be done in collaboration with Nvidia. 'cause it does require further, uh, evolution of their kuda driver to support that. Uh, likewise with cross generational GP migration, right now we're limited to like, to like, as, as we mentioned earlier, a 100, a 100, et cetera.
Uh, so some of tho some of those constraints will be relaxed over time, uh, in our, in our collaboration with nvidia. In addition, we're also working on what we call transparent, uh, uh, clustering, uh, uh, clustered checkpointing. So right now we can do individual nodes, but as you know, a lot of things, especially on the training side, are clusters and even inferencing is become a distributed, uh, architecture.
So, uh, the notion of gang scheduling to be able to transparently checkpoint that is, is also on the way. Uh, so the notion of doing what we call job set, uh, actually that's what Kubernetes would call it, job set, uh, uh, coordination across multiple pods, and then also supporting these backend networks. There's, there's RDMA networks, there's infinite band networks, uh, NV link networks.
And, uh, some of that will be done in collaboration also with, uh, Nvidia. Uh, and, uh, this checkpoint technology has other, uh, use cases besides just this basic, uh, scheduling and platform engineering, uh, management, uh, hybrid cloud scheduling and migration is another use case. Uh, heterogeneous pipelines.
So a lot of these pipelines, uh, are gonna be pieced together from different, uh, components. Software components can imagine a multi-agent, LLM rag pipeline could have several stages, and there could be situations where we want to, uh, checkpoint and migrate or hibernate the entire pipeline, uh, and then restart it somewhere else, or, or just let a higher priority work, uh, go by it. Uh, h HPC is also a, a good application for this type of, of, of, of clustered checkpointing.
Uh, and as mentioned earlier, model, model loading and model restart. So also kind of putting, uh, high, kind of accelerating further turbocharging, further, uh, serverless architectures and things like that can be used with, can, uh, be, uh, a beneficiary of this kind of technology. And also, we will be working with other schedulers and more, uh, uh, directly integrating, uh, our stuff into things like storm.
So the next checkpoint, uh, for this checkpoint technology, it'll be CubeCon in April. Uh, so this kind of gives you a clue where we're heading, as I, I said, uh, we are working toward, uh, doing these cluster wide, uh, checkpoints. And, uh, so it's been very exciting and we, we'll be presenting it jointly there with, uh, Microsoft.
And, uh, again, to remind you, we have this pioneer program open up for, uh, for, uh, anybody that wants to be an early adopter of our AI infra. Thanks. Not Steven.
I don't know if this is appropriate question my first time. That's a terrible question. I won't.
So do you have any reference accounts yet, or any reference accounts you can talk About? Uh, no. We're just starting this early adapter thing as, as we speak.
Yeah. Okay. Yeah.