Driving Storage Efficiency and the Impacts of AI in 2026 with Solidigm
A discussion between Solidigm and Vast on the efforts in the last year, from the all-flash TCO Colloborate to the way our technologies have synced to solve AI Market demands, discusses some recent Context-related impacts to the year of Inference for 2026. With the evolution of DPU-enabled inference platforms, the value and capabilities of Solidigm storage and Vast Data solutions drive even greater customer success. Solidigm’s Scott Shadley initiated the presentation by highlighting the immense power and storage demands of future AI infrastructure, using the “1.21 gigawatts” analogy. He projected that one gigawatt of power could support 550,000 NVIDIA Grace Blackwell GB300 GPUs and 25 exabytes of storage in 2025. This scale requires extremely efficient, high-capacity solid-state drives (SSDs) to stay within power envelopes, making Solidigm’s 122 terabyte drives a key enabler.
In 2026, the presentation introduced NVIDIA’s Vera Rubin platform with Bluefield 4 DPUs, which fundamentally alters AI storage architecture. This new design introduces an “inference context memory storage platform” (ICMSP) layer. This layer, positioned between direct-attached storage and object/data lake storage, is critical for rapid access to KV cache data in AI inference workloads. The new hierarchy distributes the 25 exabytes across high-capacity network-attached storage, the new 6.4 exabytes of context memory storage, and 6.1 exabytes of direct-attach storage. This evolution, while reducing the number of supportable GPUs within the 1-gigawatt limit, requires faster NVMe storage to improve performance and is projected to drive a 5x or greater compound annual growth rate (CAGR) in high-capacity storage demand.
Phil Manez from Vast Data then detailed their role in driving storage efficiency for AI. Vast’s disaggregated shared-everything (DASE) architecture separates compute from storage, utilizing Solidigm SSDs for dense capacity. This design enables global data reduction through a combination of compression, deduplication, and similarity-based reduction, achieving significantly higher data efficiency (often 3-4x more effective capacity) compared to traditional shared-nothing architectures, which is crucial amidst SSD supply constraints. Critically, Vast can deploy its C-node (storage logic) directly on the powerful Bluefield 4 DPUs, creating a highly optimized ICMSP. This approach accelerates time to first token, boosts GPU efficiency by offloading context computation, and dramatically reduces power consumption by eliminating intermediate compute layers, enabling AI inference workloads to operate at unprecedented speed and scale with shared, globally accessible context.
Presented by Scott Shadley, Director of Leadership Narrative & Evangelist, Solidigm, and Phil Manez, Go to Market Execution Lead, VAST Data. Recorded live at AI Infrastructure Field Day in Santa Clara on January 30th, 2026. Watch the entire presentation at https://techfieldday.com/appearance/solidigm-presents-at-ai-infrastructure-field-day/ or visit https://techfieldday.com/event/aiifd4/ or https://www.solidigm.com/ for more information.
Transcript
So I'm Scott Shaley, director of Leadership Narrative with soy. 5 architectures. And then I'm gonna hand it off to Phil and he's gonna fi wrap up this section and go on to the, the, the last section of it.
So you only have me for a few more minutes. I'll be back soon. So what does one point 21 gigawatts solve?
Come on, somebody's gonna smile about that. Oh my word. I'm not that old.
Alrighty. That's shocking. It can send Marty back to the future.
21 gigawatts. It's also what it takes to power San Francisco for a day. It can also deliver power to 550,000 Grace Blackwell, GPUs, GB 300 platforms, and it enables 25 exabytes of storage in a one one gigawatt environment.
Now, how do I know this? What did we do to be able to tell you that one point 21 1 gigawatt can do all this. Again, looking at the ecosystem, looking at our friends, doing the research, these were all announced in 2025.
These are all the platforms that are gonna go live sometime after the announcement in 2025. Stargate Meta Core Weave, XAI, all these guys. And we did some math and that's how we got to the one gigawatt for 550,000 GB 300 platforms.
And if you look at the math required for how many GPUs you have and how much storage you need to go along with those, you get your 25 exabytes of storage. And so the next question, of course, is really how do you get there? Well, we did that math for you.
Um, again, 550,000 GPUs direct attached. These are the E one s performance drives going right next to that, that GPU sitting in the server. That's the nice, uh, rack based design, currently has eight drives in it because they're air cooled or potentially, uh, direct to chip, liquid cold plate cooled.
And that nets out to eight and a half exabytes of storage next to the GPU. 5 exabytes of supportable storage using 1 22 terabyte drives to be able to get up to 550,000 GPUs. So this math is, is like all good TCO models.
It's a tit for tat. So whatever you put in, you get out. So we focused on, if I use just my solid state drives at the highest capacities for both knobs, it enables us to get to that many GPUs.
You have any other product that consumes different amount of power, or you put sixties in instead of one 20 twos, you're doubling the power footprint. You reduce the number of GPUs available in that gigawatt. So the gigawatt was our bar here.
I'll be 100% honest. There's math that shows if you ignore the gigawatt and just look at the GPUs, the amount of exabytes of storage will be just mind blowing that they're expecting to use. And this is Grace Blackwell.
This is 2025 data. We're now one whole month, literally last day of the month into, um, 2026. And the whole ecosystem has changed because when I showed you guys this graph last time at field day two, it had a quote from our good friend Michael Dell.
This is our friend Jensen at CES. This is the market that never existed in the market that will likely be the largest storage market in the world. Gotta love the fact that we finally have Jensen talking about storage and not just memory.
I love it. Now the next trick is that GTC to have mentioned our name, we'll see if we can get there, right? You signed our drive, of course, you know that whole thing.
5 IC MSP layer in the Vera Rubbin platform that's tied to the Bluefield four implementations. Now, the next slide, I'm gonna tell you how that works from a SSD hardware point of view. And then a little bit in just a few minutes, Phil gets the lovely chance to show you how that actually looks from a system level implementation point of view alongside what we're talking about.
So what I'm doing here is we have the KV cache exceeds the HBM spills over to dram. We still have a limited DRAM footprint only. So many dims, only so much capacity falls onto the local storage, those nice directly attached products, and then it falls over again.
And every hop is a connection and a distance and a time. And so by bringing it closer and closer and closer, that's what this whole architecture is about, is access to data faster mm-hmm. In a more confined environment.
And so if I take what I had before on the one gigawatt example, and I throw in the Vera Ruben platform or the Ruben platform as it's being called, they haven't officially given it the GB nomenclature. We don't have three banks of stor of storage products. Now, again, I'm constrained to one gigawatt.
So I'm still at 25 exabytes, but this is now how it splits out. 4 exabytes of this new context memory storage, which can still be the high capacity drive. It's just closer to attached to that blue field forward architecture.
1 exabytes of direct attached storage. 'cause they're changing the amount of local to move it just a little bit further out past the blue field. So the number of drives in that initial server is actually coming down as the capacities are going up.
And so the interesting thing here is we're with the one gigawatt. We still have 25 exabytes of storage. We split it out three different ways now instead of two.
But note the number of GPUs that are supportable, we're down to 440 or down to 400,000. We lost 150,000 GPUs. But the performance of the system doesn't change at one gigawatt.
And the reason for that is because you're using NVME storage for the direct detach and for the context structure, you have to have the fast storage products in those two layers to overcome the gigawatt problem in this environment with, uh, this new added layer. Because fewer faster GPUs need more access to fast data. Therefore, you get ICMS with, uh, solid state drive.
You can't put traditional rotating hardware in that layer. You just can't. So if I, when we come back, uh, at our next field day, AI field day, uh, we're planning to, we're gonna have even more details on this and we're gonna blow off the one gigawatt and just show you the capabilities of what the storage looks like.
And we're talking a five x or larger CAGR year on year from 26 to 30 on just the demand for high capacity storage over what we were already talking about. It went from where it was at about a 20% cagr. It's now 30 40% CAGR because of this introduction of this.
And that's for the NVIDIA only based systems. So the storage platform is now the shining star on making the success of the next layer of AI as we get into the inference context scenarios. So we're gonna do a little bit of context switching here.
Um, we wanna talk about efficiencies too. And so I just gave you the hardware centric ICMS one gigawatt constrained environment view efficiencies. When we partnered with Vast, we came out with this amazing TCL model that talked about replacing your SEF hard drive infrastructure with R one 20 twos and Vast in an efficiency play, talking about how to make your systems more effective.
And when we were at sc, we put up a bunch of slides. This was an SC carryover for you guys. You wanna talk about it from here.
Um, the Computer History Museum from the museum to the 1 0 1 is the SSD implementation on equivalent of this graph. The hard drive implementation is going from the Computer History museum all the way up to Oracle headquarters, where, where they were up in the NICE four, I know what an Oracle headquarters, but that's how, that's how far our distance is that you do when you put a drive end to end to end and how much reduction in overall ecosystem environment you can drive. So what we're gonna do now is I'm gonna hand it over to Phil.
He's gonna help explain a little bit about this and then jump back into the context memory and give you some more fun, uh, topics about, uh, the wonderful VAs platform. So I-C-M-S-P is inference context, Infe Inference, context, memory Storage Platform. That's what they called it.
And It's behind the bluefin. So it's effectively a, a storage solution out there that's doing something for the context management. I'll dig into it.
It Looks like great lead into the next little section. And it's different than the rest of the NAS object data lake that's behind it. It It re-architect For you.
Okay. Yeah. Thanks.
Yeah, yeah. We'll dig into it. Okay.
Thanks for having me guys. So philes, I'm the go-to-market execution lead at Vast. I've been there, uh, I think it's three or four days while I hit my six years at Vast.
So I've kind of got to see the company, uh, grow. We'll do a quick introduction, but I really needed to play off Scott's, uh, metaphor here, which is one, really happy to be here with solid Diamond and our partners. But vast, we build software, right?
We can't run well obviously without any hardware. So peanut butter great, you know, but it doesn't work so well. It's not very portable or, uh, reasonable to eat if I'm gonna spread it on my hand.
So really, the jelly and the bread to our peanut butter is solid. I'm, I couldn't go without that piece. Um, for those who, you know, not as familiar with Vast, we actually launched really the company to the public here at Storage Field Day back in 2019.
And when I was interviewing, that's really how I learned about the company and whether I wanted to work here, right? And I saw some obviously compelling things, um, since then, right? We've really become a significant portion of the storage market as we look to this year.
We're gonna drive a very significant portion of all enterprise SSD utilization, uh, with storage, expecting dozens of raw exabytes. That's before our data reduction, which we're gonna talk about that efficiency and also doesn't count any of the data going into the cloud. We've made some really big announcements on cloud partnerships this year and extending the platform, which was primarily on-prem into the hyperscaler space.
Sales have really gone well for us. We're roughly tripling year over year. Our quarter's gonna finish tomorrow.
So pay attention as we start to, to announce some of those new things. We have our, uh, customer event first ever user conference for Vast at the end of next month. We'll talk about that.
And I would say just like Tech Field Day evolved from Storage Field Day to AI Field Day Vast has really evolved, uh, from being a storage company to building many more things on the platform, which we'll talk about, really allowing you to capture data, contextualize it, and then act on it with ai. So I would say the founding principle of VAST is that really a few things. We were very bullish in 2016.
AI was gonna change the world. We were very confident that AI was gonna change the way we computed on data, right? I think both of those things proved out to be correct.
And then the third is that the architectures that got us to where we are, were not the architectures that were gonna take us forward, right? And this is really the main culprit, the shared nothing architecture really invented by Google in 2003 in a white paper, basically defined the internet kind of application cloud era where you've got these node based architectures, right? I've got a node, got some CPU in memory, I've got some kind of storage in there.
Originally it was disc. Now we swapped it out for flash in a lot of circumstances. But the only way to get to the data on that node is through that nodes controller, right?
And I personally storage guy, like I look at this is every scale out na, every scale out object platform. But ultimately it's also the architecture for every data lake, for every distributed data warehouse, right? It's all over the place.
Eventing infrastructure, it is everywhere. And it really does create a lot of scale problems in the AI world. When we look at it really from a storage view, there's some challenges around flexibility, right?
I've gotta create nodes or pools of homogenous node types. Uh, not designed with flash in mind, right? We'll talk about some of the challenges around things like data reduction.
And then finally, one of the big things that shows up everywhere is just this east-west traffic. There's so much communication between these nodes that even though I can scale my resources linearly, I'm not scaling performance linearly. And a lot of times these architectures work well small, and these problems show up more and more and more as the clusters grow.
So we're looking at it now from the, the really the TCO and efficiency perspective, right? We know we're in a supply crunch, right? Customers have been trying to move steadily from, uh, spinning disc space architectures to SSDs.
When we look at the AI deployments that we see in practice, you don't see any spinning disc, right? You got power challenges. I I was wondering why you actually had the round things with the floating heads on them for showing describes here Because our marketing team likes that image, I guess.
But these are all right. The, the world of, uh, architecture based on spinning, right? It's got the arm.
It's more like a record player. Oh, uh, older records. You old school, okay, well, I'll, I'll, I'll give the marketing team the feedback.
What's A record player? I love it. Okay, so in the AI world, right?
I think as you look at what's in practice, solid states required, right? I think now we're seeing a rise of companies coming up around saying, Hey, tierings cool again because we have an SSD supply crunch. But if it was not a good idea before a supply crunch, I don't see how it's a good idea after a supply crunch.
So what we need to do is help customers be a lot more efficient with the way they use SSDs, right? One of the big problems with this shared nothing architecture is how data reduction works, right? And when you look at a lot of these architectures, a lot of them have given up on things like deduplication.
You have compression only, right? And a lot of the data in the unstructured world, it's already compressed. So that kind of takes away a lot of opportunity to drive efficiency, right?
2 to one is kind of what you're gonna get if you're using compression only. So what is the challenge with ddu? It's around having a global view of the data, right?
In this world, I basically have to chunk up my DDU index because the other option would be to put the entire index on one node. Everyone would just hammer it and that would not work very well, right? So we said, okay, we're gonna shard this up.
Essentially create a distributed database. And in that world, every node has a piece of the index, right? So I get a limited view there.
As I start to scale this, all these nodes are talking to each other, looking at who's got the data that I already might have as I grow it adds to the EastWest traffic, adds to the performance limitations. And then ultimately the kind of bandaid there is to create limited DDU domains, right? So I'm basically only de-duping within a pool or within a few different nodes within whatever architecture that you're building around.
But it's always very local in this world. So vast, again, looking at the architectures, brought a new architecture to market that we call date very quickly. We call it disaggregated, shared everything.
Because essentially we kind of broke the idea of a node apart. And we have two independent scaling layers. We have our logic compute layer, we call those C nodes.
It's essentially container running on an X 86 server. And then we've got where all of the state of the system lives down in these enclosures filled with very dense, solid, IM 122 terabyte drives or whatever the right, uh, drive is for the customer. Now, some different things.
And unlike the shared nothing world, in the day's world, every one of those containers has direct access and actually sees every one of the devices in the system as a local device connected over NVME over fabric, right? Architecture impossible without NVME over fabric, which now makes it allow that I can have remote drives, feel local from both how they're mounted and performance. We also have a layer of storage class memory in the system where all of the systems metadata lives.
So that means I can create a shared global index that every single one of these containers sees. And that means I can do global data reduction at an exabyte scale without any of those different challenges, right? So fundamentally unique architecture that allows us to look at the data in a very different way from a data reduction perspective.
Any questions? High level to architecture, how it works? Okay, that'll be a theme that we hit on, right?
So step one, can we give an architectural, I would say, advantage to how we look at global data reduction, step one. But again, the problem is we're talking about unstructured data here, right? Not as friendly of deduplication as things like VDI and virtual machines, right?
Not as friendly of compression, maybe as a database that hasn't been compressed already. So we had to look at some different things, right? And really move beyond duping compression alone.
So very high level, you look at compression, right? I'm looking for commonality, repeating data at a very granular level, right? That's gonna be a small chunk, I don't know, eight to 60 4K usually could be anywhere in between, uh, deduplication.
Right? Now I can have a global view, assuming my architecture allows it, but I'm looking for more course matches, right? Two chunks of data exactly the same.
I find that a lot. VDI, virtual machines, I'm copying databases, whatever. Uh, but again, I don't always find identical matches in unstructured data.
If I chunk up and try to do DDU on a big pool of unstructured data, what you actually end up finding is a lot of chunks of data that are mostly the same, not exactly the same. DDU misses that every single time. 'cause that would be a hash collision that's corrupting your data.
It's terrible. So what we do is introduce a new type of data reduction, again, enabled because we have this giant metadata structure living in storage, class memory, the architecture that we will identify if two chunks of data are mostly the same, compress them together and essentially store the differences, kind of like a snapshot, right? And ultimately, we don't just use similarity, we use all three of these, right?
So we're looking for the best opportunity compression. We actually use a couple different types of compression. We will look at the data, take a sample, what's the best type of compression, and use that deduplication.
We have, again, global deduplication. We have something we call adaptive chunking. Chunking, which means we'll actually change the dedup window to find the best opportunity for deduplication.
And then similarity is kind of that icing on the top where we're gonna find that next level of similarity and drive out even more savings. Right? And ultimately, you're in a world where we could easily get two or three times more data reduction than the next biggest competitor because of what's happening here.
Do you, uh, I wanna know if you're a believer or not. Uh, yes, absolutely. Uh, I, I'm chuckling because you're, you're giving the exact description of what I would've been describing with solid fires architecture 10 years ago.
Got it. Okay. I knew about it.
Right. But I would say solid fire in the shared nothing world, right? A little bit.
Absolutely. I, but I mean the, the, the things you're describing are just like, yeah, this is exactly what we were doing 10 years ago. No, it makes so no, that, that's, I'm sorry.
That's why I was chuckling. No, no, it makes sense. And by the way, I think it's interesting that, you know, in the block world, the hard drive died like immediately, right?
You know, I was part of the extreme IO team at EMC. We had pure, we had solid fire. Everyone.
The, the hard drive in the, like the block world, virtual machines, databases. VDI died immediately. That was like 12 years ago.
And there's still so much of the world's unstructured data on spinning disc because they haven't been able to figure this calculus out, right? So it's actually a great point. Um, some actual data, right?
So if you were gonna say, I don't believe you, I was like, look, I have data, um, average data reduction by the way this is pulled this month. Because as we've looked at the supply crane crunch, we're like, let's start digging into like where we've come and what the results are. 4 to one, right?
Again, these are not VDIs, these are, this is unstructured data. Some of our customers have hundreds of petabytes of highly compressed video. Some of it's encrypted, um, some massive estates.
The weighted average. 87 to one. Exactly.
9, looks prettier on the slide. Um, and then we have 27% of our customers get better than three to one. We have some customers getting eight to one.
We have some customers getting like 30 to one depending on the data type. So where typically, again, in the world of unstructured data, you'd say, if I get anything at all 10%, I'd be happy. We're talking about getting you three times more data for your flash.
And that gets combined with something that I'm not gonna nerd out on today because of this time. But our erasure coating is also incredibly efficient. So our erasure coating at scale under 3% overhead, it's actually 146 plus four stripe that we use enabled by our architecture.
So, um, when you look at that compared to, you know, traditional kind of shared nothing where you're gonna have maybe 27, 20%, we have a lot of customers moving to Vast that are still using like das Data Lake technology and they've got their data triplicated, you'd be shocked about how much of the world's capacity is still triplicated. And it's because it's in these monster data lakes, um, where again, they're getting, you know, for every 10 petabytes they can store three petabytes of data. And those systems don't have any data reduction, right?
That's all over some of these large data analytics environments. So you combine these things, a lot of times our customers, even if they don't get good data reduction, they're getting four times more effective capacity per, you know, petabyte that they buy. And even if they're buying something that is more kind of enterprise, then maybe it's more like double the capacity that you can store for every raw petabyte that you're gonna buy.
So we actually just launched this, uh, something called Vast Amplify. So in the SSD Crunch vast over the years has really gotten a lot more flexible. Again, uh, when we started, we had to run out of every specific hardware build.
Now we're working with pretty much every major OEM vendor running on more, uh, traditional servers. We're in the cloud. So we actually have a program where we're going to customers and taking their SSDs that they already have in their systems and repurposing them into vast systems to amplify the capacity.
We actually had a cus a couple customers come to us and said, Hey, we've got SSDs. Your technology is way better than what we're using. Can we reformat these and use them?
And we have in, in some very large scale environments. I'm talking at this point, we've repurposed hundreds of petabytes of data, thousands and thousands of drives. So, Just a question.
This is all really great statistics and y'all are doing really awesome, but, um, we're talking about ai. So I would love if you could tie this back to ai. Can you, does it matter if I have a data lake that's not deduped?
Maybe I want that and I just want the, I just want the data tagged in a different way so I can find it for different reasons. But like what, how does this tie back to ai? Yeah, so I would say how it ties back to AI is right now what we've seen in practice, any large scale training environment, any large scale inference environment that's actually in production at scale is a hundred percent based on SSD.
Right? Okay. That's, that is standout.
We have no world where customers are gonna struggle to get as much SSD as they need, right? So what I need to be able to do right now is make more use of my solid state devices because AI is driving tremendous demand, right? And we'll get into more how it fits in the architecture, but the point is, looking at bringing spinning disc into this world we really think is a terrible idea.
If having a tier miss is going to kill my GP utilization, destroy jobs, destroy performance, then I can't use that as a lever. I need to figure out how to make most use of my flash in this AI world right? As I wanna deploy agents and inference over a much broader set of data that data's hitting on spinning disc, it's not gonna work.
Well, you're Not gonna get an argument about that here. So, yeah, right. It so go ahead.
I was just gonna say, but when it comes to draining training data specifically is, uh, as de duplicable, if that's a word, as traditional data sets ha have been in in your experience so Far? Yeah, so I would say in training data, um, two to three to one, okay. Is common, right?
If you look at some of the bigger neo clouds that are our customers, two to one's pretty typical on training data sets and, and more. And You're doing this all in line ddu, right? So it's, I would say it's kind of the best of both worlds between inline in the old world, inline men in memory with vast, our inline memory is storage class memory, right?
So what happens is the data lands in storage class memory, it's acknowledged up to a host and then we date, we data reduce it when it migrates down to QLC. Mm-hmm. Okay.
So yeah, kind of at a band, I'll show you what that looks like Actually. Yeah, but I mean that's, that's not terribly unusual way to do it where you, you actually need to do the hashing at some later point to, to be able, you duplicate it, but you end up using much less storage later on it. It's, right.
Yep. What I would say the difference is we don't land it on the capacity tier, right? So we don't land it on QLC and go back and mess with it again.
It's not good for wear, it's not good for performance. What we do is we leave it in storage class memory where it's very fast access, it gives you a lot of opportunity to move. And then we don't need to plan to have non de-duped and non-used data on the capacity tier when you're, and I promise by the way, the most of the rest of the presentation will be specifically on ai, but we wanted to bring in the TCO of making SSDs affordable and and hacking the supply chain crisis.
Was that okay? Thank you for saying that. 'cause that was not coming through.
Okay. Sorry about that. Appreciate that.
When you're, uh, repurposing SSDs, are you having to, to migrate the data off the SSDs and migrate back on from a vast perspective? Or are you assimilating We do need to simulating The data directly. I mean, We do need to move it.
Yeah. So if we take a file system, we can't convert it to vast and data in place. So a lot of our customers were either working with swing space or we're creating clusters and failing nodes out and growing into it.
You're seeing a lot of usage of the, uh, the new capability to, uh, reuse SSDs. Yes. So again, we customers brought us the idea originally to say, Hey, we've got SSDs, we wanna repurpose it.
So that was how it got rolling. And since then, yeah, customers are all over us to say, we know we're looking at the year, we've got more demand, we're looking at rolling more ai. We've lived in a solid state world and we, we know that we can't have capacity that's 30% utilized, right?
Or giving us one third of what we're buying to store. Okay. More AI stuff, right?
So now I promise the rest specifically on ai, right? But again, we think flash is that kind of first step is to enabling your data on fast access. So we're just gonna talk about the context challenge and KV C and why, right?
And again, you guys probably have been paying attention to what's going on with Nvidia, but for people maybe, you know, more infrastructure folks, essentially the thing is here, right? I ask a question to whatever large language model, the first thing that it does is trying to figure out what do I actually care about, right? There's different words in a statement.
The what, the, the, uh, what, what is this guy actually asking about versus some of these words that don't make sense? So I calculate that, turn it into key, uh, key value stores. And that's essentially the context of the conversation.
Something else that adds context is maybe a document or a video, right? Someone says, Hey, I wanna just summarize, you know, solid, I'm in vast tech field day, I'm gonna upload the video into my favorite large language model. But if I ask a question, again, it used to be I had to calculate all of that context again, right?
So a question, maybe not the end of the world, but if it's a document or a video, then I'm calculating that a lot, right? Think about some big enterprise organization dumps a new document out to the world and all of their employees are asking questions about it. I'm recalculating that same context on that document over and over and over again, right?
And then obviously I need to make sure the decode phase is the answer part. That is where I'm creating an answer that makes sure it's related to the question that you ask. So there's some big problems with this context piece, which is one, if I'm recalculating over and over and over again, I'm burning GPU cycles on something that's not adding a ton of value, right?
And honestly, GPUs are too expensive for that, right? We had, I think NVIDIA's customers are like, we can't just keep dumping all this CapEx and scaling forever. You gotta help us use these things more efficiently.
That's step one. Problem two, user experience. If I'm a user and every time I ask a question about a document, it's going to do a bunch of work that's so annoying, I wanna engage in a conversation with you, not have you forget what we're talking about every time I ask a new question, right?
I think, you know, some of these large language models, they know everything about you because they have all of that data. So that's KV cash. I wanna store that context so I can condu continue to reuse it, right?
And Nvidia has this hierarchy, which Scott talked about. So step one, stored in high bandwidth memory, right? Obviously there's some challenges there.
It's really expensive and hard to come by right now. Uh, the other piece is it's local. So if I'm engaging in a conversation just on, you know, this one session, that's fine.
But if my friend is trying to have the same conversation, you know, do I have access to that memory? Then I can move it down to dram, right? Then I can move it to local SSD.
Again, everything's local. And then the next step is shared file and object, right? When you're gonna have a massive drop off in performance there.
EastWest traffic, all those different things we talked about. We shared nothing. So we were like, we need something right here, right?
5. Something that has a performance closer to local but is more global in terms of its access, right? And that is essentially what I-C-M-S-P is, right?
How do I create that local field? Now again, I'm not gonna spend a ton of time on this, but one way to do that is to basically take a shared nothing architecture, right? I can either put, you know, a client on the blue field or I can deploy my software, right?
On essentially the CPUs in these g um, GPU servers. Now again, problems there is I'm bringing the problems of that shared nothing architecture up into my most expensive assets, right? I have EastWest trap happening, right?
I might have a hotspot, right? Where everyone's asking about the same piece of context. That means I have all my GPU servers attacking one essentially and asking it for information.
I don't know that that's a good idea, right? So what we said is we've got, um, a, a different architecture, right? We walk through this shared, shared, uh, everything architecture where I've got this stateless layer.
Now I would say, you know, some potential challenges with this instance of the deployment, really two, right? And again, I think, I don't wanna say challenges, but optim areas for optimization one, right? I have this layer of CPUs that is essentially kind of in between my access to SSDs, right?
Um, and as we know, right? The CPU is always gonna be the bottleneck to SSD performance, right? You think about the world's most powerful processors.
How many do I need from a thread perspective to saturate one single 122 terabyte drive? It's a lot. So we have that problem.
The other problem is, you know, I have to essentially create a copy of data, right? I've got an RDMA operation to our front end, and then I've got another RDMA operation to the SSDs, right? So you kind of have this hop that's happening.
What we're able to do with I-C-M-S-P on Vast is actually take our logic, our C node, and move that up to run on the blue field. Now, we actually, uh, introduced a prototype of this style architecture, uh, I think with AI Field Day, um, earlier, but that was with the previous generation of Bluefield, right? Bluefield have gotten dramatically more powerful from a core count.
So now I can run my C node, the logic of the system up in those blue fields. This is a paradigm shift, right? I no longer have a host going through other CPUs to basically get in line to get access to data that's on fast media.
Now, every node has its own little friend. That's its protocol server. There's A storage class memory here, Phil.
It's still in the dbox down there. You just can't, it's like, yeah, there'll be two layers of storage in that dbox a storage. The old Way, the storage class memory was also in the dbox.
Yeah. Nothing changes there. It's just how I give access directly from the host to that device.
Okay? Yep. So we don't have to change anything there, right?
Which again, now, instead of having to need to use the local SSDs and introduce potentially, you know, conflicts and all the different things that we might have by putting software on those servers, I can just have JBoss, right? Full of flash with dense, solid M SSDs and everyone has direct access directly to the metadata structure and directly to the actual data itself. And again, scaling and everyone sees everything.
So there's no problem in sharing context, right? If there's a hot piece of context, everyone can access it with a, a whole bunch of parallelism. Uh, but I'm not gonna have any hotspots up top.
So Phil, excuse me, Phil, Jack Pollar with Paradigm Technica. It sounds like what you're really doing here is you running storage controller software on the GPU because you've got spare GPU cycle. It's on the blue field.
So I have a GPU server, I'm putting a blue field, which is like a smart nick now it's got 40 cores in it. We're taking those cores, which you don't need for network performance 'cause it's just more cores than you'd ever would. And we're running our storage software there.
So it's not in the CPUs, it's not on the GPUs, it's on this little server essentially that's mini, that's a smart nick, but a lot more powerful than that. Okay? And the net effect of this is, So ultimately what you get, right?
So we talk about certain things, right? One much faster time to first token, right? So if someone's asking a question, I now already have that context, I'm sharing it globally, right?
So if anyone's asked about anything, So in, in, in a traditional architecture, then you are making a request from the GPU to a storage controller that's off host. Yep. Right?
And then that storage controller goes fetch as the data feeds it back. Correct? And so in this case, what you're doing is you're moving that storage controller on host correct.
Or a little bit closer to the GPU. Correct. And that's getting you sign, that's accelerating significantly.
Significantly, okay. Yeah. And I think there's a few things, right, that come into play.
So you've got the acceleration of taking out an RDMA operation in the middle, right? Uh, you have a, a scaling advantage of the fact that now every time I add a new host, I'm adding compute specifically with that host. That is its own storage resources, right?
Essentially it's dedicated. So I'm taking out all the potential conflict, right? You get resources for you.
You don't have to fight over them with your partner. When I have that shared CPU pool, we're all fighting for the same resources, right? So it's a scaling, it's a parallelism and it's efficiency perspective.
The fact that the data is now shared means that I can have more GPU servers able to share more context. They're much less likely to calculate things again, right? So that's why I get a faster time to first token because I can pull that context without having to recreate it.
I get much better GPU efficiency because my GPUs are not recalculating the same things over again. They're actually doing inference instead. Right?
And then the final piece is I'm taking out that entire compute layer and that all is power that's drawn and power is precious now. So by taking out that entire server CPU group, I cut power by 75%. Got It?
Make sense? Mm-hmm. Okay.
So can can, can you tell us again what I-C-M-S-P was? It's context management something something. I think it's inference.
Context management storage platform. Okay. Inference context.
Memory storage platform is what they called it. And if you Google it, just be careful. There's a whole bunch of other uses of the acronym ICMS.
I-C-M-S-P. Just think of it as really cool close storage. Okay.
And I keep thinking about the image that you put up that had the different layers and had context as one of those layers. So, okay, so is this vast I-C-M-S-P, that's what's being attached to the blue fields. So essentially it's I-C-M-S-P is, um, think about it like NVIDIA announced something called Dynamo, right?
And we were working closely with them on that. You've got like all these different problems in terms of, you know, how do I manage where inference jobs run on GPUs, right? How do I make sure that I am, uh, intelligently using the different tiers of me of you know, memory and storage just for context, right?
So that's just for context. Um, and then how do I scale and run those things? And basically the I-C-M-S-P is a tier of storage and essentially a standard way that Dynamo's gonna interact with that storage.
So I'm basically saying we're lining up to saying this is how NVIDIA expects to use extended, you know, off, um, or shared storage for context. So can you go back to your diagram that shows? Yeah.
Okay. So where is it on this chart? So essentially this is gonna be used for context in this world, right?
The notes? Yeah. So all the da, all the, the, sorry I'm now, I'm not supposed to point to the screen.
All of the context gets stored down in that D box layer in the same way we would store any type of data. Okay. And so that is y'all's I-C-M-S-P is gonna be in the dbox.
Correct. Okay. Thank you.
And the data lake that's also underneath that is stored there as well. You easily can. Yep.
So, you know, something that I would talk about, 'cause I, I, let me go to the next slide and maybe it will, it will help. So one of the things that's going on right now around I-C-M-S-P and KV cache is the question, should you use any data services? Right?
Because could they impact performance? Right? And we went through that whole shared nothing thing.
And certainly if you do use data services on a shared nothing architecture, you're gonna have challenges. Data reduction is one of them. Another one is something like encryption, right?
So right now we're not sure should we encrypt that data. I think it's a really bad idea to not encrypt that data. It's a giant shared thing which has everything about every conversation that all your employees are having with ai.
You might want to encrypt that, right? And you kind of have to save it again. But there's also a chance for data reduction from what we've tested.
3 to one and two to one data reduction, which means, again, And again, this is an inference solution, not a training Solution. Correct. It's all inference at scale.
Um, so the point is, you know, in our world, this is how we would do it, right? But what's unique and flexible about the vast world is those Bluefield controllers don't have to be the only CPUs that that cluster has access to. We can create actually a sidecar pool of compute that just does data services.
Because remember, those blue fields are gonna write data down in let's say storage class memory. We can then have this pool of compute, take that data, reduce it, store it back down to the QLC, right? But I can also attach other workloads over here, right?
And what I know is that these blue fields all get their own dedicated amount of storage performance. Every blue field has 40 cores sitting. So In the other prior solution, the blue fields were actually responsible for the d duplication.
Yeah. We Without the other compute side correct. Cluster.
Correct. Which they, you know, at this point, this is a very new solution. We're gonna have to do a bunch of tests.
Will they run hot? Will we want to augment? Will it be enough?
Uh, but the point is, we're the only ones that have this flexibility to say we're gonna add compute that you can then leverage for data services. So essentially, I just care about the fact that these guys can access data and then let our friends over here take care of all the data services In the old world that these sort would've had to have been homogenous nodes. But in this environment you're taking, you could put any, any cluster of compute services out there to be your C nodes, correct?
Mm-hmm. At this point, it's very flexible. We have customers with multiple different generations of compute running c nodes in the same giant environment.
Mm-hmm. We can actually take, we can pool them, right? So we can say basically, you know, certain applications can use two C nodes and the rest can use 20.
There's a lot of flexibility in how we can carve this up, that disaggregated shared everything piece gives us flexibility in a way that wasn't possible before. And you mentioned storage class memory down at the Dinos. Those are, uh, different types of SSDs that are tailored to, you know, read, write access and things like that.
Right? So essentially it's an SSD, you know, lower latency, they're more expensive, uh, much better endurance profile. Mm-hmm.
Right? So essentially, you know, when we came to market in, uh, Intel, I'm losing word pcm, you know what it is? P pc MO Optum.
Yeah. Opt Optane, Optane. He, it died such a long time ago.
It was all we had. Soy has a, uh, P 58, 10 SLC based SSD that is used as a storage class memory solution for these guys. Mm-hmm.
So there you go. So yeah, it's all about endurance profile, cost latency. Thank Marian.
I have a quick question. Sure. Uh, going back to similarity, how do you keep the similarity decisions as the data and the models change?
Yeah, so that's really just the underlying data structure, right? So it's totally abstracted from models or anything else, right? So as data comes in, essentially we're hashing it.
And normally with ddu you use a really strong hash, right? That 'cause you wanna make sure that I never accidentally mistake two pieces of data for the same. That's why DDU uses a strong hash.
All we're doing is taking a chunk of data and using a weaker hash. And that weaker hash basically says that we don't use it to actually store the data. We use it to say, Hey, this data is very similar to data we already have.
So it doesn't matter if that data came from an inference job, a backup job, whatever. We will look across any data that's been stored on the system, um, doesn't matter what protocol it landed in. And we will identify that there's commonality and we just won't store it.
And the system has no idea this is happening. And you're using a weaker hash for, um, Comparison. Why?
Comparison? Comparison? Yeah.
Because if you use a strong hash, you can't tell, because when I use a strong hash, a small difference in the data, a results in a very different hash, right? That's why you do that. So a we cache basically says if the data's slightly different, then I'm gonna get the same result.
So now we know we're in the zone, right? That these two pieces of data are very similar. And what are you hearing from your customers who are in highly regulated industries?
They have no issues with it. Right? Um, again, data's now typically on a, on a system, it's chunked up, it's erasure coded, it's spread around anyway.
So at this point, you know, it's all generally pointer based. Uh, I haven't never heard anyone have issues that they would have to turn it off because of some kind of, Okay. You sure?
Thank you. As far as the KV cache and actually doing any special caching of the data coming off the blue field versus it's all going to storage class memory when it's written and it'll be get to QLC or whatever the backend is. Uh, it's not like you're holding that data in storage class memory or anything like that.
We are Not. So, you know, in general, KV cache, we'll use the different types of media available, right? So it could land in the memory, the high bandwidth memory, right?
Or it could land in, you know, a local SSD, uh, this is essentially another tier of KV cache. Um, and for us, yeah, we're, we're gonna keep all that metadata in storage, glass memory, but we, you know, there's, there's nothing different about how we have to store the data. Essentially.
We've just created a more optimal data path, right?