What is AI Ready Storage, with Hammerspace
AI Ready Storage is data infrastructure designed to break down silos and give enterprises seamless, high-performance access to their data wherever it lives. With 73% of enterprise data trapped in silos and 87% of AI projects failing to reach production, the bottleneck isn’t GPUs—it’s data. Traditional environments suffer from visualization challenges, high costs, and data gravity that limits AI flexibility. Hammerspace simplifies the enterprise data estate by unifying silos into a single global namespace and providing instant access to data—without forklift upgrades—so organizations can accelerate AI success.
The presentation focused on leveraging existing infrastructure and data to make it AI-ready, emphasizing simplicity for AI researchers under pressure to deliver high-quality results quickly. Hammerspace simplifies the data readiness process, enabling easy access and utilization of data within infrastructure projects. While the presentation covers technical aspects, the emphasis remains on ease of deployment, workload management, and rapid time to results, aligning with customer priorities. Hammerspace provides a virtual data layer across existing infrastructure, creating a unified data namespace enabling access and mobilization of data across different storage systems, enriching metadata for AI workloads, and facilitating data sharing in collaborative environments.
Hammerspace addresses key AI use cases such as global collaboration, model training, and inferencing, particularly focusing on enterprise customers with existing data infrastructure they wish to leverage. The platform’s ability to assimilate metadata from diverse storage systems into a unified control plane allows for a single interface to data, managed through Hammerspace for I/O control and quality of service. By overcoming data gravity through intelligent data movement and leveraging Linux advancements, Hammerspace enables data access regardless of location, maximizing GPU utilization and reducing costs. This is achieved by focusing on data access, compliance, and governance, ensuring that AI projects align with business objectives and minimizing risks associated with data movement.
Hammerspace aims to unify diverse data sources, from edge data to existing storage systems, enabling seamless access for AI factories and competitive advantages through faster data insights. With enriched metadata and automated workflows, HammerSpace accelerates time to insight and removes manual processes. HammerSpace is available as installable software or as a hardware appliance, and supports various deployment models, offering linear scalability and distributed access to data. A “Tier 0” capability was also discussed, which leverages existing underutilized NVMe storage within GPU nodes to create a fast, low-latency storage pool, showcasing the platform’s flexibility and resourcefulness.
Transcript
I'll give you a little bit of an overview of what Hammer space we'll be covering today, as well as our agenda of speakers. We have some other Hammer Space team members in the back, um, that are highlighted here on the screen. Um, but what we're gonna focus on for the AI infrastructure field day and really for how the Hammer Space data platform is focused, is what are customers looking for when they're, we're building out their infrastructure or perhaps thinking about their infrastructure strategies and how might it need to be different than how they architected their traditional IT infrastructure?
Most of the customers we work with in the enterprise AI space have a lot of existing infrastructure they've invested in. Um, there's a lot of data, of course, sitting in that infrastructure. And the problem that Hammer space is really focused on in AI is how can you leverage what you have today, both the infrastructure as well as the data and make it AI ready, um, leverage it in your AI workflows the way you want to.
And so that's really gonna be the focus of this presentation. Um, there's a lot of different elements of that, and the title of our presentation implies this, but customers in enterprise are also looking for that simplicity tool that AI is new. Um, most AI researchers, AI infrastructure folks are under a lot of pressure to deploy, get results quickly, and do things with, um, high quality results.
And that's a challenge for the entire AI f factory end to end and Hammer space does everything we can to help making the data readiness, the data that you're gonna use, that infrastructure project easy to get to, um, easy to leverage. And we really hear that from our customers a lot, that there's so many choices they need to make across their entire stack, the ones that are proven and can make it easy for them to deliver their results to their organization. It's really critical to them.
Um, so we all like to get into the technical weeds and talk about the neat kind of widgets and wingdings that we can do in our platforms, and I'm sure we will do a bit that, um, especially with questions that we'll get from the delegates here today. Um, but we'd just like to emphasize too that what customers want is something that will deploy easily, get their workloads up and running easily, and get their time to results very quickly. Um, and that's a big focus of our platform as well.
So how will we spend our time today? Uh, I'm gonna focus on our data platform software now grad. We're at the AI infrastructure events, so the majority of our focus will be on the infrastructure and how we deploy our software and why we have the model, the deployment options that we do.
But starting out with some explanation about what our technology is able to do at the software level is where we'll start. Then Floyd Christofferson, who's well known to this group, as well as to a lot of the folks who are probably li listening in, um, we'll be coming up to talk about Tier Zero and what is the Tier zero deployment option for, um, the Hammer space data platform and some of the coolness of is probably sitting in most customers data centers today, and an untapped resource that you can put to work in your AI environment. And then Kurt Kine, um, who is also well known to this group recently joined us from DDN, is gonna be talking about a future infrastructure, um, project that's underway.
Um, it's being highlighted at the OCP conference coming up. We launched a multi-vendor initiative around this about two months ago. It's called the Open Flash platform.
And he'll be talking about what is it, why is it really important to future looking AI infrastructure designs? So with that, I'll start with something just a little bit fun. Um, probably the cutest slide that we have in our deck and the funnest, um, what is a hammer space and why do we call our company a ha call our company Hammer space?
It makes it more memorable. I think sometimes kind of know some of the origin story, but if any of you are into comics and are delegates in the room, got a bunch of comic gear, and may have been wondering why they got a data dynamo action figure, like the gentleman, um, the sitting behind me, what is the idea of the comic books? And it's really that idea that if you are into comics and think of, um, you know, pulling a big mallet, a small mallet out of your pocket and becomes large or the pocket universe where mm-hmm Mary poppin can pull large lamps out of her purse, those types of things, that's actually a hammer, a hammer space.
The analog to hammer space, the technology company is we create a virtual data layer that lays across all of your existing infrastructure structure. And pro logically provides a unified data name space so that you can imagine a space like, um, enterprise AR architectures where they're looking to use data that sits in their NetApps that have been deployed for, you know, and built over the last 20 years. And then cloud instances have been developed for specific reasons.
Object stores that maintain a lot of their archive data, creating a hammer space that lays above that and provides a unified data layer that you can, you know, magically jump into and see what data sits and all those different storage systems mobilize it to where you want to be able to use it, enrich the metadata with maybe information that you didn't need when it was just a home directory. But you do need now that it's part of an AI workload, that is what Hammer Space provides. And so it really matters in these environments where you now have very siloed infrastructure that was designed for a one-to-one ratio of your home directories or your HPC system, your, um, financial services environment that was just accessible by the finance team.
Now all that data needs to be shared in an AI workload, and it's challenging. What Hammers praise provides is the ability to leverage that infrastructure you have, make your data AI ready, so the correct interfaces, the correct metadata access, the ability to move it to where your GPUs or your models are located, and make this kind of magical environment of using data that was created elsewhere for something else and available effectively immediately in your AI workloads. So, um, when we did our prep conversation with a few of the delegates from today, they suggested we really start out talking about use cases.
Um, so I'm gonna start there and use cases in ai. Um, hammer space kind of hit the radar of the AI market with some of the work that we did in large, the foundational model and large language model training. We did a lot of work with Meta, which is one of the public companies that we can talk about, but others of the very large foundational models.
And being able to just deliver screaming fast performance to GPUs. This is thinking of that kind of classical environment where you have tens of thousands of GPUs local to the data and local to the model that's being trained. And it removing as much network interference, network hops as much latency from that path to keep your GPUs as busy as possible was the primary focus of Hammer space early on in the language model training space.
That's not really where we're gonna focus today, and it's not where most enterprise customers are most focused. They're thinking more about inferencing training agents, and we're gonna spend most of our time today talking about that. But that's not to discount the importance of fast benchmarks, parallel file system performance, feeding fast GPUs.
Um, the market has really moved to where they really are thinking about more of these other areas where they have, from a global collaboration perspective, they have multiple data scientists. They maybe have agents and models sitting in different locations. They have AI scientists that are working with data sets with different interests than maybe the data scientists that are involved.
And making it possible for them all to share their data set in a single name space is really critical to making it where the different AI research and projects that are going on can go on simultaneously with the same data set. So we think about ETL that, that kind of extract my data set, then I'm gonna go transform it, do what I'm gonna do and then load it back in. That doesn't work very well in collaborative workloads.
You want people to be able to work in a single name space if data is updated or new data is being generated. It's not being done asynchronously, but it's being done synchronously within the platform. So that global collaboration is not just about the humans, but multiple models and the people who originally created the direct the data, for example, the home directory owners can continue to do their work.
It's a ReadWrite environment for them while the AI researchers are going, oh wow, there's some new data that's relevant to my project, and they can collaborate real time in that sort of way. Molly, um, yeah, to, to, I could see your customers experience all of experiencing all of these all at once right now. But is there any, are there any patterns?
Two questions really. One is are there any patterns like from customer to customer? Some are, are, is it, is the problem that they can't invest in all of this?
Or is the problem that, uh, some of them are more bigger than others or that they're all happening at once? That's the first question. The second question is, are different teams within the customers seeing different parts of this?
Um, those are all good questions and very real, um, issues that our sales team has to work through. So I'll start with a, a lot of enterprise customers don't have a big enough budget to just go build a whole new separate AI infrastructure environment. And that's where we, um, target from our differentiation.
You know, there's, each of us in the infrastructure space for AI have different approaches. We think about the customer who has the problem of, I really just, I don't wanna have to do a big data migration project and I don't have the money to stand up a new project. How can I leverage what I want to today?
That is not the primary buying reason for every single customer, but there's a lot of customers who are kind of like grappling with, I don't know how to do this. I and I don't have enough budget. So we like to start there for our ideal customers, simply because that's where we're most differentiated.
There are the customers, you know, like meta, you know, the cloud, the big neo clouds who are standing up net new infrastructure. And that's very competitive, um, with, you know, some of the other infrastructure providers and we do very well there. Um, but that's not where we're the most unique, if I were to say it that way.
Um, and then we do definitely have the customers when we work with our partners in the, the AI ecosystem who are trying to get a project off the ground and we get pulled in to help them get that project off the ground. So they have a new agent they're trying to train, um, they don't really know much about the infrastructure, they're thinking more about their agents and just need to get their data in and they'll pull us in from that perspective. So it's a big space, a lot of, um, different types of influencers and buyers and pain points, and we focus on trying to get to the ones who have existing data infrastructure they wanna leverage and then being in, you know, kind of the competitive route race for the net new deployments, if that makes sense.
Okay. Um, but yeah, every customer and the buyers are all different. Um, sometimes it's the IT team buying, sometimes it's the AI team buyings sometimes.
We were at the Reuters, uh, AI infrastructure event a couple of months ago, and that was focused more on the C-suite. And what they're looking for are kind of different. So we have to train on which parts of the technology to focus on and why.
Um, and this is more of a sales thing, probably not bad, interesting to everyone, but having the sales team stop and listen and understand what is the pain point for the customer before diving into all six of these use cases when the person they're talking to really only cares about the one thing. How am I gonna get, like, overcome data gravity, get my data into the cloud, that's all I care about, you know, so you have to stop and listen to. So along the lines of guys question, are there specific verticals you do better in than others or the needs pretty, um, similar across, across all verticals.
Um, as a company, we have a pretty wide mix of verticals and, but when you think about our AI deployments, we have in particular really focused in the areas of life sciences, financial services, EDA, um, some of the environments that have larger amounts of data that they're trying to put to work in not just an enterprise agent type of way, but they're maybe actually gonna use AI to help, um, further their drug research and those types of things. Um, so from a deployment model perspective today, we've had more early suc notable success in the larger data set environments. Um, this new, all the new workflows that you see coming out of the AI factories, um, whether it's our partner with Hitachi or work that we're doing with Nvidia or whoever that is, I think that's helping the environments where they have a little bit less data.
Um, their, their projects are still being developed that are kind of, I'd say more horizontal, but vertical wise, life sciences, we do a lot in HPC where, for example, Los Alamos is building their first supercomputer that also has to do ai. And how you build a supercomputer for h super computing versus how you build an AI machine is typically different. Um, and they're building their, their next one to do both.
Um, so we do a lot of work there where you're kind of spanning traditional compute models into AI compute models too. Okay, thanks. Mm-hmm.
I mean, it seems like what you're advocating for, what your product's doing is to take traditional siloed storage, maybe with other modern storage capabilities and maintain those capabilities while you continue now share that storage to be used for AI workloads. Do you not find yourself creating a situation where you've got a noisy neighbor problem? Well, um, no, because we end up taking over the namespace at that point.
So what ends up happening, we call it data assimilation. So let's say you have a nice salon, a NetApp and oh, the cloud and object store. We go in and we assimilate the metadata so that, you know, the, the basic metadata, this file, this object, this owner, this state into a single unified metadata mm-hmm.
Um, control plane. At that point, you redirect your home directory, use your home directories to hammer space instead of to the NetApp underneath, for example. Right?
And at that point, then you have a single interface to your data and your infrastructure. You don't have another side one coming in from the previous file system or namespace. Um, of course sizing is important.
How active is the dataset, how many users, what's happening with it? So sizing our metadata servers properly for the workload is very important. So as A value add, hammer space is doing storage io control as it comes into that namespace.
Totally, Totally. Um, and it's important. So do you kind of think about in the past, let's say you had 10 storage systems, you had to manage the IO for each of 'em and you, you know, how am I providing quality of service and all those things on each of 'em?
Now you do that through Hammer space and you say, everybody to my right gets really fast storage. Everybody to my LE left gets really cheap storage and we'll automatically place the data and provide that across the existing infrastructure instead of it being done silo by silo. Okay, great.
So it is a different kind of way of thinking about operating, but it's actually much more simple because you have one place that you're connecting your users to and one place you set the business objectives as far as performance, data redundancy, and then we place the data in the system for them. Okay. Thank you.
Mm-hmm. Alright. Um, so we talked about data silos.
Uh, I I'll just emphasize even though today we're gonna be talking primarily about data center type infrastructure or hardware type of infrastructure, um, the cloud has certainly exasperated the data silo problem that now even every region of the same, you know, whether it's uh, um, file system or whatever it is that you're running in the cloud or in as an object store, those are all silos too. So when we work with our cloud partners, our cloud partners like Hammer Space for a variety of reasons, um, we've have a partnership with Oracle related to their GPU compute environment where they had all these GPU servers that they had bought from Nvidia, and they had a bunch of, um, NVME storage that came with them, but they weren't actually using the storage. And those NVME servers, camera space makesa possible to add that as storage that they're presenting to, um, their customers.
When they sell a GPU compute environment, they also get a fast storage tier. That's super cool. But then when we work with Amazon and they're pulling us into an S3 type of environment, they're looking at, well actually each S3 region right now is disconnected and is a silo, but the way our users want to run is where they can have local, local data in different S3 regions, but present it to the applications as a single namespace and hammer space provides that.
So mm-hmm. Think about cloud as well as data center or hybrid cloud, multi-cloud that we're unifying across all of those. It's not just about data center infrastructure.
Um, overcoming data gravity, and this can be one of those terms that especially people have been in the industry for a while, will say, oh, is this kind of like a marketing thing? Um, but it's very real when you think about the issue of data has gotten very large. We've all had tech field days a while back talking about big data and things like that.
Network pipes that connect 'em tend to be smaller than the data sets. But now we have this very real problem of, that data needs to be used with GPUs that very likely isn't local to all of the data. So organizations really do need to figure out how do I identify data sets that I'm going to move to my GPUs 'cause I can't afford, it's not practical.
Put GPUs every single place where I have data I need to get my data to my GPUs. And so how Hammer space does this, and Floyd and Kurt will probably talk about this a little bit as well, is really around present the metadata in a unified way. That's what you're interacting with when you're figuring out which project you're going to go, um, and actually need, move some data to the compute.
And then when the data does need to move, do it very intelligently. Do it in a file granular way. Only move the data set you need.
Don't say, oh gosh, I have to take this entire project and move it and then analyze. And I actually moved a lot more data than I needed to only move the data on a file or object granular perspective to the GPUs. You can work with that data while it's in flight because you're actually working with the metadata and not the data itself.
So we can't come overcome physics, but we can be intelligent about which data is moving and how you interact with it while it's moving, you know, pre warming data sets and things like that. Um, and this is used all the time now that organizations are using neo clouds for compute organizations, maybe have one GPU cluster and one data center. And this is very common in enterprises.
The rest of their data centers don't have their GPUs. You have to get the data to them. So really, really important in that data mobility piece.
So what's the difference? You talk about moving just the metadata then you talked about moving the data. How do those two relate?
Yeah, so we are a parallel file system, um, from a, if you know, file systems, well Linux based NFS We are. Yeah. So we, um, that's, that's a good question.
So a couple of the Linux kernel maintainers work for us, the Linux storage server kernel maintainer guy named Mike Ster, and then the NFS Colonel Maintainer Tron, Michael Bust. Um, we put all of our intelligence from the client side, so the parallel file system client we build into Linux. Um, we do that so that any modern distribution, whether it's a or a Red Hat or whatever it is that's running on the, in the Linux machine, can tie into the hammer space data platform and do things like, um, sharing data across multiple locations.
You don't have that one-to-one map that you would otherwise have with NFS, that you can take the metadata out of the data path and have that screaming fast parallel file system. So that intelligence, which if you look at Luster or CCA or any of those parallel file systems is proprietary software you have to put on each client machine. All of our client machines is already built into the Linux that's running there.
Okay. And then to your question of metadata versus data, so metadata, as we all know, is small compared to the data itself. You know, it's a fraction of a percent of the size of the overall dataset.
That's what everyone interacts with, is the metadata server, the metadata control plane. And that has all the information about this file was created this date, it lives in this location. It has to be subject to these business rules.
You know, it can't move out of the country, it can't be shared for language model training. That's all done in the metadata and we replicate that locally, but it's very small. The data itself stays where it is until you have a reason to move it.
And that's the heavy thing to overcome from a data data gravity perspective. Does that make sense? Yeah, that Does.
Okay. And you divorce the, uh, data location from data access to some extent. So I very much data can actually be moved underneath, uh, a file access while it's being accessed, uh, to optimize it or, or to mm-hmm.
Whatever other requirements might be there. Yeah, And I mean, we could spend a lot of time, we do a lot of Linux community, um, educational events in this area is something called flex files, is what the technology is. Um, but that's what makes it possible that there's something called flex files in Linux where the file has isn't, is nimble enough to not have a one-to-one mapping to the network interface.
Um, and so you you can actually do that. You can have, Ray is connected to his files that he's always been connected to, but actually it's being moved over to A GPU to do something by an AI researcher. But as far as you're concerned, your files are still in your home directories, just the way you set 'em up, you can still access 'em, you can still read write 'em.
The fact that it moved off of the NetApp and is now in S3, you don't even know that that's happened, is, is very different way of thinking a mini to one use case of your data. Mm-hmm. And it's because of a lot of Linux advancements that we've done.
All right. Um, I'm gonna kind of skip past the maximizing ROI in GPUs because that's really the same thing as, um, being able to stream GPUs quickly. But the other piece is, you know, get data there fast enough.
One of the issues that we see in enterprises today is they've invested in a lot of GPUs, whether they're renting a cloud instance from one of the neo clouds, or it's their own GPUs, but they don't have enough data or they can't get their data there fast enough to take advantage of them. So there is a piece of this of just being very fast, um, direct IO as much as possible to the GPU, but it's also having access to enough data to keep those GPUs busy. And that's definitely something that organizations are grappling with right now.
And then as far as reducing cost, um, we saw early on there was a race to building foundational models and they were buying net new infrastructure. Um, those companies had a lot of money and they ran into two problems. Um, the next round of ai, they're trying to do more efficiently with maybe a little bit less budget, um, available.
And then power is also an issue that we all know, power constraints for GPUs is a problem. Anywhere you can get eek power out of the infrastructure to make it available to GPUs is a good thing. Um, or just eek power out.
So there's power available for the work that everyone's trying to do is a big thing. So we spend a lot of time on efficiency of the infrastructure, leverage what you have today, which is not just your storage systems, but what's in your, um, GPU servers, but also how can you do this more power effectively? And Curty Floyd will spend a lot of time on that, so I'm not gonna go too deeply into that right this second.
So those are our key use cases, um, and hopefully that sets the right foundation for which area of the AI workload hammer space focuses on. And this is an area that data really is often overlooked. And early AI deployments, there's been so much focus on which GPU will I use, um, which storage system do I need, which network do I need, which model am I gonna use and actually gain access to the data?
Um, wasn't the foremost consideration. And that's a lot of the metrics that you see here of, um, lots of these projects never make it to production. Um, enter only a, um, portion of the data that the enterprise has access to is being used in ai.
And that's, I run a podcast, which a couple, at least one of you has been on the podcast. But, um, this is the big issue that a lot of organizations have is they have their AI infrastructure teams building out their, how will I do ai? But then their compliance and government is, teams are like, Uhuh, you're not touching that data, or I don't feel comfortable with how we're gonna be able to audit what's happened with your data.
And so when you think about why a lot of enterprise projects aren't getting launched is the other, um, people in the business are putting the brakes on until they have an understanding of the risks and how they can make, ensure compliance with their overall business needs, that is really slowing things down. And so that has a lot to do with data readiness for ai. It's not just having your data in a fast enough low tier or in the correct location.
It's does it meet your business objectives for your AI policies? And so within what Hammer space does, we have, um, what we call them business level objectives. So you can set essentially their scripts that will run that meet your business needs that identify what should happen with your data in a variety of different situations.
And it could be you're in Europe and you have GDPR concerns and only certain data can pass country boundaries. Instead of your GDPR team having to manage that, you write a, uh, an objective within hammer space that says data of this type can move to another country. Data of this type cannot move to another country.
And then we know which infrastructure is in another country and can place the data in the right location. You can imagine in like we do a lot with pharmaceutical research, certain data can be used in the AI pharma research. Other data that has privacy protections built into IT cannot.
And that's different by comp Country company organization. And it's the same business objectives of where your data can be placed is a really important piece That wasn't early on, conceptualized a lot in AI because what was happening is they were building out an entire new infrastructure, identifying one data set they would load in that infrastructure and then they'd go off and run their AI workloads. Now that organizations are wanting to have access to more data, train their agents with certain data mm-hmm.
These kinds of, um, data readiness and topics have become more prevalent in getting a project launched. And Hammer Space does a couple of cool things. We automate it, great business people using plain language can say what they wanna have happen with their data, and then we automate that happening.
But we also stop the very risky issue of copying your data out of a system. So once you copy your data from the regulated system, how do you know what's happened to it after that? Where, where did it go?
Was it modified? What is your gold copy of data? So keeping it all in a single system where you can view where did the data go from an audit perspective, you can always know what your gold copy of data is, um, really makes things a lot less risky for an organization as well.
So don't copy data sets into your AI system and then up to the model up in the cloud. Keep it all in a single data platform and then ensure from your compliance team's perspective, you're doing what you want with your data as well as providing the speed and performance attributes of your data sets. So can an end user as you determine where the underlying data lives?
Or is this something that is pretty much some, uh, uh, like a hammer space admin would be the only person who could actually determine where the underlying data lives? So, uh, IT administrator of Hammer space at the end user? Yes.
So like if I, I was the guy who had system level access, absolutely. I can see where all my data is moved around, what are the attributes of the storage it's sitting in that type of thing. But if you were a scientist, you wouldn't know, like if Ray was, was a not technical person, but was actually a, I'm picking on Ray today, a genomics researcher who didn't wanna know anything about it.
He just wanted to know where his genome sequences were. He probably wouldn't know where it was, but his IT team would. Okay.
So he can't tell if it's in the data center next door or if it's, uh, like across the world. He, we, he would just see on his, whatever he is using his, his home directories, he would just see a file directory tree. And just like today, probably he doesn't actually know where that's provisioned, um, in the data center to what kind of storage system.
It's just he's given a share to leverage. It would look that same way to him on camera space. Okay.
Because I think in some cases, the end user is the only person who actually knows what the governance rules are for the data that they're manipulating. Mm-hmm. They need to fully communicate this with their IT department to make sure that it, the data lives where it could, Well, their data is still their data.
Um, so they map into it and their home directory is the way it always has been. Um, just like today, if you think about it has certain things they promise to the scientists, for example, of redundancy, um, retention for seven years. Um, they still get all of that.
It, it's just they don't just like end users don't manage their filers today. It does. It is the same concept, but they still set their rules on their own data as far as how long do you keep it, how available, how fast, that type of thing.
Well, but I feel like what Andy's trying to ask is, let's say you've got a distributed system where you've got, you know, a follower in Chicago and potentially an S3 bucket in, uh, Paris. You need to make sure that person X's home directory stays with data gravity into that Paris bucket. Mm-hmm.
The underlying system. Absolutely. Is that policy driven?
It is. Exactly. Yeah.
So, And how, how We guarantee to our end users, you can have, you'll always have this, um, level of, you know, latency and performance. That's still the same with Hammer space. Okay.
And how granular can you get with that? So let's say I create home directories. Okay.
Is one name space, do I need to go through and say, okay, everything that's in that name space is all going to be, have this policy applied to it or can you know, no, it's home directory slap Craig quite granular lives in Singapore. You know what I mean? Yeah.
It's quite granular. Do you wanna speak to that Floyd, on the variety of, I'm gonna ask Floyd. Yeah.
The, the other aspect of this is the, the power of custom metadata. Mm-hmm. So It's file granular, object granular, but, and it's not just one policy.
You could have multiple layers of policies. And so an individual user could set a custom metadata tag, which is automatically inherited through the hierarchy below that, that would determine this has, this particular data set has a, as a rules associated with it. And those would complement other maybe it level rules for compliance or sovereignty or other types of things.
So it can get very granular. So it's a balance between giving the individual user specific file granular control so they can do what they want, but then giving it the ability to create, you know, uh, enterprise level rules that apply across all of them. So it's this balance, and this is where all of the various metadata types that we aggregate into this unified data plane mm-hmm.
Is what makes all of this possible with minimizing the disruption to the user. Because any of that data movement in, in the background, which I'll talk about a little bit in the Linux portion of this, is completely transparent. Okay.
Thank you Molly. Um, the 87% of projects never making into production, are you claiming that it's all data read readiness or data compliance? No, that's a component of it.
These are just overall trends, but we definitely see that data is a portion of it. I mean, everything in ai, not everything, that's not fair to say. A lot of AI is still experimental.
How, which one adds value, which project is gonna move the needle? All those things. And data is a component of it.
You may have said this, but what's the source of this? Is this a a hammer space study? No, it's not.
This is different. Um, pulled from different, just publicly available sources. Okay.
Um, a lot of it is work that we do with some of our ecosystem partners like Nvidia and other of those is as we're all tackling, what does enterprise AI need to look like? I do have it cited in the, um, notes of the slides if you're interested. Okay.
Yeah, I'd be interested. Thank you. Yeah, yeah, Absolutely.
That's really interesting data. Yeah, it is. Yeah.
I'll, I'll make sure you guys get copy of the slides with the, where the dataset points were pulled from. And it's dynamic, you know, I mean this market's changing so fast. Um, every six months things get a little bit more, um, evolved.
There's better reference architectures and blueprints for how to do things. So, you know, the market is definitely maturing quickly, um, as new capabilities are also coming out really quickly. And that's, I I will point that that point of market is evolving so quickly is another reason why Hammer spaces relied on so heavily in these environments that organizations are reluctant to make big bets on AI that, you know, which GPU do I buy because it might be obsoleted in two years by the next biggest GPU, which storage system?
Am I gonna be running my AI in the data center or in the cloud? Will I be using cloud GPUs or my own own GPUs? These are all things that are just dynamic and changing very quickly.
On the hammer space side, we make it very possible to be flexible. You put your data into a single data plane data platform, use the storage that you have today, and then when you want to add storage capacity, that can be in the hammer space in another location or in the same one you can use cloud and flexibly come back into the data center. And if your GPUs are in other locations, again, your same data platform or data storage system is providing it your data to your cloud environment.
So that flexibility is really important to our customers that they don't really know what their AI environment's gonna look like a year or two from now. And we have flexibility that you're not putting all your, you know, rocks in one basket right now and wishing you had done something different a year from now. Um, so that, that really is a big portion of it.
It'd be interesting to see trends or predictions for these numbers which are gonna go up and which are gonna go down. It would be interesting. Yeah.
So, um, the way we, when you start to think about an AI workload, what we see becoming kind of in vogue is building out or accessing these a, um, AI factories. So many vendors, Dell, Hitachi, Nvidia Penguin, all of them are building out AI factories, which are this top part. I know I'm not supposed to move, but I'm gonna move for a second of where, how are your models interacting with your tools?
There's a lot of the upstream workload piece that we're supposed that organizations are trying to put together quickly and having nice ties into your data and your storage is super helpful to these environments. So you deploy your AI factory and yet you need to get your data to it and think about your data. It has different interfaces, different locations.
There's data being created at the edge that nobody really knows what it is until it's moved into a, a data center or, or it managed system. And yet these AI factories are ready to go. You have a fast way to deploy your AI environment with the AI factory.
What customers are looking for is the analog for their data. Great. I'm gonna go buy my AI factory from someone and I want my data easy and quick to load.
And that's really what Hammer space is providing us on the other side, connect your AI factory to hammer space will make it seamless that no matter if your data's being created at the edge, it's already existing or you have new data sets coming in, your AI factory has access to it immediately. Um, one of our customers that we have does, um, they build rockets and they have the desire to be able to make sure that their, um, team who's looking at failure analysis, which is different from their engineering team, their new product development team, which is different from the team that, um, is doing the AI research, has the ability to see the data coming off the rockets as quickly as possible. So you imagine you're test firing an engine, whatever you're doing, your AI team wants to be able to access that quickly.
This is considered edge data right now. Mm-hmm. Um, what we do is we assimilate the test fire metadata instantly, the AI team knows, oh gosh, there was an anomaly that occurred immediately and has the ability to say, I want that data of the edge as quickly as possible, get it moving.
And that's really kind of the types of things that you can see that's really important to be able to know and ingest your data as quickly as possible to get it into these AI factories as quickly as possible. And as a big competitive edge, you hear that's the types of words exact set our customers use is, is a competitive edge to know what data I have and be able to use the interesting data as quickly as possible. Um, and that's, that's something that we're enabling here.
There's New, new edge data they collected. The AI team wants it for training, tuning, rag, whatever their application is, or their service is based wherever. And what you're doing is finding the quickest way to get that data there first, exposing it through the metadata.
So exposure is instant, right? But let's say it's three terabytes of data gotta get next to, well, one's gotta move the, the services gotta move or the data has to move. Exactly.
And that's the cool part first exposure. And you said it perfectly. That's the first thing is like, oh my gosh, there's interesting data that exists.
Cool. I act actually know it exists now. Where in the past it could have been who knows how long by the time data was collected, copied, analyzed, and made available.
Mm-hmm. Um, if ever, you know, a lot of edge data always just sat on the edge. Um, so that's, that's the main thing.
But then Correct, you get the shared metadata instantly and then can make the decision. Do you use workload orchestrators to orchestrate the workload to the data? Or do you use data orchestrator or like us move the data to the workload and AI environments are doing both, right?
There's the ability to containerize and orchestrate a workload to an edge server that has enough compute to do it. Or do you orchestrate the data to a big, you know, centralized environment And we see our customers do both and we do some neat hooks with those workload orchestrators, which is the super cool that the decisions made upstream from us. You know, it's actually like a slum job underneath, but the workload orchestrator is saying, I'm gonna run a workload here and then it will through our A API trigger get this data there.
And so some of it may be, but there's three missing files, get it there too, or 30 terabytes, whatever it is. So that tie between the AI factory and the workload orchestrators, or GPU orchestrators to hammer space to invoke data movement is a really important, elegant part of the solution That that is super cool. Mm-hmm.
Is Hammer space enabling that decision though? Or, or is it The decision is not usually Here's like, here's here's, well no, but like here's an estimate of the amount of time it's gonna take or what we Can provide that's gonna be Yeah, it's pretty intelligent as far as, um, should I create a move? You know, let's say you have duplicate files, which is my fastest tier, which one do I pull from?
Mm-hmm. Like, there's things like that that we can do, um, as far as providing estimates based on our notes. I was giving an example.
It's just making the decision of whether to move the workload or move the data has a lot of factors. It does. And In your platform you have a lot of metrics that you can use to We Do, yeah.
We have a lot of metrics that we can expose and a lot of automation that can be done based on those metrics. Ultimately it's the business objectives of course. So who has control upstream of us of making the decision of where something should run and then us making the best decision automatically or exposing that to the administrators so they can make the decision.
Yeah. And that's part of the objectives, these business rules. And it's, it really is plain language.
And some of the metrics in there would be performance, performance versus cost performance versus, uh, priorities and other things. And so yeah, how the data moves when it moves, if all of it moves or only subsets of it move, all of that is a background operation that can be driven by a slum job through the API. It can be driven by a business rule, it can be driven by other automated workflows.
So it's really across the board. Well, I'm not asking if Hammer Space makes the decision or helps make the decision, how the organization makes the decision. There are more factors than that Yeah.
That, that you've described. Mm-hmm. But to make better decisions, you need, uh, more information.
And some of this information has to do exactly with the, the, the, the, the Hammer space platform would have a unique visibility into Yeah, exactly. And so you're saying that that information is, uh, retrievable. Absolutely.
And a nice ooey, a nice visuali visualization go. Well, Is it also available programmatically? It is, yeah.
Like through an API and that's one of the differences in the types of customers we work with. The, the big language model builders, they wanted to do everything programmatically. Okay.
We see you have the ability to go from like, on-premise to cloud and whatever. Oh, Sure, sure. No, I can, I I I can see all of that.
So yeah. And we fundamentally believe that most AI environments, there's so much to manage, you know, so many objects, so many files, so many use, so many requirements of the data that most of it needs to be done. There Was one more question over there kind Of related to the conversation we're having.
I think you said something really interesting that just really picked my curiosity. You talked about, um, the system identifying interesting information, and then we had this whole big rolling conversation about, you know, what's it gonna do with it and where's it, where's this data gonna get set? How is it gonna get to work?
How do you make that determination that the data itself is interesting? Can, can we back the truck up a Bit? Yeah.
And I, I'll probably ask Floyd to come back up. 'cause he's, I, I've always called him Mr. Metadata, he's like a fanatic about metadata.
Um, but it, it really is a, um, like he was, I learned about metadata from him ages ago. Um, but it really does come down to encryptable metadata. So okay, we will grab all the readily available file and object metadata, but it can be enriched.
Um, and it can be enriched by the end user. You know, so Ray, the genomics researcher that he's, is taking on that persona today could put in his own information. Um, but we will als we can also get it off of the instrument who created the data, or you could set a parameter that says, I'm interested in any data that meets these attributes.
And so there's a, it's all through enriching metadata. Yeah, but you wanna add to that? Well, I mean, specifically your question is about what determines the interest level.
And it's gonna be a combination of all the above. Typically when people think about storage metadata, it's age of the file, size of the file path and these types of things. But there's other telemetry involved in this case, like we were ta discussing earlier.
It's the performance of the particular storage that that lives on today. Or if there's multiple instance of the data. So map that back to a policy.
Or in the genomics use case, a big problem in research organizations is the data will come off of an instrument, it'll land on a Windows server somewhere, but then it needs to be moved somewhere else. And how do you know, you know, how do you know what it is? How do you know if it's important by being able to throw a, a, a custom metadata tag, which might be the, the grant number.
It might be a cost center, it might be the researcher's name, it might be the instruments, you know, firmware version, all of these things. And you don't have to rely on an individual to remember to tag. You simply put it on the folder.
The data lands there, it inherits that metadata. And that now triggers objectives that say, okay, this cost center data is for a chargeback, goes to this accounting, but for compliance reasons, uh, you know, a gold copy needs to be saved and in a worm copy on an object store somewhere. And so it's like you, it's a multiple layers of information that each individual, um, use case, whether it's the central it, whether it's the researcher, whether it's a compliance officer, they will all have their priorities.
And all of these things mesh together in the aggregated metadata that is part of that unified data plane. Did that answer your question? So if I can dive in there 'cause I Go ahead.
Yes, go ahead. Um, so back to the, um, the rocket launching or, or in test, uh, use case, right? So is the, is the interesting fact about that data that it exists or is there, right?
Like if there's a failure, then that data becomes more interesting? Potentially. Yeah, we wouldn't necessarily know.
We are not interpreting the data, but what we're doing is exposing to the applications that would analyze that data to be able to see it in place through its metadata. And then if it's a failure scenario that might, that might then trigger an action to move it somewhere else for post-analysis. And we have partners who their job is enriching metadata.
That's, that's their entire business. That's, there's ones that do like sentiment analysis or what's happening within a video, and all of that can be enriched depending on the vertical and the use case, which type of data enrichment, metadata enrichment they want. We have partners who do that.
So, you know, if it's a video, there's AI tools that will dig into the video and say, here's e you know, every clip that, you know, there was a fire or an explosion hard hat, and then somebody's not wearing a hard hat. And then we can trigger an objective based on that enriched metadata That enriched metadata. Then Anytime somebody's not wearing a hard hat, yeah, this needs to happen.
You're providing The information channel through this metadata. That's Right. It's basically an aggregation point where rather than each individual application having to go get a part of the information over here or part of the data over there, this unified data plane gives them from both the metadata, but also the data down below where it is, it's a unified control plane.
So any application can plug into it, no client necessary. It's, it's a, any application can plug into it at the API level or at the, at the protocol level. And that gives them the power.
So they're focusing on the data and not on the infrastructure behind the data. Quick question, Where does Hammer space sit? Does it wrap the pipeline?
So Hammer space sits, um, and I'll have another picture, but all of this here is the presentation layer to the AI factory. So, so Wrap all of that. Mm-hmm.
Okay. Yep. And I have some visuals here.
I think that will show that a little bit better. Um, so we, this is Hammer space right here, this blue section. So those of you who don't see me walking the screen, that's the blue section underneath.
So we incorporate all of this data and storage systems into one namespace that has the interface that the researchers as well as the AI factories want. So it could be, you know, like S3, it could be NFS, it could be MCPA, KB C, you know, like we present the same data set with the, to the in as an interface that's appropriate for whoever's interacting with the data. Does that make sense?
Mm-hmm. Um, so one data namespace provided in the way that the applications or users want to interact with the data. Does it have a cap?
Do you have a capability, let's say I do want that forklift upgrade. You know, I'm, I've got a field full of, uh, follower arrays in my data center and I'm starting to move everything out to the cloud. Mm-hmm.
Does it have a migration capability underneath the covers as Well? It does. Um, and it is a really good question.
Migration is one of those things that usually when companies have a budget, they're expecting a migration, you know, a storage refresh or I'm going to the cloud, or I'm getting outta the cloud. There's something else they're trying to do too. Mm-hmm.
Um, what we do is make it transparent to the business and the users. So first deploy hammer space, get all your data into a single namespace. And then if you do have infrastructure that's being refreshed or a goal to get to the cloud or whatever that is, users applications are up and running.
They don't have downtime. So think about, and I think most of us have probably been in the industry long enough to remember week long, month long, whatever migrations where you had to plan, plan this downtime as the data was being moved. There's no downtime to user in the business.
When you're with on Hammer space, you simply, you know, take your old system, whatever it is, present it in the hammer space data platform, add some more capacity or point to a new cloud instance and it can say, I wanna move the data over here, but the users and applications are run up and running and we get used for that quite a lot actually. Um, okay. You know, like, think there's a isilon's, a good technology, um, but there's some ice lawns that are 10 or 15 years old and the downtime to migrate it is intolerable.
What do I do? Well cool. Put in hammer space and I can do that migration without having that painful IT experience.
Okay. Um, and then we do some cool stuff too, you know, still make that hardware even if you've moved off your old Isilon, but you kind of still wanna use the hardware. 'cause really the disk drives are still working.
There's some life left in the SSDs, you can continue to use it. Maybe you'd use it for a less, um, important use case and make it like a little archive or something like that, which is kind of nice too. Mm-hmm.
Thank you. Mm-hmm. Alright, so I do want to, um, highlight this is kind of conceptual of what organizations are going through, if they are using a, like storage or infrastructure centric approach to building out their AI factories versus a data center.
So I'll just take a minute there. Think about historically the way storage systems have been built is you say, I have a certain workload and I need a certain amount of availability, cost per performance, cost per capacity. Um, there's some features that I need for, you know, snapshots, whatever it is.
And all of data of a certain type went into that storage system. Now when you think about these environments where you're, you do need those things, but you also need your data to be mobilized and available for other use cases. Having a data centric, um, environment and how you think about infrastructure is really important.
So customers historically have thought, okay, I'm gonna buy my Isilon for my video data, my NetApp for my home directories. You know, some kind of, um, like DDN for very fast performance. You've now got your data silos.
That's how that happened. And there was a good reason for it. Um, now they're saying, I have these data silos, but I need to use my data in other places than what was originally designed for.
I no longer have a one-to-one mapping of my HPC environment to my DDN and my home directories, to my NetApp and my video production, to my ice salon. I need to have a mini to one relationship. That typically means I also have the need to mobilize my data, whether it's to go to compute or it's to get to people who are living in other environments.
I need to do a data migration and I have the need for more speed. Um, a lot of these use cases, performance wasn't all that important. It mattered to your HPC environment, but pretty much every vendor had good enough performance for home directories and things like that.
Now you look at an AI factory wants to touch all that data, you need to have a solution of how do I get the data to the AI factory, have it move fast enough and u unify it and it teams are trying to figure out. What do I do? Do I have a big ticket?
A system where I'm manually copying data all the time and moving it to different environments and trying to identify what it is. I have some tools to do data discovery that's super complex. It's manual and it's fraught with the ability to introduce risk and errors, the wrong decision being made, data gain out of the auditable system, that type of thing.
So having a data centric environment. So instead of designing your infrastructure on I need my fast pass price performance, I need my low price per capacity, and then I need my high availability enterprise data features, you can really design it more on what are my data needs? This data needs to have a certain performance.
This data needs to have be available in four different regions. 'cause I have developers around the world. And then we'll place it wherever it needs to be.
So it's just a different way of thinking of no longer having silos, but instead having your data automatically where you want it. And it's a new kind of way of thinking for IT organizations. And Hammer Space was developing this technology before AI became such a big focus.
Mm. And it was interesting that AI seems to be the driving factor for why to change. I've done this for 10, 20, 30 years, I don't wanna change the way I do things.
Now. AI comes in and you need access to your data and your organization is gonna use agents and they are gonna do inference. You need to figure out a way to make your data available.
And doing it manually is just too hard, too fraught with errors. And so Hammer space, like it is, there's a moment here of why organizations are shifting over to this data centric architecture now. Yeah.
Um, and then that automation of the mobility is really a big deal. Um, the idea that you can manually extract, transfer, copy, move data sets in an AI environment where now you maybe have multiple models, multiple data scientists, and they're all trying to use the data. If you try to always move data to the data scientists, but the models are trying to train on it and the data's not where it needs to be.
Doing that manually also doesn't work. So the data automation is a big deal. And it's interesting when Floyd and I started at Hammer Space, we've both been here for a few years now.
The idea of orchestration was really new. You heard about it in the database world, but you never heard about it In this unstructured data world, orchestrating data was most people looked us and didn't really know what we were talking about. So we had to come up with other terms.
But now workload orchestration to your question earlier has come up too, that orchestrating these environments, that the applications, the compute and the data are all in the same place. And doing that automatically is how these workloads are now being built. Um, so it has to be automated and needs to be done intelligently and done through software.
And so before I hand things to Floyd, I'm just gonna talk about a few of our deployment models. Um, we do run as software only and we're built in with our, a lot of our client intelligence into the Linux funnel. And meta, literally all that you had to deploy when they were using us was putting metadata servers, we call 'em Anvils is our branded name for that.
We put metadata servers into the meta environment where they already had their compute nodes, they already had their infrastructure. All they needed was our metadata to aggregate, um, the data and the information and to pull the metadata out of band from the data. So if you're trying to, and for like scale out NAS systems like Vast NetApp Isilon, the reason that they're so slow in an environment like this is they're in band trying to shard and manage the metadata at the same time that the data is trying to stream between the compute environment.
We pull all that outta band so that data flows directly, like as fast as the network and infrastructure can provide it from the storage to the, um, compute notes. And so that provided them the ability to power whatever it was, 26,000 GPUs or whatever their number ended out being and keep them all busy, um, and use unifying their data sets. So, and that was a software only deployment model.
Um, taking advantage of our Linux capabilities. You shift over to Los Alamos and they are a little bit more of a traditional HPC environment. They have their supercomputers, they have both InfiniBand and ethernet networks.
They wanna use both of those and they want to be able to have the performance attributes to keep their computers available. But they also have the need to have an archive file system. They need to have a distribution file system for when they collaborate with other national labs.
And now they need to do ai. Um, which is another thing. And what they're trying to do is, as we know, our government funding has gotten difficult and challenging with, you know, the policy of, you know, policies that we have under Trump right now.
They're looking at what are new efficiencies, rethinking do I actually need seven file systems and now an eighth for my ai or can I collapse that all into one environment? So in this case, hammer space is deployed in an appliance with, you know, not just a software, but with a hardware appliance that has the capacity and performance attributes it needs and can actually displace both the luster as well as the GPFS, the NetApp they're using for home directories and provide all those different services with one data set. So one of the important things for Hammer Space is that we do have to, to accomplish this, we can't sacrifice on, I still need my benchmark fast performance for ML perf.
I also need my high availability across multiple regions for my archive. So we've built all that into our environment and in there's cases where customers really don't want to deploy a software. They want an appliance that you just take the thing, it shows up on your dock, you plug it in and, and it works.
And that's something that's newer for Hammer Space is offering ourselves as a integrated appliance versus just as software. Um, and different kinds of customers have different products, Data server and the storage is co-located in that environment. In that environment.
So you can, on our website, we call them, um, ready nodes and we have different configurations of, depending on how much metadata performance, how much capacity you need. Um, we have recommendations of which types of configurations based on super micro, um, is our key partner there right now, Hitachi as um, also offers them a Hitachi gear Ready, ready nodes. Mm-hmm.
Yeah. And you can build out like distributed networks of these of the ready nodes. Yeah.
As it is it one node, two nodes with a failover cluster? Yeah. So it's, these are just the traditional building blocks we have.
You can have multiple, you can have your multiple metadata servers, you can deploy 'em in different environments. Yeah. I mean it's really just, we had originally thought that customers would prefer to deploy themselves.
We'd tell 'em, you, your metadata nodes has to have this kind of, um, attributes, you know, this type of processing, memory footprint and your storage nodes should look like this. Um, and you, you can source it or use what you have. We're just finding that they kind of wanna be able to go to their VAR and say, Okay, the question is that if you have multiple nodes mm-hmm.
They actually are have distributed access to Absolutely, yeah. The The data Okay. Yeah.
To very massive scale or different redundancy levels and that type of thing. Yeah. Okay.
So both on the metadata level. So you could have multiple clusters, um, that multiple clusters, uh, at the metadata level. So you size it as deep or as wide as you need at the metadata level, um, for highly paralyzed workflows.
And at the data mover, the DSX, our data services nodes level, and it still sits in front of whatever storage they may have. But you've got, basically it's just another easy button to deploy hammer space and you scale it customized to your Okay. Particular use case.
Okay. You can just stay here. I'm about to hand it over to you, but to your point, um, so Hammer Space has linear scalability of, or up to massive amounts, thousands of storage nodes, which we've did, which we've exhibited as like meta and some of our larger customers.
So the scale, because we take the metadata out of band and the number of storage nodes we can support is, I mean, kind of just conceptually limitless. Um, and you don't run into that performance clip as you start to scale across and have more and more storage nodes. Um, you can continue to add a massive frame network of storage, storage nodes into Hammer space environment without running into performance bottlenecks, which is what most people think is gonna happen in these really large environments.
Um, and that's what parallel file systems do. Well, we do it well, luster does it well, technologies like that do that well. Okay.
But it is using like a distributed database for the metadata across the Yeah, it's not exactly a distributed one single distributed metadata database. Each anvil instance, each hammer space cluster has its own database within it. And then we synchronize across clusters, either through referrals or through a direct synchronization.
So if you've got, you know, we've got a customer that has 12 data centers around the world, all one single global namespace for build environments and for other things, while each one of those 12 sites is continually synchronizing the metadata, the data doesn't need to synchronize unless it has to move. So it's, it's not one single distributed database, but a cluster of interconnected databases. Okay.
And are those databases like eventually consistent or, Yeah, it's an eventual con consistent model because you can't have, especially in a, in a global deployment, well, Yeah, I understand. There's, there's no way to, But it's within seconds. And then within the hammer space, uh, you know, we have versioning and version control and other things.
So that eliminates collisions. Okay. Okay.
Mm-hmm. Thanks. And so the last item as I hand this off to Floyd, um, one of the big foundational models that released around Christmas time last year that was built on Hammer space, um, they took advantage of our tier zero capabilities, and Floyd's gonna talk about what that is.
Um, but I just wanna introduce kind of, this is a really interesting, um, way that we see AI environments, um, evolving. That they had a deadline to get a new model trained in a short window. Um, they probably didn't have budget to go buy a new storage system, but even if they did, they didn't have time, they didn't have time to go source the storage, have it shipped, spin it up, get it configured and delivered.
You know, the, the, their training timeline was just too short. And so they went, reached out to the vendor community and said, Hey guys, we have this problem. Um, our boss, a very well known guy in our industry who builds AI models said, I need this by the end of the year, what can you do to help me?
Um, we took a look at their environment and they had, I don't remember how many GP units No had, but thousands of gp No, that actually all had but NVE in each of 'em. And we said, well, how about we use your NVME and those GPU nodes, we'll assimilate them into the hammer space namespace, create one namespace across these thousands of GPU nodes. So now you have all these, think about, these would historically have been silos of, you know, we would call it hyper-converged or direct attached storage or something else.
You know, it would be terms we've used. Now those are all just members of shared storage that we were presenting to their training environment. It took them a couple of days to get up and running.
They didn't have to ship hardware, they didn't have to configure anything. And now all of a sudden they were able to use that NVME instead of a bunch of like local small store SI storage silos as one big, super fast, low latency storage pool. And they were able to start training within days and get this model deployed without buying any new infrastructure.
'cause they actually already had it there and were largely, it was unused.