Taming Data Estate Chaos for AI with Hammerspace
Hammerspace introduces itself as a “data company,” distinguishing itself from traditional storage vendors by offering a solution that addresses the complex data demands of modern infrastructure, particularly for AI workloads. The core concept behind Hammerspace is an instantly accessible, infinite virtual space that disaggregates data from its underlying infrastructure, enabling it to reside in any location, across any cloud, and on any storage backend, thereby eliminating data silos. This is achieved by assimilating metadata from existing storage systems into a single, global namespace, managed by metadata servers outside the data path. This approach not only accelerates data pipelines but also enhances existing infrastructure and enables rapid, easy integration of new technologies, providing users with visibility and access to all their data within minutes, rather than requiring lengthy, costly data migrations.
Hammerspace extends its capabilities to address critical challenges in AI infrastructure, including the current tight market and rising flash memory costs. The solution leverages underutilized flash storage within existing environments by aggregating systems and intelligently orchestrating data placement across tiers. It introduces “Tier Zero,” which consumes and aggregates local flash within compute (CPU and GPU) clusters into the global namespace, providing extremely high-performance storage by eliminating network latency. Hammerspace also treats cloud storage as a direct extension of on-premises infrastructure, not just a destination for data, thereby maximizing the use of available flash resources. The software-defined platform ensures data portability and access through a parallel file system (PNFS v4.2) and multi-protocol access (S3, NFS, SMB). Importantly, its policy-driven orchestration automates data movement and ensures data durability and availability through redundant metadata nodes and erasure coding across storage systems. It also centralizes privileged access and security policies, allowing permissions to follow data regardless of its physical location, critical for cross-border data compliance and auditability, and supports rich custom metadata beyond basic POSIX attributes.
Customer examples illustrate these benefits, such as a digital payments company that reduced storage costs by $5 million and simplified workflows for 3,000 data scientists by providing parallel file system access over object storage and enabling hybrid cloud agility. Another customer, facing a 3-4x increase in performance demand from new NVIDIA servers, leveraged Hammerspace to maintain existing NAS systems while deploying high-performance NVMe storage, avoiding significant new infrastructure investments. For inference workloads where latency is critical, Hammerspace can use policies to preload entire projects into local NVMe (Tier Zero) directly connected to GPUs, maintaining high performance and data consistency across globally distributed inference farms. Ultimately, through its integration with platforms like the NVIDIA AI Data Platform, Hammerspace goes beyond merely unifying data access; it truly unlocks the value within data by automating data preparation and orchestration, moving organizations from data chaos to a state of AI-ready data, often allowing interaction with the system via natural language for streamlined management.
Presented by Kurt Kuckein, Sr. Director AI Product Marketing, Hammerspace, and Sam Newnam, Sr. Director – AI Solutions, Hammerspace. Recorded live at AI Infrastructure Field Day in Santa Clara on January 29th, 2026. Watch the entire presentation at https://techfieldday.com/appearance/hammerspace-presents-at-ai-infrastructure-field-da/ or visit https://techfieldday.com/event/aiifd4/ or https://hammerspace.com/ for more information.
Transcript
Hi everyone. My name's Kurt Kine. Uh, I am Senior director of AI Marketing for Hammer Space.
Um, and joining me on the call today, Sam, if you wanna introduce yourself. Good morning, Sam Newham. I, uh, run our AI solutions practice here at Hammer Space.
We appreciate the time today. Yeah, really excited to be here with everyone here in the room and online as well. Um, here's a brief agenda.
Uh, so I'll be going over a quick recap of just what is Hammer Space? What is the hammer space and what is Hammer space, um, and how we do what we do, because it really takes a little bit of time to understand how we're differentiated from the typical storage company out there. Um, we see ourselves much more as a data company.
And so, um, it's an important difference, especially with today's workload demands and, um, infrastructure setups. Uh, then I'll kick it over to Sam and he'll be covering our Nvidia AI data platform integration and why that extends what Hammer Space does into a whole new arena for folks who are deploying AI factories, um, and are looking to go from data chaos to AI ready data. And then finally, um, I will give an update on the open flash platform.
Uh, we talked to the delegates at a previous field day, um, about OFP and the initiative behind that. And we want to give you an update on practice. And rare for a software company like Hammer Space, I'll even have some hardware to feel touchy.
Yeah, it's exciting. Uh, and I'll be joined by, um, two folks, uh, Ted and John from Hammer Space to talk about why they're a key part of the initiative and our, uh, first reference design. Alright, so, so, uh, the hammer space, probably not so much for the people in the room.
Maybe a few people in the room haven't gotten this, uh, background before. But what is a hammer space, right? And it is that infinite space that you see in an animated show where somebody's carrying a bag and then they maybe pick, pull a giant hammer out of nowhere, right?
It's this instantly accessible virtual space that is infinite in size. And so that's a pretty good analogy for what we do for our customers, right? We disaggregate the data from the underlying infrastructure and allow that data to be in any location, in any cloud, on any infrastructure backend, and eliminate the silos that are so common in today's enterprises.
And we accelerate pipelines by a aggregating the metadata that is attached to all of that data across the data state, making, you know, the infrastructure more valuable because we can continue to use whatever data you have in place and continue to use those systems, but then also rapidly deploy newer, cheaper, faster technology, integrate those really easily into your environment. And with AI being such a big part of where we're generating data today, um, you know, the need for this massive repository that sits across, um, locations, clouds, and, um, systems is really, really key. So a little bit of, you know, background on ai, right?
I don't think this is a question too much anymore. Um, but I think we often focus as vendors on our really cool system, right? And we don't focus so much, especially in the storage space on just how are people dealing with their data, right?
And AI has really driven us as storage players to think differently, and we're not alone in that, um, at Hammer space. But, you know, really going from where we were just serving files out, serving objects out into being able to serve something that's really useful for these AI applications, right? We're not just repositories anymore.
Now we are really, you know, full data management systems. And what Hammer space does is go beyond the system's view, right? And we want to give people visibility into all of their data, not just the data sitting on one system.
We want to be able to give people access to all the data across locations, give people access to all the data and visibility across their clouds. And we need to be able to automate those data pipelines to provide AI ready data, not just raw data out the front end. And we need to be able to do that in real time.
And so I've covered these problem areas, but I'll, I'll dive into each of them a little bit more, right? So storage folks have been deploying systems for all sorts of rational reasons all over the place, right? And each system might have a unique characteristic, right?
You might deploy object storage over here because you have this data repository that you know is gonna grow and you don't want it to, you know, have onerous overhead. Uh, you want it to be easy to deploy, scale out. Then you have NAS systems that are scale up maybe over on another side because you had more, uh, enterprise, you know, requirements for data protection, um, and things like that.
And you wanted it easy to share and easy to access for folks, um, that didn't know anything about S3. Uh, then you right, started deploying pieces in the cloud because there was this infinite bucket in the cloud that allowed you to ship data up there. But each of those new systems introduce complexity, additional data management overhead, um, you know, all the things that comes along with distributed data.
Um, similarly, right? You stand up a western region, uh, you've got another office in the east coast, you've got another office in Europe, another office in Asia. All of these disconnected sites and need to be able to collaborate, especially now that we're applying AI applications to be able to gain better insights into, you know, everything that's spread out across the globe.
So how do you unify that without massive migrations and multiple copies spread all over the place? Additionally, we encounter more and more customers. Yes, Ray ese, Silverton Consulting.
Um, you mentioned that you disaggregate data, but you aggregate metadata. You wanna kind of, I will go, I will dive deeper into that for sure. Okay.
Um, but essentially, right, that is hammer space's superpower, right? So we assimilate the metadata from the underlying storage system. So whether it's an object storage or NAS system, um, from other system vendors out there.
And then we have metadata servers that sit outside the data path that manage all of the data across those systems and aggregate it into a single global namespace. And then we can have data writing to those systems. So you can leave the data in place, or you can move the data based on where it needs to be, when it needs to be there.
Um, and you can continue to utilize that underlying system under hammer spaces ownership. That assimilation process happens instantaneously, or, uh, It's of course not instantaneously. It's, it's a day's, a day's versus month thing.
So when you look at, you know, what is a typical system vendor gonna come in and tell you is, you know, what, how are we gonna solve this problem? Well, we have the best system, we have the best new silo that's gonna solve all your silo problems. So just deploy our system, migrate all your data into our system, and then you have a single global namespace.
Instead, we say, leave all your data in place until you need to move it. So yes, there are cases where you're gonna have to move your data and that's, you know, physics, but you don't have to move it all instantaneously. So that's a nice thing that when Hammer space comes in, you know, we're talking days to be able to stand up a global namespace as opposed to a gigantic migration process into our system.
And so you get access to this global namespace really, really quickly. And I talk about that a little bit later in a couple of customer use cases too. Okay?
Um, so, you know, too many storage systems to manage too many things out there. Um, none of them meet all the needs, right? I mean, as we're approaching, um, you know, infrastructure modernization, I think what is everybody trying to get to?
And that's a single converged infrastructure. You need to be able to serve S3, you need to be able to serve, um, a parallel file system. And then also, you know, N-F-S-S-M-B-C-S-I, um, MCP, all of these things need to come from a single system.
And we do all of that natively. And then we're very, very cloud-friendly. We were built a software with a cloud first mentality, so that to be able to, rather than viewing cloud as a destination where you just spill bits out and then hope that you never need to bring 'em back because that's a really expensive product process, you know, use the cloud as an extension of your on-prem infrastructure, right?
And that speaks to, you know, the whole first to cloud type of, um, activity. And we can really help customers with an on-ramp to hybrid cloud settings. And then finally, data management across systems has often been another new tool that you need to bring in.
And so instead, you know, we bring data management into the storage layer and help people automate that via orchestration. And I'll go into that here in just a second. Is Marian Newsom?
I have a quick question. Yeah. So you talked about aggregating the data from the metadata.
How do you detect the drift if there's any between those two? Um, I will kick that one over to Sam, Sam if you're available for quick. Yeah.
The, the idea is to prevent data drift, right? It's one of the most POIs ous things that happens in ai. So by owning the metadata, right, by separating that particular thing, we synchronize those changes real time, right?
Across our global namespace. And so we don't have some of that acknowledgement wait for this to write type situation. And as part of our secret sauce is, is how we can resolve some of those discrepancies in multi writer situations and those types of things as well.
But our system was tuned to deal with this first studio in the sky where we're doing multi editing, right? Where we're doing AI pipelines, where data needs to live in multiple places, but stay and think. And so the beauty of that is that by suffering metadata, we know those changes can happen instantaneously while we deal with the payload of those files moving to those geographic locations.
So we can actually do, say a Ls on both sides of the world, see everything perfectly, but the data may still be in transit from a physical standpoint. Thank you. So one other emerging problem that we've got in real time, um, is dealing with flash memory, right?
And there's currently a pretty tight market for flash memory, whereas maybe a year ago storage vendors were out there saying, oh, disc is dead. Well, now that nobody can actually buy flash or flash is going up to X in price because of this really tight market, um, we've got issues, right? And so if you can't even acquire new flash, but you still need high performance storage, what do you do?
Right? Um, and you know, there's all sorts of, um, ways that the system vendors are going about this. You know, let's go out and do a buyback program so we can get some more flash in and then resell that to other customers.
Um, you know, I I think it's a pretty, I don't wanna make light of the situation 'cause it's a pretty dire situation for, um, customers who really, really need that capacity. Um, but you know, our view is that there is actually underutilized flash storage probably out there in many environments, right? You've stood up these multiple systems, they're all disconnected, they have an application running on system one, you have application B running on system two, application C, and all of those consumption patterns are not the same.
So if you can aggregate those systems and use orchestration to make sure the right data is sitting on the right tier of storage, then you can free up some flash space within your environment. You can make sure you're utilizing all of that flash space that it's already sitting there. Um, within the systems you have, we also have the capability, and we talked about it at the last, um, field day presentation, uh, a relatively new one that's, uh, was introduced about eight months ago, and that's tier zero.
So that's the ability to consume the flash that's within the compute cluster, CPU and GPU co clusters and aggregate that within our namespace. And so we set that up as a tier zero. It's a high extra high performance tier because you eliminate all of the network latency.
And that's sit, that's the flash that's sitting in the GPU and CPU servers. And we can incorporate that as part of the namespace and use, utilize that as a tier of very, very high performance storage. And if you look at, um, most of the GPU and CPU clusters out there, they do have underutilized local flash storage that we can access and add to the global data file system.
And then finally, I mentioned this already, um, but treat cloud as an extension of the infrastructure, not a destination. So, right, the cloud vendors have pretty good allocations of flash storage, but they aren't running into just availability problems. Now, probably not cheap, but at least you have access to that cloud instance, and you can treat that as an extension of your namespace that's both on-prem and in the cloud.
So how do we do it? So you can deploy hammer space, it's software defined. Um, we do have reference architectures and look for more news coming soon on other ways to deploy hammer space.
Um, but we can deploy on virtually any server in front of any storage, including local NVME as well as anywhere, right? Within an environment. So data in place assimilation here was Ray's question, right?
So data stays in place and we assimilate the metadata into our anvil server. I'll have a little diagram of that here soon. And data is visible to users within minutes, right?
So again, we're not doing an onerous long-term migration proc process. We're assimilating that data in place and then users see everything that they're authorized to use. And from there we provide a parallel file system access to a cluster or standard S3 N-F-S-S-M-B access, um, to other clients that are, you know, heterogeneous within the environment.
And so we can accelerate data that has been sitting on legacy file systems or object storage that's not tuned for high performance workloads that the GPUs demand, while also managing all of that storage underneath. And then finally, right? I think the most important point here, and the thing that we are going to exploit when we talk about, um, the A IDP integration with Nvidia is all of the automated orchestration that we can do underneath the scenes, right?
So as Sam mentioned, we can move data right across locations, between systems, even doing so while it's being written to, and, um, we can set up policies based on, you know, normal business language that allow people to tier their storage as necessary based on their individual requirements. Hey Kurt, it's Ray Ese, Silverton Consulting, you mentioned earlier MCP. Do you provide, um, MCP server for agents to be able to move the data based on where they think they might need it, things of that nature?
We do. And so Sam, that's exactly what Sam's gonna be talking about with our A IDP, um, solution. So that's where it really turns into something truly, I think, mind blowing, where you go from this really, you know, complicated management schema and, you know, none of this is easy, right?
I don't wanna make it sound like, you know, under the covers everything is super seamless, right? But we go from something that, you know, takes some storage management and storage intuition and experience and are able to turn it into va basically a, a system that can interact with you via natural language. So we'll talk about that a little bit.
So, um, here are the building blocks, right, of the hammer space deployment. So you've got the metadata control plane with the anvil servers, and then essentially the data service nodes. And those are the things that provide file access as well as do the data movement.
And so we can use PNFS, right? 2, which is integrated into every current Linux version that's out there today for folks to do high performance parallel file system access to the data within a hammer space namespace as well, again, as your standard multi-protocol file and object access for folks who are doing heterogeneous clients, you know, for all sorts of, A little bit about how you've clustered these various different types of nodes as well. I mean, I I assume that you have some way of clustering them so they, there's redundancy and, um, you know, Yeah, Sam, you wanna The more availability Yeah, exactly.
There, there's h chain durability, right? Built into the architecture of these platforms. And so part of this thing kind of like a, a ring architecture, right?
And so part of those anvils will live in each site, right? Controlling local metadata and the storage that's there and those, you know, and then replicated it at other areas in those particular sites, right? And so that's kind of how Kerr talked about how we do this out of band, right?
But yeah, we, we cover all sorts of reliability and durability, right? 999 however far we get in that particular architecture, right? Like most systems are always a trade off between, you know, how transients that data, how much do we care, where do we want to do, especially with cloud resources, but our architecture is fully redundant a across the entire platform, I guess, you know, follow on to Andy's question.
Um, so the redundancy, so some storage systems have redundancy built in, some do not. Um, you're talking, I assume about the anvil metadata being redundant. I understand how that would be completely under your control.
But let's say I have some, some dumb server storage and I want to have some sort of, uh, high availability for that. Do you provide, uh, uh, I guess I'd call it external rate across? Yeah.
So we do eraser coating across those stor. We can do it across storage systems, right? To provide redundancy Okay.
For the data. Yeah. The, the other point real quick is that all this is done via policy and orchestration.
So by share, right? So not even by storage box, right? But by share by data type, we can choose to keep multiple copies of that data spread across different geographic locations or to Kurt's point, whether it's a ratio coded or whether it's stored locally across a handful of servers.
So we've got a lot of flexibility in how we do that very granular, right? So it doesn't even have to be a master policy. We can do that specific to projects specific to shares or specific to organizations.
Yeah. I my question was more specific, specifically how you handled the cluster of your, your metadata nodes and you know, the, I I think that you answered it adequately that, um, you know, there's, they're fully resilient. You, you, um, if you lose any one of them, you still find, Correct.
Yeah. They're, today, they're in an H eight care. We're changing that to be fully containerized and scale out.
So we'll be able to add performance as we add containers. Um, same with the DSX, right? Those data movers and some of those components, right?
We can scale those today. In fact, we already have some customers that do an early access in that containerized format. But today it's ha But to your point, that's fully replicated across our environment.
Marian Newsom, I have a question. How do you prevent, uh, privileged access on those different sets of data? I'm sorry, can you say that again?
How do you prevent, how do you keep the privilege access the same across those different sets of data like the S3? No, exactly. So, so this is kind of the beauty of a single namespace, right?
Is that right now those are all managed silos, right? We see people sometimes do direct permissions into an individual NAS active directory, those sorts of things. By actually tying those directory systems into our platform, we a inherit those default permissions, right?
So that's part of understanding the POS landscape perfectly. So we can pull those particular things in. But, and I'll talk about this in my section a little bit, but the beauty of what we're, we're aiming for, right?
Is that security, is it tied now to where data's born that that security can now follow that data? And that's the beauty of separating the metadata from the storage itself is that now those permissions can follow that data whether they live in the cloud, live on a different storage platform, et cetera. And so we move that security plane into our metadata global namespace realm.
And so if I'm a, a company with cross-border data, uh, concerns or AI regulations, uh, that's what that solution will provide Exactly. We can set up all sorts of policies, whether it's GDPR to prevent things across country, whether it's certain patent information, we can do some custom metadata tagging. In fact, we're actually doing some integration with some, some digital security posture management companies and stuff as well, right?
The tiki, even our file system attributes well beyond that to be able to really control at a granular level what data should or should not live in certain places. Ever audit it. Could I ever see the logs to that?
I'm sorry If I get audits, can I pull the logs from my regulators? Correct. Yeah.
The, the whole thing is fully audited, right? Where files have moved, who's access to everything, that entire premise is built into our solution. Thank you.
And I think there was an important point, um, that, you know, uh, Sam made there that may have been lost and I hadn't mentioned yet. And that is, we definitely go well beyond a storage system in terms of the amount and the type of metadata that you can attach, um, to any object within our, um, global namespace. So that allows you to take additional steps, right, based on business rules to be able to interact with that metadata.
So it's not just, you know, your posits metadata of file create and permissions and things like that, but you can attach other attributes onto that data, which again, becomes important when you get to managing a system for ai. Um, I'm gonna go through this quickly here, um, just so we can jump over to, um, the next section. But, um, I want to talk a little bit about how our customers are using Hammer space today.
Um, here's an example of a large digital payments company. Um, and they had a team of about 3000 data scientists and research engineers who were building AI applications, and they were looking to move off of their fairly expensive up for renewal n systems, and they wanted to deploy object storage, cheap and deep storage, you know, cheap and wide storage, I guess, um, for their data scientists. And they started this process and then figured out that, oh, our data scientists don't speak S3, they speak NFS, you know, so what do we do here?
Right? And, you know, there's systems that you can put a, um, bridge in front of an object storage, you know, those are integrated to more or less extent, but typically don't provide you very good performance when you're accessing that object storage through some kind of gateway, but instead with hammer space, right? We can put that in front of the object storage, still take advantage of that scale out architecture, but now provide parallel file system access or a cluster as well as all of the other, um, access protocols natively, right?
And so we also extended then into GCP so that they had burstable capacity for additional GPUs, um, when and if they need them, um, as well as now a fully tiered environment that was NVME or tier one storage, and then this object storage on the backend and cloud as needed, reducing storage costs, their renewal license basically, um, came down by $5 million. Um, you know, the other thing is now their users are only looking for the data in a single place, right? They aren't looking across file systems, across object storage, it's single global namespace, no matter where the data lives greatly simplifies their workflows.
Um, and now they have agility to be able to utilize both on-prem and cloud resources. Um, here's another customer who ran into the performance issue, right? They were looking to deploy this new set of NVIDIA servers that now demanded, you know, three, four x the performance that they had been providing from their storage systems.
Um, but they weren't looking to, you know, up their storage by three to four x. Um, so instead of buying a bunch more storage systems or perhaps, you know, some kind of bespoke parallel file system solution, right? They brought in hammer space, we could still continue to leverage the systems that they had invested in their NAS systems.
So those stay in place and be used both with data in place and then also as an extra archived here. And then we just deployed our own NVME storage behind. So I did want to emphasize that we don't just use other people's storage as a, um, backend, we can deploy our own storage as well.
So we do have, um, storage boxes to be able to support customers in their deployments. And that's just standard commodity based storage servers. Yeah.
Product Van Hern from, uh, ENS Consulting. So, um, I, I do definitely see the value on the training sites, you know, with the metadata being available all over the place, but your data might be delayed. Mm-hmm.
Uh, how does that work for inference? Do you see the product working well on, on inference considering that your data might be slightly delayed? You know, where we talk about a little bit latency, where on the inference side, you know, latency is kind of a key, uh, key factor.
Yeah, I think, um, I mean, I guess fundamentally, um, what are you changing from data being remote or local, you know, in the situation, right? We are gonna bring the data as close as possible to the GPU, right? For, um, you know, again, under the covers can do that based on policy as a object is, you know, accessed.
We can then bring the entire project into the local NVME. Um, so I don't think we're necessarily added disadvantage. I mean, yes, if you had migrated all of your data into a bespoke file system, right, and then connected that up to your GPUs, um, then yes, but I think then you're looking at a massive one-time migration, um, which takes time, right?
So it depends on where you wanna spend your time, I guess. So if I understand it correctly, there is a way to create a profile where you kind of preload the data. So Sam will talk about that.
Okay. Here in Jeff. Yeah, Frederick, we can spend some more time over this.
We do have a number of customers that are using us for global inference, right? So when they're using tier zero for incredibly pa fast performance, right? Holding model repositories information that's been created, we then use a bunch of policies to drag that back, right?
To continue to do some fine tuning or some updates of those things. But they also use us to, to sync vector database information to those sorts of things. And at that point, we're up against the speed of light problem, right?
But they still find value in the fact that this all happens through policy and automation rather than this like spray and RC problem they've been running against. So I can show you a couple of designs where we're helping those particular inference farms, especially when they're trying to, to your point, keep latency low by moving inference close to the user, but still keeping data sets intact right across those particular farms. Yeah.
F**k you. All right, I'm gonna skip over this slide. It's what we do.
We're helping tame the data chaos with a single unified global namespace. Now you'll notice here, right, I'm just talking about the unified namespace plugging into an AI factory, but there's an important part of the prep and rest of the data flow, um, or the pipeline, right? For AI data that maybe we hadn't touched in the past.
And that is getting data truly AI ready, right? We were a storage system that provided data, raw data at the front end. Now with A IDP, we're truly unlocking all of the value that is within the data.