90. Modern Data Mobility is Challenging the Laws of Physics Presented by Hammerspace – Tech Field Day Podcast Spotlight
Modern data mobility is challenging the laws of physics; the speed of light is a fundamental limit for moving signals. This episode of the Tech Field Day podcast features Kurt Kuckein from Hammerspace discussing data movement and management with Jim Jones, Jack Poller, Andy Banta, and Alastair Cooke. The challenge is that the distributed nature of data, spread across the globe, creates significant obstacles for AI, particularly regarding the speed of light and power consumption. We delve into overcoming these limitations through technologies that facilitate data access and movement, touching on concepts such as efficient storage solutions (Open Flash Platform), the importance of centralized data management, and the agility required for evolving AI workloads. While the underlying principles of data management are not new, the scale and complexity of AI necessitate innovative approaches to ensure data can be accessed and utilized effectively, regardless of its physical location.
Transcript
Do you know where your data is? Can you make your data come together into some useful place? Is it spread all around the world?
How's your AI gonna work with things that are all around the world? Does the speed of light get in the way? What about power, cooling energy, all this and more on the Tech Field Day podcast.
Welcome to the Tech Field Day podcast, where we'll bring together a group of it technical experts to discuss a single idea about key concepts in the industry. This podcast features a variety of perspectives from members of the Tech Field Day delegate community. It's often record an association with one of our events, of course, tech Field Day, part of the FU and Group.
And this podcast is also published on our sister company Site Techstrong tv. On this special episode with, uh, hammer Space, we'll be discussing how modern data mobility is challenging the laws of physics. But before we have the discussion, it's meet who's gonna be on our massive panel today.
Hi, I'm Kurt Kine. I am the senior director of AI marketing here at Hammer Space, and I am super happy to be joining today and I'll go next. My name is Jim Jones.
I am the senior Product Architect at 1111 Systems. Jack Poller, founder and principal analyst for Paradigm Technica. And I'm Andy Banta, storage janitor.
And I'm Alistair Cook, of course, the event lead here at Tech Field Day. And it's nice to get the band back together again, because the last time the five of us were together was at AI Infrastructure Field day three. And what we've seen is that getting data in the right place is absolutely vital for building your AI applications, your AI infrastructure.
Getting business value out of that data is very much dependent on getting the right data in the right places. And yet we have this fundamental lore of physics. It's the speed of light.
You can't move anything, particularly data faster than 300,000 meters per second. Um, I'm remembering the speed of light, right? Uh, it's been a long time since I've needed to use it.
Uh, yet when we've got this data being generated all around the world by large organizations, we need lots of it in the same place, or at least access to it from the same place to build out our large applications. And I think that's one of the challenges is that the speed of light can't be fixed, but there's some technologies around that can make it less of an impact for us. And we've seen those across a few of the different technologies that we've seen for building these AI infrastructure data centers.
Um, Kurt, I kind of, uh, brought this, uh, this topic because it's, it's very much the sort of thing that you are working on with customers on a regular basis. It is indeed. Um, and, you know, figuring out how you overcome each of the variables, um, that are influencing how we're working with all of this data is definitely, um, something that we're thinking a lot about, right?
The fact that power is now the most expensive thing in the data center, um, that, you know, being able to acquire accelerated computing is really tough. Um, you know, most folks aren't even considering building out these large infrastructures on-prem because it's just far too expensive. And our data center is just far too constrained.
And, uh, then add onto that, the fact that everybody has this distributed nature these days. How do you unify everything if your own organization is out there in all sorts of regions generating data on the edge, uh, but then your compute might be located somewhere else as well in somebody's cloud. Um, and so being able to bridge that, um, and figuring out, you know, what is the representative data set so you aren't moving all of your data all the time.
Um, it's, it's a big challenge for sure. Right. And I, I think, uh, you know, you touched on the other aspect of physics that we really need to pay attention to here, where it's not just the speed of light, but it's also the density of energy necessary.
Uh, one of the things that we talked about at AI Infrastructure Field Day was, uh, was your open flash platform, which is part of the Open compute project. Uh, and shortly after the, uh, that field day event, I did attend the open compute project, uh, um, conference. And there were jokes around there that it should have been renamed the Open Cooling project because most of the things that were being displayed on the floor were actually had to do with cooling, uh, cooling of the compute pro, uh, the compute systems rather than actually the data itself.
Uh, I did have a chance to stop by the Hammer space booth there and, um, put my hands on one of your open flash platform, uh, trays, which was kind of interesting. I didn't actually have a chance to go anywhere else, or I did not see any other, other open Flash platform trays while I was there. Uh, but it's, uh, I know that you guys were working with several other companies, uh, um, was ex site and, um, I don't recall who else you were working with, but, uh, it's, uh, it'd be interesting to actually see one of those in, in action at some point.
Yep, yep. We're getting close. Um, that, uh, Trey that you saw, there was actually our initial POC, um, and we've had a ton of learnings by building that, uh, right.
Things around airflow, um, things are around serviceability. Um, and so we're, we are working to ref, refine, refine that platform, refine that design with a lot of input, um, from various flash vendors, uh, from various hardware vendors, um, as well as looking at some requirements from, um, different software vendors in addition to us. So, um, it's something that we've, um, heard a lot of interest from, and yeah, we really hope that that, um, coalesces around this design and, and, um, starts seeing some, some broader adoption.
But yeah, the whole concept behind, behind that there was that there is a more efficient way, um, to just store data. Um, and part of that is just getting out of the way of the data. Um, you know, the more layers that you put in between the data and the processing, um, is going to introduce additional latencies, uh, and additional power requirements.
And so, uh, in the case of OFP, right, our goal is to just remove the storage server completely, right? We're putting, um, the infrastructure software directly on A DPU, uh, that's low power, still high, a fairly high performance, uh, and then connecting that to the flash, so, you know, a very direct connection. Uh, and, you know, we're kind of labeling it nothing but Nick.
So, um, you've got your nick, you've got your flash, you've got a direct path from your processing. Um, and, and I think that's a big, um, part of it. Uh, that's the infrastructure part of it.
Uh, and then I think just dealing with the massive data challenge is the other big rock to move. I think one of the things that makes the Open Flash project project work, uh, with, especially with Hammer space, is that you don't actually need to have any of the data services on the plates or on the chassis themselves, on the trace themselves because you're actually taking, taking care of the data services through the Hammer space data services, Right? Yeah.
We've got a centralized, what we call our anvil server, um, and that's the thing that does all the metadata operations, um, and is outside of the data path, um, and allows us to utilize, yeah, open standard Linux, uh, and NFS, uh, on those, uh, systems to be able to handle the, the data transactions. I think Andy, one, one of the things is that we love playing with hardware and, you know, hardware cool. And we're all geeks, and that's all good, but thinking about sort of the problem at a higher level, one of the things we've learned over the course of the, the AI infrastructure field days, not field day three, field day two field day one, is that the, arguably the most expensive component other than the power plant itself is the GPUs.
And the big problem is the GPUs right now, while being hard to obtain, they're also sitting idle 70% of the time, right? And that's because they can't get the data. I think we focus on we being not only us here in the podcast, but we as an industry like to focus on the hardware and throwing hardware solutions at it.
And what if we tweak this component, make this component faster, and we've heard from how do you make the ethernet faster and how do you make the switches faster, and how do you make this, and how do you make that faster? But the architectural problem is how do I have terabytes or petabytes or multi petabytes of data? How do I get that from where it lives into the training environment, and then get it, then feed that into the G-P-U-D-P-U environment.
And I think that's the, the sort of the physics problem that we're sort of wandering around, talking around, but not talking about is we have, how do you get a petabyte of data into a processing environment that can consume that almost instantly? And then what do you do from there? I'd say it, it, it's interesting from a couple, couple different points because you're gathering that data, you know, probably not where you're processing it initially, but from a application, you know, we're gonna spend millions of dollars to create a model, or we're gonna spend, you know, probably millions of dollars, even if we're not doing, you know, quote unquote AI creating a model, et cetera, et cetera, et cetera.
We're just gathering it. It costs lots of monies to develop those applications regardless of what they are. And every time we have to change those to allow for data migration, data movement, whatever it is, or, or data expansion, we have to, you know, that incurs costs that include incurs downtime or, you know, maybe from the, you know, PHY physics of the, how fast can I get this capability out to the market or to my consumer?
It slows things down. And that's where I was really impressed with what Hammer space was doing with the unified namespace, was it lets me hide any number of sins, if you will, underneath the covers. You know, I may, I may have a NAS floating over here, or I may have an object storage in 17 different platforms, but as far as my developers, as far as my consumers are concerned, that's one, that's one endpoint, that's one set of credentials and things just work.
And that's, you know, where that gets important. Um, you know, in my day job, I'm dealing with customers that are wanting to move from point A to point B, um, to chase cost the physical as the physics, as in the, how much does it cost physical, uh, side of things every day. And it's incredibly painful to do that, um, one because it costs money to move the data, but more importantly, your application has to be able to support that kind of thing.
Um, and so I, I really love that kind of capability to, to further us as technologists to be able to not have to worry about, well, where is my data that may well be at the edge where I've gathered it or maybe my data center or maybe somewhere else and not have to care about it. Well, I think that's, uh, fundamentally been a big part of what's been missing from the conversation, uh, in ai, right? Is, um, mainly because the industry has been focused on these really, really large developments, right?
This initial training of these massive LLMs and, you know, everybody's been focused on just that interface between the storage and the processors, and how do I get enough data into the processors, um, from the storage system. So peak storage performance, right, has been a lot of the conversation, um, and, you know, has served that side of AI very, very well. Um, you know, and we see parallel file systems and really specialized systems, um, being developed for that.
Um, but for the other 98% of people who want to do something with ai, we haven't talked about, okay, how do I get my arms around my organization's data? What do I have? And then once I realized that it's yeah, out there on 50 different silos, how do I then identify our representative data set and make sure that's what we're moving rather than, oh, the way I solve my silo problem is by implementing a new silo and migrating all my data into that new silo, which is gonna take forever, is not going to really help you identify for on an ongoing basis, oh, these are the key pieces of information I need to ingest into my representative data set, and then process against that, whether that's on premises computing or somewhere in the cloud, right?
We need to reduce the data problem in the first place to be able to take advantage of all these things. That's, that's one of the, the other things that was talked about at that AI field day was the, the tier zero concept that you, uh, you've put together where you take advantage, the ME drives actually on the GPU servers and use that as, uh, a very local tier where you can migrate the data into and have, have very fast access to it, uh, and just have it presented out of each one of the GPU servers as an NFS store that's shared, shared and consumed by Hammer space. And, you know, it's, um, what, uh, you know, Molly kind of early on and talked, talked about, um, how the name Hammer space came up, which is, uh, the, the space outside of frame and the cartoons where the, where people reach off and they, uh, you know, find something that's off frame and, and make use of it.
One of the things that I really liked about the, this most recent presentation by Hammer space is that you kind of did talk about, uh, where you're breaching off frame to grab the various different things where you, you talked about the Tier zero and whatever, uh, interesting concept. And it also was a, um, it, it was also an interesting way of talking about how to get the speed, the, getting back to the speed of flight issue that we were talking about, had to get the speed you need to feed the GPUs. The Jack was meshing where, you know, GPUs are always hungry, you always need to feed them.
Yeah, I think, um, you know, there are a few, um, key things that we're addressing with Tier zero. Um, and it's not just right providing high performance, um, access to the, the local GPUs with that and VME, but, um, there's some other market realities that are developing, uh, that make this very, very interesting, right? If you're, um, actually purchasing right on premises, GPU computing, you've got these NVME devices that are coming with it and are inside of those servers, um, and we're seeing or beginning to see a really, really constrained flash market where, you know, the hyperscalers and really large, um, organizations have really cornered the market on available flash.
Um, and so if you are struggling to find that flash within your environment to be able to create a high performance tier, well, hey, we can help you reclaim that flash that you already have on premises in those GPU servers. Um, another, um, advantage is, um, you know, some of the hyperscalers have really, um, adopted high performance networks for their AI computing. Others are still using, you know, standard enterprise networking, um, with, you know, abstraction layers between the storage, um, and that AI computing and what we're able to do with folks like, um, Oracle and, um, Microsoft is actually place the data within the servers.
So you as a customer to those clouds know that you're getting the maximum performance, um, out of those systems rather than having your, you know, data sitting in a blob somewhere and hoping it's getting into that GPU system fast enough. Yeah. Growing on my VM world or my VMware background, it's kind of like what vsan was 15, 20 years ago, or 10, 15 years ago, uh, where you're making use of local storage because it's there.
Yep. And we don't love the virtualization word. We aren't storage virtualization.
That's not what we're doing. Um, but VMware analogy is very, very close. The, the, the funny thing is, you know, the, the trite saying is actually true, which is everything will this new again, right?
And it's, this is not new stuff. As Andy mentioned, we've been reclaiming local storage for a long time. Uh, databases used to have database administrators who sold job.
It was to, uh, place your data in the most advantageous spot for speed of access or speed of update. And, you know, local tiering and caching was something that, uh, 40 years ago, uh, I visited a Ford plant where they used A CDC cyber nine 60 supercomputer as the local cache for a cray supercomputer, because that was the only thing that could feed the cray quickly enough, right? And, you know, and a lot of what we're talking about here is that management of data, both the physical placement of the data as well as the management of what data do you need and when do you need it in order to feed the beast.
Yeah, I think that the whole systems thinking of this is not a, a set of isolated decisions, isolated components, but it is actually a, a system that interacts together and that you've gotta get coherent across multiple pieces of infrastructure, whether it's the networking, the storage, the compute layers, but also across those multiple locations. And this is one of the challenges we've had as enterprise organizations have grown progressively, right? Back when Jack was working with these credit supercomputers, you, you couldn't have 20 or 30 of them spread around the world.
Now we have a 20 or 30 data centers spread around the world, each of which is generating vast amounts of data that we're trying to get a coherent view of and get coherent value outta. I think one of the really vital things to think about that Kurt touched on was that generating foundation models, taking that common internet for ingesting that and training a foundation model is completely different to an organization that wants to fine tune a model or to use a, a rag solution to do inference, the actual workload, the data flows, the messy enterprise world where we, we have one of everything, sometimes three or four of everything, uh, and we need to somehow get the intelligence out of all of these things. And as Kurt says, we don't wanna just pour all of that into a new silo and pay again for storing all of that data in the new silo that's gonna, uh, supposedly unify and show us everything.
Of course, we know that by the time we built that silo, we've got 15 other sources of data that maybe our silo can't accommodate, and now we've gotta build another silo of silos. So I, I think there's some very different sets of challenges for enterprise organizations, but as Jack says, these are not entirely new things. These are things we have seen over years and years.
Well, an AI is just the current variant of the conversation of the, this is the application that's driving this need. You know, it's as, as you were speaking, Kurt, the thought that came to my head was, we, you know, one of the things that you're trying to do is you're trying to loosely couple the loosely coupled things. You know, we've been talking about making our more, our applications less, you know, aware of what's going on in the infrastructure for literally decades at this point.
And as we abstract out of that, all of that is just trying to make it to where, okay, so today it's ai, the next conversation may be quantum computing or the conversation past that may be, you know, you know, something from the Jetsons as far as I know. Uh, and, and you know, the, anytime that we can abstract, you know, the, the interface from the infrastructure, we're just making our life much easier as we iterate from this thing to the next. For sure.
And I think flexibility and agility are two big driving, um, characteristics behind Hammer space. Um, because right, it's not always about bringing the data to the compute either, right? Some, sometimes it's gonna be about bringing the compute to the data, right?
So our friends at Nvidia continue to churn out different solutions that, um, you know, go all over the place. You've got 'em in the robots themselves, you've got 'em in all of the sensors, you've got 'em in region, um, you've got 'em in your data center, and then you've got 'em in the cloud, right? And it's horses for courses.
And so, um, I think even getting back to Jack's point, right? Each of these steps has been done before, right? Um, we've done this in cycles over and over again.
Um, what we're really looking to do is, yes, we're solving it for the unstructured data challenge, right? Which is, has its own unique headaches, um, where prior, right? With, you know, maybe the structured, it's been much easier.
It's been a, a smaller challenge. Um, but now we're dealing with disparate types of data. We're, you know, the size of the data generally, um, is, is much, much bigger on an individual file level.
And, um, the, the demands are constantly changing. And how we process that data, um, is constantly changing. And so, you know, introducing agility and flexibility is key so that you aren't locked into a single architecture where it's monolithic in the data center or even monolithic in the cloud.
Um, you need to be able to tune everything and do it without doing it manually. Yeah. This, this topic actually came up at open compute, uh, OCP and at Super Compute where the idea that, you know, what is currently being talked about is AI was simply HPC or High Performance Computing three years ago, and that was the exact space that Hammer space was playing in three years ago as well.
And the, the use cases are the, the infrastructure needed is incredibly similar between HBC and ai. Uh, with the, the difference being the HPC actually strives to come up with the correct answers, and AI just makes stuff up. I think the, the four letter word you're looking for, Annie, is Hadoop, which was all about moving compute to the storage rather than storage to compute.
But that's a four letter word, and we won't go there. But it, it's, um, you know, I, your Hammer space goes in trying to address a problem that is needed in these types of compute environments, and you've, you, um, you have a variety of innovative ways to actually accomplish that, For sure. And you know, the, the other important point being there is, you know, we are not looking to develop our own walled garden either.
Um, we want to be able to give our customers choice when it comes to the different things that they do. So, um, you know, things as subtle as you know, are how are you embedding a Vector database into your solution? Um, we see some people saying, we'll be the one appliance for everything, um, that you need to do for ai, and you don't have a choice, right?
Um, that's one perspective. And yeah, in many ways that makes some of the decision and implementation simpler. Um, at the same time, um, it locks you into their vision and their silo.
Um, whereas we're looking to give you that flexibility and choice when it comes to decisions like that, as well as the decisions behind, you know, the actual infrastructure that you're using and the location that you're using it. Um, and having a solution that sits in the middle that then automates the administration of all of that data movement is where we see our key value. And we're gonna have, uh, plenty more conversations around some of that key value and the problems we're trying to solve for people through ai, both through further episodes of the Tech Field Day podcast, but also at future AI infrastructure, ai, uh, field day events as well.
There. This is definitely a place that we like to have our conversations, as you can probably tell, but we don't have time for all of those conversations today. So, before we close out for today, where can people connect with you to carry on this conversation?
Maybe learn a little bit more about the things you, you've been thinking about, the things that are important to you? So I'm always available, uh, on LinkedIn. Um, you, you know, my name's Kurt Kine, K-U-C-K-E-I-N is the last name.
Um, so you can always connect with me there, um, as well as I'm, you know, always publishing stuff, uh, on our website, blogs, things like that. Do feel free to reach out through that as well. Yeah.
And I can be found on the internet at Coolaid Info or pick your social media platform up to and including LinkedIn. And I'm Kool-Aid it there. com, our corporate, uh, site as well as on LinkedIn Security Boulevard and other social media sites.
And you can find me on LinkedIn, uh, at Andy Banta, uh, on Blue Sky, and on my, from my own content. You can find that at andy banta blue uh, substack com. And of course, I'm Alice Cook, and you can find me all over the tech Field day, as well as on social media, uh, LinkedIn.
Uh, you may also find me as demi tess nz nz since I live here in New Zealand. So thank you for listening to this episode of the Tech Field Day podcast. If you've enjoyed our conversation, our discussion with our delegates, please subscribe on YouTube or your favorite podcast application.
So don't miss a single episode. As always, give us a rating, nice review. Tell us how wonderful we're we.
com/podcast of us on Techstrong tv, maybe even on your smart device. Uh, thanks for listening, and we.