UT07x04: Maximum Performance and Efficiency in AI Data Infrastructure with Xinnor – Utilizing Tech
Cutting-edge AI infrastructure needs all the performance it can get, but these environments must also be efficient and reliable. This episode of Utilizing Tech, brought to you by Solidigm, features Davide Villa of Xinnor discussing the value of modern software RAID and NVMe SSDs with Ace Stryker and Stephen Foskett. Xinnor xiRAID leverages the resources of the server, including the AVX instruction set found on modern CPUs, to combine NVMe SSDs, providing high performance and reliability inside the box. Modern servers have multiple internal drive slots, and all of these drives must be managed and protected in the event of failure. This is especially important in AI servers, since an ML training run can take weeks, amplifying the risk of failure. Software RAID can be used in many different implementations, with various file systems, including NFS and high-performance networks like InfiniBand. And it can be tuned to maximize performance for each workload. Xinnor can help customers to tune the software to maximize reliability of SSDs, especially with QLC flash, by adapting the chunk size and minimizing write amplification. Xinnor also produces a storage platform solution called xiSTORE that combines xiRAID with the Lustre FS clustered file system, which is already popular in HPC environments. Although many environments can benefit from a full-featured storage platform, others need a software RAID solution to combine NVMe SSDs for performance and reliability.
Transcript
Cutting edge AI infrastructure needs all the performance it can get, but these environments must also be efficient and reliable. This episode of Utilizing Tech, brought to you by soddy features David Villa from xor, discussing the value of modern software raid on MVME storage, SSDs with a Stryker and myself. Welcome to Utilizing Tech, the podcast about emerging technology from Tech Field Day, part of the Futurum Group this season, brought to you by soy focuses on the question of AI data infrastructure, all of the ways that we have to support our AI training and inferencing workloads.
I'm your host Steven Foskett, organizer of the Tech Field Day event series, and joining me today as co-host is ACE Stryker from Solid I, welcome to the show once again. Hey, Steven. Thank you.
A pleasure to be with you again. So, ACE, we have been talking about various aspects of the AI data infrastructure stack, uh, on this show today. We're gonna go a little bit nerdy, a little bit deep, uh, on the, the topic of storage and raid.
I know that most of the companies that are deploying, uh, AI infrastructure, especially for training, the last thing they wanna do is invest a ton of money and, and effort and precious, precious PCIE space and, uh, data center space in a big, fancy storage, uh, system. A lot of these companies are trying to create a, a system that, that that does all the things that need right there in the chassis. Yeah, a lot of, uh, uh, folks, you know, looking in at, at AI data infrastructure from the outside may not appreciate, you know, the, the challenge this that come with sort of coordinating all the, all the pieces of the system, right?
Uh, it's easy enough to say, oh, you know, you can buy more drives to add capacity, or you can, you know, uh, pull these levers to increase your performance if you need. But it turns out that, you know, sort of optimizing the way these pieces play together is not, uh, easily done, right? And there's, uh, there's a lot of interesting innovation happening in that space.
Um, in particular, we're seeing, uh, you know, a lot of, um, this, these kind of coordination efforts, whether it's in, uh, networking or storage or other parts of the system. Uh, we're kind of entering a world where a lot of that stuff is being done by software, right? Uh, where, where historically, you know, we had these sort of purpose-built, uh, pieces of hardware that were responsible for that kind of work.
And, uh, and in the new world, we're transitioning to, um, software defined solutions for a lot of this stuff that, uh, it's really exciting to see because it can, what, what, what you can get out of that in, in a lot of cases is not just, um, uh, more performance, but you can also do that, uh, uh, in, in a more efficient architecture. Oftentimes you're saving power space at the same time. And so, uh, it's definitely an area to watch going forward.
Yeah, and it, it seems as a storage nerd myself, um, it always makes me sad when people underestimate, uh, what they need in terms of storage, uh, solutions. Uh, they maybe will try to deploy on just bare drives, or they'll try to use sort of out of the box software that isn't really up to the task from a performance. Uh, and, and even from a reliability perspective or, uh, they'll deploy something that's just way overly complicated and huge.
Um, so refuting all of this, uh, we have, uh, xor, this is a company that makes a software raid solution essentially. It, it lets your server manage storage in a way, uh, in internally, in a way that an external storage array might do. So we're excited to have David Villa joining us today to talk a little bit about the, the world of software raid and the world of the practical ways that companies are, are managing storage.
Welcome to the show. Thank you Steven, uh, for inviting me. I'm really excited to be part of the show.
I'm, uh, David Villa. I'm the Chief Revenue Officer at, uh, Zinner, uh, the leading company in software raid for M-V-M-E-S-S-D. So tell us a little bit about yourself and about xor.
Uh, Zinner is, uh, is a startup, uh, in, uh, software raid, as you mentioned. Uh, we were founded, uh, a couple of years ago, but, uh, we, in reality, we inherited, uh, the work that has been done in the last 10 years by, uh, the previous company, uh, that was sold and created by our founder. Uh, so now we, uh, we're a young company, but, uh, uh, we are leveraging more than 10 years of, uh, development in, uh, op optimizing data path to provide very fast, uh, storage.
Uh, we're about 45 people, uh, disperse around the globe and very much an r and d company. And just to be clear, uh, when you talk about, uh, optimizing the data path and software rate and everything, you're talking about building basically enterprise grade reliability and performance on, you know, within the server without having to have a bunch of expensive add-on cards or a separate chassis or anything. You're talking about basically building a server with a bunch of NVME drives and then using the power of the CPU, uh, to provide an incredible amount of performance and reliability, right?
Yeah. There are, uh, enough resources within the server that we don't need to add any accelerator or any other component that might become a single point of failure at some point. Uh, so what we do, we combine a VX technology available on all modern, uh, CPU, and we combine it with our, uh, special data part, we call it lockless, uh, data part.
And, uh, what's unique in our data part is the way we distribute the load across all the available cores on the CPU, uh, by minimizing, uh, spike of load on a single core. And by doing that, we, uh, we avoid spike and we, we can, uh, get stable performance, not just in normal operation, but also in, uh, degraded mode. One of the things that we have, uh, set out to explore on this podcast is, uh, you know, within the, the context of AI specifically, um, how this boom is drawing a lot of these, um, technical challenges and opportunities for solutions into, uh, uh, sharper focus, right?
So can you talk a little bit about what impact the acceleration, uh, in, in the AI world has had on the problems that that xor set out to solve? Yeah, that's, um, that's a very hot topic. Uh, today's, uh, as, uh, we, we see that, uh, our main market is definitely becoming, providing very fast storage for, uh, ai uh, workloads.
So what, what, uh, we experienced by working with, uh, with our customer is that, uh, uh, traditional HPC player at the university, the research institutes, they're all, uh, uh, now, uh, facing, uh, some level of, uh, ai, uh, workload. So they, they're all, uh, moving. Uh, they all keep themselves, uh, with some, uh, GPU very powerful gpu, uh, that, uh, require a different type of storage than, uh, what they traditionally used to deal with.
Traditionally in the HPC space, uh, HTD uh, rotating spindles drive were good enough for many use cases when it comes to AI workload. Uh, they are not sufficient. Their performance is not sufficient any longer, uh, because, uh, of the very high, uh, read and write, uh, requirements, uh, that those, uh, uh, modern GPU, uh, have and, uh, those modern GPU, they are, uh, expensive systems.
So they, the customer cannot afford to keep them busy, uh, to keep them waiting for data. So it's absolutely critical that, uh, uh, the storage that is selected to provide data for, uh, AI models, uh, is capable of delivering a stable, uh, high performance in the tens of gigabyte per second. That's certainly something we hear a lot about in our conversations with folks, uh, in the industry, is the, uh, sort of primary importance of GPU utilization, right?
Nobody wants to spend tens of thousands of dollars per unit and, and, uh, in some cases, uh, even more than that, uh, you know, to run something at 60% utilization, right? And so feeding the data to the GPU in, in, in something like the training stage of the, the data pipeline becomes really important to make sure that you're getting, um, getting the bang for your buck on the, on the compute side. Right?
C can you talk a little bit about, um, you know, what I see if I open up a box that has xor running in it, if I, if I, uh, take a conventional architecture, you know, I'm probably used to seeing an array of NVME drives, and then there's a, there's a raid card in there that's doing a job. Uh, can you talk a little bit about how your, uh, solution is, is different? Yeah.
Fir first of all, our solution is software only. Uh, so we use the system, we leverage the system resources, and when I say the system resources, I'm referring just to the CPU, uh, because we don't have a cache in our, uh, uh, rate implementation. So we don't need, uh, uh, memory allocation.
Uh, that's the, the primary difference. Uh, but the reason why we came up, uh, with, uh, our own, uh, software rate implementation is because traditional hardware rate architecture cannot keep up with the level of parallelism of, uh, new MVME drives. Uh, so the level of parallelism that you can get on PCE, gen three, gen four, and even more on Gen five is such, uh, that, uh, you need a powerful CPU, uh, to, uh, be able to, to run the checks sum calculation.
Uh, then the other limitation that you face, uh, with, uh, hardware rate is the number of PCIE lanes, uh, hardware rate connected through the PCIE, uh, line, uh, can only have 16 lanes. And, uh, each MVME has four lanes on its own, meaning that, uh, you are saturating the PCIE bus with just four MVME drives. And for, uh, ai, uh, models, uh, and workloads for MVME are not sufficient.
So we have customer deploying cluster of, uh, multiple tens of server with 24 MVME, uh, per server. So we believe that, uh, for, uh, um, uh, MVME uh, drives, but, uh, and, uh, uh, for AI workload, there is only one way to go, which is a software rate. Well, it's true because you look at these servers and, um, you know, you talk about four NVME drives, most of these servers have a lot more than four NVME drives.
Uh, most of these servers have a pile of them. And even though those drives are pretty big, and each drive provides a lot of performance, you still don't wanna manage those individually. I, I don't know about, uh, our listeners, but I'm an old school Unix systems administrator, and I don't want to be dealing with, uh, 20, uh, individual drives.
I wanna be dealing with a, uh, a combined drive. And not only that, these drives are incredibly reliable, um, very, very reliable, but nothing is guaranteed, especially when it comes to things like the mechanical components of the drives and, you know, insertion and removal and things like that. It is possible for drives to fail, um, you need to have, uh, reliability as well, and predictability of performance.
That's another thing I I think that, that, that occurs to me too, is that, you know, if you, if you lose a drive and there's a rebuild or something like that, you don't want to lose all the work that you've done, uh, so far in terms of, uh, training workloads and so on. So all of this points to the need, I think, for a system that manages storage. Now, I mean, RAID isn't really storage management, but it is drive management, and it does definitely help in configuring these systems, right?
Yeah, you're, you're absolutely right. Um, so when, when you run AI models, uh, those models can take several weeks, if not, uh, multiple months, uh, to run and, uh, while running the model, something can, can go bad. So you need to provision, uh, and make sure that you're able to deal with potential failures or drive this connection from, from the array without, uh, having, without losing data for sure, but also without being impacted in the performance.
And that, that's where we, we, we step in by providing, uh, rate capability, uh, so by providing, uh, data integrity, but, and making sure that, uh, we keep, uh, very high performance even in, uh, degraded mode. I'm curious, as you talk to, uh, your customers who are engaged in AI work, what's, what's the sense you're getting of, um, how those customers are viewing their storage needs? And do you see any trends there?
You know, do you hear from folks about, Hey, we really need to get more sequential throughput out of our storage subsystem, or, Hey, you know, random, random performance is really important for us, or, or capacity needs continue to grow and grow. Um, you know, what, what do you see as kind of like trends in, in, uh, the way folks are viewing and, and their requirements that they are, are, uh, demanding from their storage subsystems in, in AI clusters? That's a $1 million question, I would say.
So we, we have been asking this question to many different customers, and we kind of get, uh, a very similar answer from most of them. And the answer is that they don't know they need to provision for, uh, uh, the extreme cases because, uh, the workload is different. Uh, it is not always, uh, the same workload.
Uh, so if we want to oversimplify, we can say that, uh, uh, AI workload, uh, it's mostly, uh, sequential by nature, uh, and a combination of, uh, read during the ingestion and, uh, right during the checkpoints. Uh, but not all the AI models, all the AI training are equal. Uh, so there are, uh, many distinction that needs to be, uh, to be made.
And, uh, we, we see that, uh, random performance plays a role as well, uh, what we experience with our customers. As I say that most of those customers, they used to be running HPC, uh, infrastructure, and, uh, they very much, uh, would like to stay with what they know. So they would like to keep on using the popular parallel file system that used to used on their, uh, uh, HPC implementation and be able to, to leverage it, uh, to leverage their competence in using those, uh, parallel file systems or file systems also to run AI models.
So, as, as a matter of fact, uh, every customer has a different type of, uh, implementing storage. So we have, we're working with many different university. We have universities implementing, uh, our zero aid, uh, with all flesh implementation based on last.
We have other university who prefer going down the route of, uh, uh, open source, uh, using, uh, BGFS. And we also work, uh, we, we did, uh, deployments with universities that they don't want complexity, they just want a simple file system like NFS and be able to, to saturate, uh, the, uh, the bandwidth. There are network bandwidth with, uh, in this specific case I'm referring to, and InfiniBand, uh, deployment, which we, uh, recently did, uh, at a major university in Germany, uh, to provide fast storage through the network, uh, to, uh, DGX to, to DGX systems.
Uh, so to answer your question, uh, it's, it's tough to give you a simple answer, uh, because, uh, we're still, uh, in the early days of AI adoption and, uh, uh, everybody's still, uh, in a learning phase. What, what is clear is that performance really matters. And, and in order to get that performance, I imagine that there may be some tuning that you might have to do as well.
So if you're the raid level, uh, the raid layer, um, you, I imagine that there might be some slight different configuration, uh, well, seemingly slight configuration that can make a huge difference on performance. Again, based on my background in the storage space, I know that, that things like, you know, block sizes and so on, can make a huge difference in performance. I assume that you guys can adapt to the needs of the higher level software, right?
Yeah, absolutely. So our, our, our software gives the system admin the flexibility to select the, the right level of geometry and the right, uh, uh, chunk size, uh, the, uh, the, the minimal amount of data that is written, uh, uh, within the rate to a single drive in the, uh, most optimal way, depending on the workload that, uh, will be run. Um, so we, we actually did a lot of activities with our partner solid D to find the optimal configuration base, uh, on the specific, uh, uh, workload that we were running, uh, in our eight five implementation.
And, uh, with, uh, uh, proper alignment to the SSD, uh, indirect, um, we see that customer, they require more and more storage, uh, q and most of the workload is sequential. Uh, so this makes, uh, uh, QLCA an a very viable, uh, technology for a AI workload. Uh, but everybody knows that, uh, QLC uh, comes with some limitation, like, uh, limited number of program array cycles.
So with our software, by selecting the proper, uh, chunk size, we're able to minimize the right amplification into the SSD, and by doing that, we can enable using QLC for, uh, extensive, uh, AI uh, projects. So what, what I'm saying, uh, is actually, uh, going to be part of, uh, uh, a joint paper, uh, with solid dme and, uh, uh, you will be able to, to see the outcome of this research. Yeah, absolutely.
com in our Insights hub. We've got some great testing there, uh, with some of our, uh, QLC drives on, on the CA, uh, solution. Um, you mentioned right applica, right, right, right.
Amplification a moment ago. Um, do you mind, for folks who are less, uh, immersed in the storage world than ourselves, maybe just, you know, give us a, a one minute version of what that is and, and why It's, um, uh, uh, uh, a, a challenge and, and kind of how Z rate addresses it differently In very simple terms, uh, right. Amplification, uh, is, uh, uh, that terms that refer to the fact that, uh, when the host is writing one data, uh, to the SSD internally, there are, there is more than one, right, that happens to the physical component, to the physical man, uh, component.
And, uh, given the fact that, uh, all the SSD have limited number of program cycles when it comes to QLC, they have fewer program rate cycles than TLC. It's, uh, very important that, uh, uh, we implement, uh, uh, algorithm to minimize these numbers. So to, to keep this number as close as possible to one.
And, uh, uh, with our software, we can do that because we can, uh, change, uh, the chunk size, so the minimum amount of data that will be written, uh, to the SSD to each SSD that are part of the rate, uh, radar array. And by doing that, uh, we can minimize the rate modify, right, uh, that, uh, needs to be done on the SSD itself. Uh, so when, when we calculate the check sums, uh, if we are not aligned with the indirect of the SSD, we might risk to write, uh, multiple times, uh, data to the SSD, uh, with our software, we, we can, uh, uh, find the proper tuning, uh, based on the workload, based on the, uh, number of drives that are part of the rate array by the level of raid.
And, uh, uh, we, we are, we're able to find the optimal configuration to keep this number as close to one as possible. I could imagine that a lot of this might sound a little concerning to someone listening and, and trying to deploy this. They might think, oh, boy, that's a lot of, uh, a lot of tuning, a lot of under the hood stuff that I don't really understand.
Um, do you have, uh, best practices for various devices? I mean, do you help, uh, customers to come up with the right configuration? Yeah, that's part of our job.
So normally when, uh, we engage with the customer, uh, we spend quite some time with our pre-sales team to, uh, understand the workload of the customer and find the optimal configuration. Then once the optimal configuration is identified, there's no work to be done anymore by the system admin. Uh, it's ready to fly, uh, and there's no additional tuning to be done.
Uh, one of the other things I wanted to ask you about, uh, we've talked about Z Raid, uh, which, uh, is an incredible, uh, solution in terms of, uh, uh, sort of the benefits to efficiency and performance at the same time. Very exciting, uh, what you're working on over there. Uh, another thing I've heard about, uh, more recently is, is, uh, uh, I think another product of yours called Z Store.
Uh, could you, could you tell us a little bit about that and, and, and how these pieces work together? Uh, so our, our core competence, as I said, is in the data path. And, and it's in the way we, uh, create a very efficient array.
Uh, and we see that for some industries, uh, array, a standalone raid, uh, at least for some customer, is not, uh, sufficient. So they're looking for, uh, um, a broader solution. So this story is, uh, one of the first, uh, of those solution that we're bringing to the market.
And, uh, it is at the, it's based on our zero aid implementation, uh, for, uh, M-V-M-E-S-S-D, but we also, uh, combine it with dec cluster rate, uh, to handle, uh, the, uh, problem typical problematics of, uh, hardest drive, which is extremely long, uh, rebuild time. Uh, so through our, uh, own implementation of Decluster rate, we can drastically reduce, uh, the, uh, rebuild time of artist drive. Then on, on top of the rate within Z Store, you will find high availability implementation.
Uh, so there is no single point of failure. You can lose a server and still, uh, uh, you get, uh, all the rate up and running. We have our, uh, control plane to manage virtual machine, and on top of those virtual machine, we mount the last, uh, parallel file system.
So it's, it's a complete, uh, end-to-end, uh, solution, uh, that HPC and AI customer can deploy, uh, without needing to combine, uh, zero eight standalone with third party softwares. Uh, that's the first of, uh, a series of, uh, solution that, uh, we will bring to the market without, uh, ever leaving our core competence, which is, uh, very much the rate implementation. Very good.
Well, it certainly seems like the, uh, the market is responding, uh, to your approach here. Uh, it sounds like, uh, the future's very bright cino. I wish the same to ourselves, and, uh, I can, I can say that, uh, that's the situation and what, what we're experiencing, uh, the trend that towards, uh, ai, uh, deployments across all the industries.
So it's not just one of two guys that are deploying AI models, but, uh, it's becoming pervasive in the industry. It, it definitely boost, uh, the requirements for very fast and reliable storage. I think the, the interesting thing here too is that all the things that we've been talking about are gonna be very familiar and comfortable for people that are deploying these systems.
So you mentioned, for example, luster as part of the Z store architecture. Well, a lot of HPC environments are already using luster and are happy with it. Uh, you know, we talked about as well, the how, if you combine multiple NVME drives into a single, uh, zade, um, system, well, that's gonna be familiar for people who don't really know a lot about storage, because they're gonna see the storage as just a big space, a big amount of space that I can use and, and, and let zar manage that.
Um, similarly, the, the entire idea of software raid, uh, it's one of those things I think where, um, people may, uh, they kind of probably fall into two camps, either. On the one hand, they think storage is just storage, and I threw some drives in and why doesn't it work? Or they think storage is a big task and I have to go buy a big thing and, and, and do a big thing.
And this kind of falls comfortably in the middle where they can get those features, but they don't have to have a, a huge, uh, investment in storage. I can see that there are probably times when people might want, uh, a big storage infrastructure storage platform as well. But, um, for many people that are deploying, uh, especially, uh, you know, ML training, uh, they may want something that's a lot leaner and, and yet still provides the kind of reliability that you're, that you're talking about.
So this makes a lot of sense. Uh, thank you so much for, for talking a little bit about, uh, the, uh, the lower level of AI data infrastructure, lower level in the stack, not, not in terms of I importance. Um, where can we, uh, continue learning more about SI bet you guys are, are doing some things and are gonna be at some industry events.
Yeah. Uh, you can start by having a look at our website at, uh, www zinner io, and then we are going to exhibit at the future of memory of, uh, and storage, uh, which will happen in August, uh, in the Bay Area. So we look forward to see you there and, uh, asking many questions.
Ace, uh, I think that, uh, you and Xena are also working on, uh, a paper together, right? Yeah. That's available, uh, on our website.
So we've got some, uh, really compelling, uh, results from the lab where we talk about, uh, what we saw putting a bunch of, uh, soy high capacity QLC SSDs, uh, into an array using, uh, ZA, uh, and, and I have to say, you know, I didn't, I didn't do the testing that was our, our solution, uh, architectured team. But, um, but reading through the results, uh, or, or very exciting to me, um, you know, anything that can be done, uh, a to to improve performance, right? And, and move through the, uh, the sort of AI model development workflow faster, that's a big win.
But also to do so while freeing up, you know, A-P-C-I-E slot and saving, you know, power that you would've otherwise spent on a dedicated card, for example, to do work like this, uh, is a big deal. We keep hearing more and more by the week about these really scary, uh, projections of how much space and power, you know, AI data centers are gonna consume, uh, in the near future. And so it's a focus area, uh, at soy certainly to figure out how can we reduce the, uh, the environmental impact there, make these things more efficient, and a solution like ZR fits right into that, right?
Doing more with less is absolutely, uh, the path forward here to, to enable, you know, AI development to continue at its sort of breakneck pace of, of, uh, advancement that we're currently seeing. Well, thank you so much, uh, both of you for joining us today for this episode. Uh, a again storage nerd here.
I'm glad to be able to nerd out a little bit about storage, while also maybe reassuring folks that they don't have to be storage nerds in order to have, uh, reliable and high performance storage in the software domain. Thank you for listening to this episode of, uh, utilizing tech, uh, part of the utilizing AI data infrastructure series. You can find this podcast in your favorite podcast applications.
Just look for utilizing tech. Uh, you'll also find us on YouTube if you prefer to watch a video version. If you enjoyed this discussion, please do give us a rating, give us a review, give us a comment.
We'd love to hear from you. This podcast was brought to you by Tech Field Day, home of IT experts from across the enterprise. Now part of the futurum Group.
It was also sponsored this episode as, and this season was sponsored by soy. com, or find us on X, Twitter and Mastodon at utilizing Tech. Thanks for listening, and we will see you next week.