NetApp Insight 2025: Keith Aasen Unveils the AI-Ready NetApp AFX — The Future of Scalable Storage
At NetApp Insight 2025, Stephen Foskett sits down with Keith Aasen, Product Manager at NetApp, to discuss the debut of the company’s most powerful storage system yet—the NetApp AFX. Designed from the ground up for massive scalability, high throughput, and AI-driven workloads, AFX combines the familiar ONTAP data management layer with an entirely reengineered backend architecture. Supporting up to 128 nodes, AFX delivers terabytes per second of throughput and exabyte-scale capacity, with automated workload balancing, zero-gravity migrations, and a fully Ethernet-based switched fabric powered by Cisco. With support for NFS v4 with pNFS and S3, AFX provides parallel file system–level performance without the complexity, making it ideal for modern AI training and inference environments. As Foskett and Aasen explain, AFX forms the foundation for NetApp’s AI Data Engine, bringing intelligent metadata, scalability, and automation to the next generation of enterprise data management.
Transcript
I am Steven Foskett here at NetApp Insight, and we are pretty excited that we just saw a new generation of NetApp storage systems announced today. The A FX is NetApp's latest, greatest, all singing, all dancing, really honking, big storage array. And I am joined here by Keith Ason, who is the product manager for the A FX.
Welcome Keith. Thanks for having me. I'm pretty excited to be here.
Uh, you've been working on this thing for a little while, right? Uh, you finally excited to see it on stage, couple Years now, and it's been, it's really weird to see it out on stage. It, it felt like the sending your kid on the first day of school right.
Is now he's out there and hopefully everybody likes him. Well, I gotta say, uh, right off the bat, I'm a longtime NetApp customer. I was actually one of NetApp's first customers.
No kidding. Um, back in 1997. Wow.
And, um, one of the things NetApp has always done is beautiful bezels. Uh, you got one man. Yeah, I looked pretty nice in that rack.
I, I, I have to say it was, uh, a little bit of work there and yeah, it's, it doesn't hurt to to look pretty in the storage world. Do you have one on your shelf? Not yet.
Oh, none. You gotta steal one from manufacturing. Oh, I do.
I do. Maybe I'll go home from the show with one in the suitcase. So tell me a little bit more.
So, the A FX, as I said, is basically the biggest baddest storage system NetApp has ever built. Uh, absolutely without a doubt. You know, uh, or order of magnitude bigger from a throughput standpoint and a capacity standpoint.
Uh, and it, and it's ready for whatever AI is gonna throw, uh, at organizations And whatever AI is gonna throw at organizations, as we heard, of course, is data and lots of it. And those of us who've been involved in AI in the enterprise have seen that one of the challenges is, um, random access, multi-client access and throughput. And it seems like even though to the front end, it looks a lot like a NetApp ONTAP filer.
Um, I dunno if you call 'em filers anymore. I still do. I still call 'em texters, um, on the back end.
Um, it, it is totally new. That's, it's spot on actually. Yeah.
So, so the, the data management layer, the thing that makes ONTAP so popular and, and, and, you know, the, uh, so strong for enterprise is all intact, but you're absolutely right. We managed to sort of rebuild the house underneath the change, how we manage storage, um, and how the cluster connects with itself and how the cluster scales out. So, you know, on the surface, it, you know, customers will find things very similar, very comfortable.
Um, but, but what it can do from a performance standpoint and a throughput standpoint, whole new beast. So let's dive into that. So number one, uh, you've got the, uh, the a FX, uh, head unit.
Yeah. Um, controller. Controller, uh, and those are basically a scale out architecture and this, and the storage scales up below them, right?
So it's up and out. Yeah. They look and feel a lot like our other storage controllers.
One difference is they have internal boot media. So, so, and then it's change is it boots ONTAP for the internal boot media instead of off of, off of shared disc. And that allows ONTAP to boot and then join into the cluster.
Now, as soon as it joins into the cluster, ONTAP then goes, Hey, I've got more resources. And we'll automatically load balance and start adding more workloads immediately to those new nodes. Um, when I add new shelves, we bring those online and ONTAP automatically scales those shelves, and all controllers have access to all shas, all shelves.
So it's a full meshed, um, fabric. Any workload can run on any disk or run on any node. Now, ONTAP has always been good at scaling up.
Uh, that was one of the reasons that, as I said, we bought it many, many years ago because we could a buy some now and add some later. But the problem is scale up only gets you so high, right? Yep.
And now for a long time, NetApp has been working on scale out. Um, give me the big number. How big can a FX scale out?
We're, We're, we're launching here with 128 nodes. And that's not an architectural limit. That's just for right now, we don't expect anybody to need more than that.
But architecturally it can scale up much beyond that. Um, so from our previous limits of 24 nodes, that's a pretty dramatic shift. And the fact that workloads are never trapped on any ha pair, a workload can move around in the cluster when or if needed.
Yeah. That's a big, uh, differentiator here, right? So that essentially, uh, the workload is sort of, um, fully virtualized in that, um, you know, the client accesses the cluster, the cluster decides, uh, which node, which disc shelf, where to put it, and it can dynamically move that stuff around to adjust to the demands, right?
Absolutely. We refer to it as, as copy free migrations or, or zero gravity migrations. Um, and, and it happens periodically as the cluster's running.
So making sure that no node ever runs out of compute resources. We can move things around and avoid that, but it also happens when you add new nodes. So we make sure we take advantage of those new, new resources right away.
It also happens if you have node failures. So as soon as a node fails, we can then take the workloads that were running on that node and distribute them across the cluster. So we're, so it starts to uncouple this concept of an HA pair.
Now we still add nodes in pairs for now, um, but the, but the data is not trapped and the workloads are not trapped on those particular two Nodes. Yeah. So you have a really flexible, and what this means, I think, is that as you get a lot of these nodes working in concert, maybe not 128, but a whole bunch of 'em working in concert, uh, then if one, uh, needs to be taken offline, I don't wanna say fails 'cause they're not gonna fail, right?
Never. Never. That would never happen.
Uh, so, so if one needs to be taken offline or whatever, or if one is added, uh, the customer doesn't really have to worry about it because, uh, the performance, I mean, probably isn't gonna suffer at all noticeably, because you've got so much resources in that cluster. Yeah. And, and, and one of the things that's quite different on a FX is how the level of automation we've put into it.
So, so you're not having to make those decisions is, is when a node goes offline or when you take a node offline, the cluster automatically finds the optimal places to place those workloads. When I bring a new cluster or a new no on the, the, the cluster automatically determines what workloads are best to place on there. Um, in fact, if you set up the cluster in such a way where you provided a pool of IP addresses, it'll automatically stand up all the data lifts, right?
All the connectivity for that new nodes actually adding nodes is so much easier in a FX And under the hood. Uh, obviously things have gotten a lot better since those big bulky, uh, scuzzy cables from the past. Uh, you know, this is all ethernet, it's all switched and there's almost a, a, an any to any relationship right?
Between disc shelves and controllers. Yeah, absolutely. It is all, uh, industry standard, a hundred gigabit ethernet.
Um, we're selling it with Cisco switches and there are new 400 gig, uh, switches, uh, really help simplify and increase the density of that switching. Um, and one other major change is, is, you know, versus our A FF that you may be familiar with how a cable that we would do the HA across that, now everything is tied into the switch. So it's, whether that's cluster communication, whether that's communication to the dish shelves, or even the ha mirroring of the NV ram, all of that happens across those Cisco switches in the backend.
And That's seems to be sort of a theme here in that what you've done is you've got done away with a lot of the sort of proprietary and Oh, but we actually need this kind of, kind of, and there's no more, oh, but we actually need this. It's, it's connected to a switch and the switches are connected with a hundred gig and we've got plenty of bandwidth, and you can just build something huge If you need it. Yes.
The thing is, is, is actually the starting configuration is not that huge. Starting configuration is four nodes and one shelf of disc. So, so, so it doesn't have to be huge, but from that one configuration, I can start adding nodes, adding nodes, adding nodes or shelves, and I can reach that exact scalability point that I need without wastage.
I'm not having to add capacity just to get the performance. And, uh, but capacity and performance you can scale, uh, very, very high. So, I mean, I don't know, what are the, the highlight numbers for throughput and for capacity?
It, it's mind blowing to be thinking that I'm talking about, you know, throughput in terabytes per second. But that's absolutely where you're getting to is, is, is the terabytes per second both reads and writes. Um, and then from a capacity standpoint, you know, depending on, on a hot versus a tiering, um, well over an exabyte.
So I think on the stage today, they talked about exabyte scale, you know, over an exabyte if needed. And it sounds like it could probably go bigger if you added more than the 128 nodes and more disc shelves and everything. But, um, so far you haven't gotten there.
Exactly. They're not architectural limits. One of the things is when we did this re-architecture for a FX, um, we wanted to make sure we engineered it for the future.
We didn't want to do another re-architecture, non-AP. And so it can go bigger, but those are the limits we're launching with, because I think that's gonna cover pretty much all the workloads out there, or most of the workloads out There. It's gonna cover a lot of them.
Yep. And even the AI workloads that get very, very big now, um, one of the things that I think we should talk about is the front end, because one of the challenges, uh, for AI workloads is that you have really high levels of parallelization of access to data. In other words, there are many, many nodes, many more than we are ever seen before, accessing sometimes the same data, sometimes accessing completely different data, just a real, uh, uh, you know, IO blender going on on the front end.
And you've built this thing for that use case as well. Well, it, it does have a full set of protocols on it. So we do support, you know, SMB, we do support, um, S3 protocols.
So we are seeing training and inferencing even being done on S3. Yep. Um, but then NFS, we do support NFSV three, but primarily the sweet spot is gonna be NFS V four with PNFS.
And that's been out as an industry standard for a while, but it's really reaching this sort of scale where I have a, you know, a, a cluster that needs that sort of capability. And when you couple PNFS with that ability for ONTAP to move workloads around and then report back to the clients, what's the optimal path to that data? Super powerful combination.
And that gives us, um, parallel file system like performance without actually having all the headaches of app parallel file system. Yeah. And, and also PNFS at this point, I, I know that, you know, storage nerds like me have been waiting for it to, to, to come alive, but PNFS at this point is extremely widespread on the client side.
In other words, most clients are PNFS ready, right? They are, absolutely. And the other thing that's happened in that same time is the performance delta between NFSV three and V four has narrowed right down, right?
So, so a lot of people are like, yeah, I, I, I like the functionality of V four, but I want the raw performance of V three. That performance delta gap has, has narrowed right down. And with PNFS, you get that incredible scalability and optimal pathing.
So moving up the stack a little bit then. So obviously this is gonna be used in AI training. Uh, it'll probably also be used in the sort of, uh, heterogeneous, uh, AI applications that we heard about on stage today.
And, uh, basically at every event that I've been to this year where, uh, we need to basically present data. Uh, NetApp also talked about an AI data platform and, uh, AI data classification. Uh, how does this, uh, a FX fit into that overall picture of AI data readiness?
It's the foundation, right? So, so, so it's, it's powerful in what it does. And for some customers, that's all they need is that, is that the, the, the foundation, but for those of are the, that are moving, the next layer is the metadata engine.
They refer to that, and the metadata engine plugs seamlessly into a FX to give customers a, a searchable catalog and browsable catalog of all the data that's consuming space in that cluster. And that can be, you know, just for, you know, showback or chargeback or just, you know, what's consuming the space in my, in my cluster. But also that is where you start to curate and understand, you know, what data do I have, where is it, and how could I use it for ai?
And then on top of the metadata engine is where we apply, we use the, the APIs from that to apply the advanced data AI services. Mm-hmm. So essentially this is, uh, really the foundation of everything that NetApp announced today.
Now, it doesn't have to be though, I mean, this is absolutely useful in a variety of applications, but if customers are looking for a system that can just scale to really massive proportions that can scale performance, that can handle the kind of parallel access that you're gonna find in AI training scenarios, um, A FX is it 100%. And, and as you mentioned, you know, we're bringing those advanced services to the whole ONTAP because this is still ontap, we didn't fork the code. That's an important thing to call out too.
This is still ontap, this isn't a separate build or a, or, or, or a separate, you know, uh, set of APIs. It's still on tap, but it does give that massive scalability. But you're absolutely right.
We can take that data elsewhere in the portfolio. Well, this is great. I, I'm really excited to hear about a FX as again, I'm, I'm a longtime follower of NetApp.
Um, you know, uh, very exciting to see how you're taking the platform forward, how you're taking ONTAP into the future. Uh, we'll be learning more about this during our tech field day sessions here. Uh, we'll be live streaming those on Thursday, and then of course, putting them on YouTube.
If you're watching this after, uh, the week of NetApp Insight, you can find all of these videos on YouTube. Uh, Keith, thank you for joining me. Thank you for giving me this little, uh, behind the scenes peak at the development and the launch of a FX.
Thank you so much for having me. We're thanks for coming to Insight, and we're super excited to share this with you.