Unlocking Innovation with Modern Data Infrastructure from NetApp
Infrastructure modernization today goes beyond simply upgrading storage. It’s the cornerstone of breaking down silos and establishing a unified data foundation that drives innovation across your organization. Whether you’re optimizing hybrid operations, enhancing cyber resilience, or accelerating your AI journey, this featured session will demonstrate how an intelligent data infrastructure, such as the NetApp data platform, offers unparalleled simplicity, security, and efficiency for all your workloads.
At Tech Field Day Experience at NetApp INSIGHT 2025, James Kwon and Pranoop Erasani introduced AFX, an advanced architecture within the NetApp data platform, designed to meet the growing demands of AI and other high-performance workloads. They explained the project’s origin, focusing on how traditional ONTAP architectures struggled to keep pace with the rapid computational advancements of GPUs. AFX breaks from the high-availability (HA) pair constraints by disaggregating storage and compute components, thereby allowing independent scaling of capacity and performance. This approach offers greater flexibility to customers, who can now customize infrastructure growth based on workload requirements rather than being locked into synchronized hardware upgrades.
The AFX design introduces a single storage pool architecture, eliminating redundant storage layers such as aggregates and simplifying both the user experience and storage management. It supports ONTAP interoperability and maintains near-complete feature parity, including capabilities like SnapMirror and FlexGroup, while delivering automatic rebalancing and volume re-hosting for seamless operation. While three separate cluster “personalities”—unified, block-only, and disaggregated—are maintained, features like zero-copy volume moves and simplified expansion reinforce the efficiency and adaptability of AFX. The new system is ideal for AI workloads given its throughput optimization, yet flexibility in design paves the way for use cases in EDA, HPC, and other data-intensive sectors. Though not yet a wholesale replacement for unified ONTAP, AFX represents a foundational step toward universally modern, scalable, and intelligent storage solutions.
This presentation was recorded at NetApp Insight 2025 in Las Vegas on October 15, 2025. Watch the entire presentation at https://techfieldday.com/event/netappinsight25/ or visit https://netapp.com for more information.
Transcript
I'll kick it off with some of the reasons why we started working on the FX project and what the goals are and what are the capabilities. And like we said, go a little bit deeper. And now I have my, you know, partner in crime pup go to, go way, way deeper.
So let's go with the show. So why, why did we embark on this project called A FX? Right?
Number one reason was that two years ago, we foresaw that with the evolution of the GPUs. That the way the current, at the time the current HA architecture on ONTAP will have trouble taking us for the foreseeable future. The advancements and timing of those advancements on the GPU in terms of performance and the scale requirements was way, way rapid.
Then the normal evolution of how we evolve our storage hardware. For instance, if you look at Nvidia GPUs, right? Approximately year and a half cycle, they increase their performance to X.
If you look at the storage side, the storage industry hardware refreshes take around three between somewhere between three and five years. And when we do a refresh it, the performance increase hovers around 35 to 80%. So there's an offset there right now with the current architecture that we had on ONTAP with the HA pair that meant for our customers to add in more nodes, more shells, because they're coupled together.
And the outcome of that will be kind of off balance and made our customers pay more. So we need to do something about that. The other thing is, we also saw that initially when, you know, LM training and foundational training was all there, it was all about scale and performance.
But also when we looked into enterprise ai, those enterprise AI customers saw more value in terms of data management. All the rich data management that we currently have on ontap. So that's the other element.
And the third is of course, security and multi-tenancy that ONTAP provides. So those were the background that we saw with a FX project. And now yesterday we announced the A FX.
It is a new architecture from an ONTAP point of view and also new hardware offering. Unlike the traditional ONTAP offering this from a architecture point of view, it this aggregates between the controllers or compute and the storage nos like Jeff Baxter mentioned before you, that allows our customers to scale independently. Now, if you take that notion and think about the GPU evolution that I talked about, a customer can buy the latest and greatest NVIDIA GPUs in year one.
They upgrade in year two, NVIDIA comes out, double the performance, then your storage evolution is three to five years. So what do you do? Right?
In this case it's disaggregated. What it allows you to do is, okay, if that is the case, then I just add more controllers or compute to keep up with the performance requirement of the GPU and vice versa is also the same. Let's say you are, you are starting out with a small amount of capacity and you're increasing your AI applications or whatnot, and you need more performance, uh, more capacity instead of performance, then you could add more disc.
Shes to increase your capacity independent of your compute. So that's one. Yes.
So this means that when you said that it's new architecture and you said it was new something else, um, New architecture with new set of hardware offerings, New hardware. So that means that, um, if I'm a NetApp customer mm-hmm. And I wanna take advantage of this, I have my normal storage sitting here, but then I have this additional piece of storage that that gives me the AI functions over here.
Correct. That gives you the scalability and the performance and the other AI functions, right? With A IDE that resides with a FX.
The more important part there is, yes, there are two different hardware products like the A FX, we call it a FX because of a different series of set of products. The important part is where is your data, right? You have your data on your, let's say on your existing A FF, right?
And you, you deployed your A FX system because they're running the same ontap. You could use all the ONTAP function features and functionality between the two, like snap mirror, like SVM migrate, like flex group, et cetera. All the replication technology and all the data management capabilities are uniform and compatible between the two systems.
Can I add a FX nodes to an existing cluster? You cannot because the layout of the storage and how the compute or the controller scale requires a different set of software architecture below. So that is not possible.
That's why we focus a lot on the unification of the feature and feature parity between the A FF and the A FX because they're all running the same ontap. So are you gonna get into like, um, use cases of when we would use an a FX versus I know, like it's, it's this high level. Mm-hmm.
I, I think, but I'm like, why can't, if I've got the GPUs and I've got everything else, why can't I just keep it on my regular old? Why, why do I have to change platforms? Yes.
You, you certainly can. And do we have a lot of customers using our A FXA 90 for AI use cases? Because prior to our announcement, A FFA 90 was the only super PO certified ONTAP device that we had, right?
So in terms of the differences and positioning between our unified or our A FF versus block only versus disaggregated is as such, right? Unified or the AFS that we have are will be the bread and butter that ca caters to majority of the use cases that we have. And it has a bunch of different series of, uh, hardware products in there to cater to.
Very small to very large, right? That will be our bread and butter. And you might have heard of a SA, our dedicated block devices, right?
So that will stay and disaggregated will be unstructured data only. So file and object that caters to more of a high performance, high scale type of use case, including ai. Or you might, one might expand to use cases for EDA, for instance, or MNE.
And just to Con go ahead. Yes. I Was just gonna say, just to confirm mm-hmm.
To work with the data, sorry, that data has to be on the A FX, correct? It can it work with the data that's already on the a FF, the, the, the AI data engine? A IDE, um, for a ID for a ID data engine pressure.
I think the next session is the A IDE team. I think that's a better question to ask them, right? I I think I could simplify the question.
If I've got data on a unified cluster, I can snap mirror it to the disaggregated cluster. Yes. Right?
Yes. So that's one way to integrate it. But am I correct in also assuming that if I start a cluster with one of these modes, I must continue that cluster is gonna live as a unified, that cluster cluster's gonna live as de dedicated block.
Yes. That cluster's gonna be disaggregated. There's gonna be three different cluster types, three Different personalities That live in that, in those personalities.
Yes. And that's it. You're not gonna mix 'em.
No. And, and for obvious reasons, but yes, I wanna make sure That's correct. Yes.
That was simplifying the question For a NetApp tuned person. Yeah. Yeah.
One of the questions that keeps coming up for me is, okay, I've been a NetApp customer forever and ever and ever, and I have lots and lots and lots of historical data, which is the stuff I wanna use to, um, to power my AI journey. Yes. So would I just be, um, snapping those over to the disaggregated personality?
So I'd have two copies of the data, my old historical version that's maybe not even on a NetApp, NetApp, maybe that's in Iron Mountain, and I have to come and get it and restore it to the NetApp and then snap it over to the disaggregated side. And so then I'm, I've got the two copies and at least Yeah, at least if, if it's definitely, if it's a third party non NetApp, then yes, you will have to kind of migrate it over to an A FX and let the A IDE do its job of categorization. And from there on then you could see the benefits of reduction of the copies of the data.
Okay. Alright. Okay.
Any other questions? All right, so let's move on to the goals of A FX. The first goal was to eliminate data silos.
What we mean by this, we talked about this independent scale, independent, um, independent scale, uh, performance and capacity, right? Where we're, what we're, what we did here is for those of you who are used to NetApp, is we have something called shelves Raid Aggregates, right? And, and on top of those aggregate, we have volume.
So we had a hierarchy of kind of containers that consists a storage container, right? What we're doing with the A FX is we're meaning significant amount of those. So we're getting rid of the so-called aggregates, and we're going with an architecture that will provides you the so-called single pool of storage.
So all the storage capacity disks and shelves that you attach to an A FX will represent itself as a giant single pool of storage, right? The other part is what we mentioned before is the ONTAP interoperability. We want to maintain all the data management and feature capabilities that ONTAP has and have those supported on a FX as well, so that we could hit close to a hundred percent feature parody.
Third is future proof storage. So this is a first release of A FX, right? We have a lot of plans to enhance it in many ways that we see that caters to an AI workload or any high performance workload that we will start disclosing them as we get closer.
But in terms of storage efficiency, in terms of new drives, in terms of, um, you know, metadata operations, those are the key areas that will continuously evolve. And also absolutely we do not forget increasement of performance and scale, right? We will continuously increase that as the releases goes on and on from the first release.
Okay. So now you said it future proof. Yes.
If ONTAP is the current unification mm-hmm. Point, uh, can you give me the date that A FX becomes the unification point? No.
Right. I'm, I'm being asked. Yeah.
I mean, like, you know, is, is that, is that something you could see whether, whether it's a FX functionality moving down, maintain the ontap or, Yeah, so You're smiling. Good Question. From a feature functionality point of view, disaggregate sounds awesome, but if you look at the workload, I Don't know, I'm sorry to interrupt as a unification point.
Mm-hmm. Right? Do we shift the APIs out or allow for, you know, right now, uh, ONTAP is a single unification point, limited functionality, so to speak, in comparison to what a FX can do.
Um, but um, you could move some of that up with a pass through, or like, there's a lot of ways that you can potentially with a lot of really great engineering turn a FX into a unification point or some, some other way in which this doesn't become such an issue. Mm-hmm. Uh, architecturally of, I wanna keep using my existing stuff, but I want to add functionality.
So, So from a feature functionality point of view right now, it's when we, now, since a FX is released, when a feature is developed, it's developed on, on tap. So the goal is, and the objective is to have it on all, not one or the other, right? So that's the benefit of having a single storage os right?
You don't have two separate ones that I have to do this on this version on, oh, I have to develop a complete separate one in there. So that's the benefit, um, in terms of where that total unification, you know, eventually will happen. I think that's kind of, kind of out there.
We don't know yet, to be honest, right? Because there are benefits of disaggregation and there are, you know, use cases where disaggregation does not make sense, especially right now is smaller deployment kind of workloads, right? Disaggregation there does not make sense because that's perfect for an ha pair architecture right?
Now, one can say, well, you could pare down your disaggregation to make it an yes, technically that is doable. But the challenge there is, is that going to be cost effective at this point in time from a dev test and et cetera, et cetera. Right?
It's really in the context though of future proofing. Yes. This could be important.
Yes. So in terms of answer, When, when I come a few slides right on the Yes. Is that why you were smiling?
'cause I'm prefetching your slides When you said, When, when you said it was a, a word there, I said okay. I, I smile because a lot of my engineers has the same question. Yes.
Okay. Good. Thanks for that.
Okay. I speak forward answer. All right.
I just have a couple of quick questions. Sure. Naps, what, just over 30 years old now, um, on Tap's been through a significant number of modernizations and things already, but, um, in approaching a FX, uh, how, how, how are you handling some of the constraints of technical debt and legacy decisions?
Um, 'cause there will be some, any code base that's 30 years old is gonna come with a number of pieces of technical debt decisions that were made that were great for the time, but you wouldn't make again now. And how do you balance that? Hmm.
I think the easy way for, for me to answer that is top down view easy button first, right? Things like, yeah, San we made a decision not to support it because it was not needed for the workload, right? Or other features that we see that has, like, um, one of the highest technical depth that we had with this new architecture is elimination or aggregates going to a single storage pool, right?
So those kind of areas, we looked at it with a magnifying glass in saying that, is this something, is this feature that ties to an aggregate that needs to move over to a single storage pool? Is this required? If so, how urgently, or is this something that has a replacement feature already on hand for us that we will elect not to do?
So we went through a series of actually a long, long series of prioritization to make that decision and started executing on it. Excellent think, I think that That probably answer is, so what contradicted you to what he was asking, which is when is that unification point for every workload, right? And, and I think a, it is an interesting question, right?
I mean, we had to start with, uh, a growth workload, right? I think AI is a growth workload. AI is gonna be in front of every other workload that we have seen for the last, uh, 20, 30 years, right?
Uh, so we are to make tough decisions to start with, but eventually it all come together, right? You know, hey, we are optimizing for throughput on the disaggregated, we are optimizing for latency on the, uh, block, right? So we are to make some decisions there, right?
Are we gonna do something like metro cluster going forward, or are we gonna bring in another zero RPO zero audio solution that is much more flexible, much more granular, right? So these are the decision points we need to make and, uh, sort of move forward. Creating these two personalities of products will help us actually make those decisions better, right?
And not confuse like existing customers versus new customers. Okay. Thank you.
Okay. So let's kind of double click into kind of each of the key areas that, or the, or the benefits that A FX provides. Number one, as we said, independent linear scale in terms of performance and capacity.
Because it is disaggregated, you could increase your node or increase your capacity independently based on your needs, based on your workload, based on your budget, whichever way. Right? And one of the things of the benefits with the architecture is when you start adding something, when you start adding, let's say more, uh, capacity or more nodes, uh, the system now has the intelligence to rebalance them as they get added automatically.
You do not have to kind of manually go in and try to kind of offload your kind of, uh, volumes or whatnot to from one node to another, because now this is disaggregated different architecture with a single storage pool. The system knows about all the components that you have from a controller and a shelf point of view, and they will automatically distribute the load as it comes in. Yes.
Are are you spreading volumes since I'm assuming we're still using flex files as a unit of measure, um, across all the discs in the disc pool? Or are you, are, are we still bound to a physical aggregate underneath some virtualization layer? Just, uh, to be clear from, So let me answer that from a more, more of a one layer down.
Okay. So all the nodes will see all the drive, right? Number one.
Right? And on top of that, the default for a FX is not a flex fall. It's gonna be a flex group.
And by notion of flex group, it spans out. Right? Now, if the customer elects to do a flex fall, yes, they can create a flex fall, but the default will be a flex group in this sense.
Mm-hmm. Right. So flex group works beautifully with the new A FX architecture because its flex group is a capability that utilizes all nodes and spreads the constituents the volumes out, right?
So by default, flex group is configured, then it will see all the nodes and it will see all the, uh, all the capacity, and it'll spread it out that way. Yeah. But now that you're disaggregated the controller and the storage level through switches, you don't need to have secondary pads for that data.
You don't have to have the non-optimized pads anymore, I guess, right? Yeah, It really helps on that one. Yes.
So You, so that's true that, that's correct. That now every node will have primary access to its own data. That's pretty cool.
Yeah. Okay. Now in terms of performance, because again, it's disaggregated and we built it.
So one of the key goals is yeah, when you add your more controllers, the performance increase must be linear, and we have achieved that, right? So ad controllers for CPU throughput add shelves for disc, disc and spindles. What happens is, in this case, if we're going from a two node, four node, eight node, right?
The performance profile is linearly going up. And that's one of the key things that we wanted our early access customers to test. And all our early access customers has verified and confirmed that.
Yes, as I'm adding note, I do see a very, very, you know, linear performance increase, which helps them a lot, which they liked, Right? Sorry, it's to the minimum size of notes that you need to have or, So when, from a offering point of view, four is the minimum. Okay?
Right? From a technical point of view, yes. Two is ONTAP capability can handle two, but from a offering point of view, four is the minimum because we want the increments of two after four.
Right? Okay. Okay.
Improved efficiency. This revolves around the new architecture of single pool of storage, like we said before, all nodes now see all drive, right? And this provides a lot of efficiencies in terms of how you use and manage space, right?
So the notion here is what we wanted to do is, and, and the goal is have the performance of raid while having the efficiency of e erasure coating, right? So we want to have more resiliency with lower overhead. We have implemented something called a zero copy vol move since now we're talking about a very large container, right?
So that, you know, within the cluster, when a vol move happens within the A FX, they're all zero copy vol moves, which is also very, very effective when we're handling with failure scenarios when let's say one node fails, then because we have this capability in the background, right? We could efficiently move that connection of that volume to a different neighboring node connect. Exactly.
We're just moving the connector, we're just moving the pointers, right? So those will be the benefits. Operationally, it's very, very simple because now you could dynamically add both compute and capacity separately.
So space management now is much easier because you do not have to say, oh, is this container or shelf a, a part of node one or node two, node three, you don't have to do that. It's all in one pot. So it's very, very easy to have that top view of, oh, oh, is this much, I have this left right?
And then these are the volumes here and there. Right? It's very, very easy in that sense.
Um, and also of course, no aggregates manage, right? Uh, also a lot of automation has come in, we put in for managing this. So ONTAP manages now manages the rate groups and all the other things kind of automatically, right?
So that the eventual goal is when a user adds in more capacity, it should be just automatically added in and it will, it should utilize that empty space that was just added in appropriately. Right. And optimally, To go to guy's question before though mm-hmm.
That first bullet under a single pool of storage is the use case for the small two node. Why you even on the smaller side mm-hmm. Those customers typically don't have the expertise to manage storage.
Correct. Any way you can make their life easier. So they don't have to worry about a aggregate on node A, aggregate on node B.
If they just had that, I think that would also be good under the small scales. So That is a good input. I also say, I'm gonna make some assumptions here, and you can correct if I'm wrong, but the assumption that you've removed the root volumes somewhere, they're in the compute nodes, some somehow different, because that was a root Nos are tough.
Yeah. Yeah. That, that's has always been a tough thing in the Root valves.
Yeah. The root valves and everything else, that's all been really tough to handle in the mm-hmm. Traditional HA stuff, right?
So make the assumption, and maybe you'll cover it later, but that's, that's gone. But having that gone mm-hmm. With the matic space management ray group definitions and all of that stuff, I can see a real interesting place for this in the smaller end.
Gotcha. Um, because designing, you know, a DP was great, but it was a sticking plaster. Um, and so, you know, removing things like that mm-hmm.
Make, just for the smaller systems as well, the storage admin's jobs so much easier. I can see like the amount of time that went into architecting and balancing systems and doing performance and migrations and other things. It's just for Lower end.
For lower sales. For lower revenue. Yeah.
Yeah. And, and it causes your support organization headaches as well. Yeah.
Yeah, exactly. So there's a huge future there for replacing some of that. I think that's good input.
Yes. Makes sense. All right.
Um, data mobility. Um, so lemme just build this out, right? In terms of data mobility within the cluster, we utilize all data mobility features on ONTAP because it's running the same ontap.
Now, these kind of data mobility capabilities have a different meaning because now we're dealing with, again, such a huge, you know, single pull of storage. So vol, things like vol move is gonna be much, much more heavily used by our customers and by the system itself, right? Uh, like I mentioned, you know, flexibility, no failure scenarios before, right?
On an HA pair. If one node fails, then yes, you have to kind of accommodate that 50% up to 50% performance loss. But on this side, since it's disaggregated and you may have eight nodes and one node fails, then it's not a 50% because of, we could shift the pointers on the volume in this case and so that all the other nodes can participate and assist on that down.
So it's not a 50%, it's minimized in that sense, right? We're also making more enhancements on our failover capabilities, and we'll kind of, you'll see the enhancements as our, you know, A FX releases kind of go on, but right off the bat, these are the benefits that we will be providing for our customers. Simplifying expansion.
The scenario of adding a node is very, very simple. You add a node, the system will detect it automatically and it will not add it in because you may, the system administrator wants, may want to have a separate cluster, whatnot. So it will discover it and tell for a system manager that, Hey, I discovered these notes, A FX, no.
Do you want to add it or not? If so, okay, add it in off you go. Right?
Similar with capacity, but more so capacity is more of a, you know, when you, when it's added in, yeah, it automatically detects it. Like I said before, it automatically adds it to the single storage pool and it'll automatically low balance once it's added in. Do, do all the nodes act as a single global namespace?
Say that again. Do all the nodes act as a single global namespace? That Depends on how you configured, right?
If you have an A FX with what, two petabytes and you start putting in a bunch of flex vaults, which are in its own container, then no. But if you have a flex group that spans across the whole a FX system, then it's a single global interface in sense right now outside of the cluster, then we have other technology and capabilities like flex cache and whatnot that will expand that global namespace to other data centers or other clusters or other nodes for that matter, right? And how are we kind of low balancing or organizing connectivities into the many nodes?
If I've got, uh, 200 storage nodes, um, what is it DS, round robin or some form of BGP or like how, how do we connect then You them? So, uh, today's DS road balancing, we are looking into on box again. Um, I mean, there was a time when we had it, but there was some challenges with it.
Um, that's an interesting idea about BGP, right? That's another technique we are looking at to, to even like consider like moving a individual volume from one cluster to the other, right? Yeah.
Um, uh, I think that in progress actually, uh, but not there yet. But for now, I think our recommendation stands, which is, uh, uh, DNS load balancing, right? And many of our customers prefer off box load balancing because they have, they do load balancing at a, you know, data center level, right?
You know, they have, uh, they have build mechanics into that to, to use, uh, you know, with the load across clusters and everything too, right? Yeah. Yeah.
Just that there's some similarities with some of the ways that storage grid and we have the low balancer nodes in front of them and it's, you know, is there a potential officer in providing something for a FX in the future? Correct? For S3, I think there's a good possibility and you know, it's being stateless, right?
In many ways, there's a good possibility to do something like storage grid, right? Uh, but for now we are leveraging outbox, okay? That is yours.
Okay. Alright. Uh, it's really hard to explain architecture with a marketing looking slides, right?
But, uh, I'll, I'll try my best. I, so, uh, if you think through, if you think through what it's today, right? I mean, there's a set of clients and, uh, data network, um, uh, and you, you, you know, our familiar protocols, right?
If you look at what I'm showing here is a unified architecture, right? Which is, which is our 30 year legacy, right? We have file, object, and block.
Um, we had A-H-A-A-H-A-A was a domain of operation for the aggregate, right? Uh, aggregate is a collection of this. Uh, and that actually, uh, where you used to provision volumes explicitly, right?
So if you had to move, if you had to move data or balance data, then you have to actually move from one aggregate to the other. Um, back in network, you know, was provisioned for, you know, cluster, uh, network vlan and cluster, uh, you know, um, even, even, even all the HS are connected through the cluster vlan, right? Uh, we could scale up to like, say 24 nodes, right?
In some specific cases we have done 40 nodes, but with aax we're gonna go whole lot of big, right? We used to, so pay attention to the color as, as I can represent the aggregates. Um, and specifically I'm representing the modular systems we have today out there, right?
Which the, which are like we distance storage are connected by the network, right? And in this particular case, you can see, uh, even though they're connected by the network, right? Uh, but the, the, the islands of storage were created for aggregates and connected to the, uh, hs.
And these are the differences with, uh, a FX architecture, right? Uh, I'll probably go through one by one, uh, you know, so that we can follow through. Um, now we have a storage vlan, uh, which actually effectively, uh, a separate network for storage as well, right?
Uh, we are trying to consolidate that as well, uh, for reduction of cost. So there was a great question about, Hey, are we really done? Can we pivot to That could be the, the central storage for everything, right?
Um, we had to really look at, for some workloads, like for example, um, tunnel clusters with no switch, right? There are a lot of deployments there, right? We are really look at it and say, Hey, this does this architecture, uh, uh, help with the cost model of that, right?
Uh, today we can, we can see that, right? But as we look to the next generation of our platforms, um, that's our focus, right? How you get into midrange and, you know, low end and et cetera, right?
Um, so interesting thing about, you know, this, uh, disaggregated ontap, uh, is Hardware, we didn't plan for hardware and software to release at the same time. Meaning in, in a sense of we didn't have a new hardware to play with this, right? We took what we released for, uh, unified on type, which is, uh, a one K, uh, 8,000, and we actually use that as a foundation for developing dis aggregator on type.
So it may not look great today in terms of the density and everything, but at the same time, now that you have a software architecture that actually can disaggregate it feeds into the, you know, next generation hardware discussions, right? Which actually can help us, you know, get to our, uh, the best price performance intensity, um, even knowing that we still optimize our A one K platform and we call it a FXA one K for a reason, because we had to make some modifications to the hardware for better performance, right? The other aspect is like shelf for the first time, I would say at least, I, I've been in NetApp 21 years, right?
For the first time we had to pay attention to the shelf performance because now, uh, every, every note sharing, uh, is connected to the same shelf, right? And what that means is that in our performance of all note, setting the shelf at the same time can actually, uh, requires more bandwidth, um, modular systems. You know, we didn't have to to worry about that.
We have only two nodes connected to it, right? So some of these critical differences come into play, you know, to answer that question. Hey, are we ready fully for making it mainstream for every workload, right?
Uh, so it's probably gonna take like few years before we get there as we look at unified and all the use cases and move there, right? For now, the focus is ai. And when I say AI as an engineer, for me, the focus is throughput, right?
It's very optimized for throughput architecture, right? It's flex group that also makes, Makes it incredibly useful for EDA and HPC, right? Absolutely.
Absolutely. Yeah. Yeah.
Well, every workload, Anything that needs real high performance in, in, in, absolutely. And a lot of work in EDED is really performance Restricted. No, absolutely.
Every workload as you, you, as you know, right? And, and has phases of, you know, uh, different IO operations, right? And one starts with, you know, latency sensitivity, then data management, and then you get to, uh, throughput, like tape outs and everything, right?
Um, so, but we are optimizing from the other end here, right? But still leading with all the features, right? Our intent is firmly to support all the workloads, right?
But today the focus is ai, right? Because that's where the biggest demands are. And another way to look at it is that, hey, AI is gonna be in front of every other workload, right?
You can consider it a vertical or a horizontal. To me, it's a horizontal eventually, right? Because inferencing is gonna be part of every workload.
Oh, It's a horizontal now. Yeah. I I, I, I've stopped saying it lately, but I'll, hey, this is a great chance to say it again.
AI is not a what might be a workload. No, it's not an application. People talk about AI applications.
The applications of the applications, ai, multiple ai, hundreds, thousands of ai, including agents are being exposed and should be exposed to as services that can be picked up, whether ad hoc or integrated by application. That's scale for us. The way you're saying it, that's scale, right?
Um, I totally agree with you, right? I mean, I, I was probably one of the first ones who used to say that this is actually horizontal more than vertical. But as a storage company, it always starts as a workload, uh, because it's easier to identify in terms of requirements, right?
And then we bubble it up, right? Uh, in terms of app integration, like a ID et cetera, right? Um, so We don't support block today, not a surprise.
Um, we all nos share the same discs. I is a very important point, and no individual node owns a particular set of discs, right? Uh, it's all abstracted under, uh, when we create that aggregate pool, uh, a single pool of storage is what we call it.
Um, there's a certain, certain layout we follow with certain, um, uh, you know, uh, data disks and parity disks, right? For, for rate configurations. Uh, the configurations will keep changing as we expand the, the capacity pool.
Um, and HA is connected by VisaNet switch. So HA network, uh ha goes over the network, I think. So to go back to the point I made about, uh, hey, we took what we had for unified ONTAP A one K, and we supported it with a FX.
Um, the HA actually had to be put on the network, right? For that. And, and what that means is that eventually we get to NVHA where three node failures or two node failures, three node failures will be absorbed by any node in the cluster, right?
That also helps us actually remove that constraint of, you know, hey, planning for headroom, 50% headroom if there was HFA lower, right? Uh, we don't doubt that a lot, but, uh, you know, hey, um, but that's gonna be probably a game changer for many customers who wanna optimize cost performance in terms of the deployment, right? The plan for like five systems versus six, right?
Now, you have that flexibility. Um, of course, computer and capacity are deport, and that's, that's, that's whole point of disaggregated. Is this, a lot of this all over r dm a, Uh, yes, absolutely.
Was there thought of using anything else? Uh, Inna Band, CXL, sorry, direct, direct CXL or any of those newer, We are think we're looking at CXL or our platforms don't today, right? Um, there's a consideration I would say right at this point.
Yeah. And, and, and you know, if you ask Nvidia, they'll tell you, you n we link, you know, can you put that in there? Everything in there, Alright.
Invested 5 billion into Intel to get it right. Exactly. We still have the concept where we have a volume that is owned primarily by a node, and that data will be served primarily by that node.
Yeah. I will say it's hosted on a node hosted, hosted on a node, uh, because it's computer again. Yeah.
So why do we have the volume as a concept? Because it's, it's an, it's our way of look, uh, enforcing data management and consistency, right? You'll still have the volume as a concept, but I think you, you touched upon a question I think I I was tempted to answer, but I was holding back.
Um, I think you were asking a question whether the same volume can be accessed from other other node, right? And, and as, uh, IP another request comes from could do today. Yeah.
So, um, we are looking at some options where at least the read part can be optimized where you, you still need right path to go through one node for consistency, right? Because they have to go through NV m uh, nv man, like in future, uh, but you reads can be optimized where I can instantly flex, clone, and then boom. Like I can serve the read through the different node.
Can't you just, uh, do, um, LACP across multiple nodes, make one virtual interface across nodes? We, Uh, I I know what you mean, right? Uh, and we could look at that.
Uh, but I think there is some overhead there. Uh, you know, in terms I think it's probably, you know, say maybe we are Optim, maybe I'm as, as architecture, we are optimizing for too much flexibility of, you know, not having to manage multiple entities, right? Right.
Um, I think one good thing about, you know, at the bottom most point there, volume move to any node without copying data, we do that automatically today. So volume move as a feature existed in unified on tap for almost 20, 25 years. 20, 20 15 years.
Right? But we never moved it automatically. Customers had to manage it for the first time.
We are able to do this because we know exactly where, how, how the load is gonna be perceived through nodes and be able to move well in your You're actually doing a volume rehosting. Exactly. Yeah.
Actually we have another operation called Volume Rehost, which is like reattaching across SVMs, but that's a different conversation. You're absolutely right. We are rehosting a volume pointer just so that, you know, the traffic gets routed to that, right?
So because we could do volume moves automatically, right? Without copying clones also can be provisioned automatically to redirect data right across. Right?
Um, I Think probably I'll given the time, I think I'll probably skip through that, but, uh, I think I just covered that IO path. You know how it works, right? So volume is still there.
iPath still goes through that today res will be optimized route where it can be sort from any node rights actually will definitely, uh, have to go through one node, right? But if you have a flex group, your data is across all volumes, right? Technically, so This is where CXL would make this so much better.
'cause you'd have rack scale access to the MV RAM directly, and so you could write from anywhere. Yep, yep, yep, yep. So sharding, It's not high enough to make it Wearable sharding memory layer, Uh, NV RAM layer across all the nodes and being able to reach consistency point to any node, I think Yeah.
That's, that's where you're getting it. Yes. Makes sense.
Yeah. So When we are looking at NVHA, we are considering that by the way, right? Um, we are looking at, so today for our, for us, the nodes are straightforward, right?
In some sense because NVM is a rash to that and consistency point is driven through that. Um, If and when we get to Statelessness, I think there are some other options available, Probably Skip too fast. And we talked about compatibility, right?
I think we take this seriously. So, uh, there were a lot of questions about, okay, hey, should I move my data from existing data asset to this one? Um, Yes.
I mean, there's no sugar coating there, right? You have to move the data, right? But, but reality is that, you know, hey, we changed the entire aggregate layer, the storage layer completely has changed.
It aggregated all nodes into single storage pool. And that's why we really cannot just, uh, in place upgrade existing systems into that, right? Um, and the Existing systems include other ONTAP Systems, other ONTAP unified systems, right?
Yeah. Um, when you have the next generation of a FX systems, I think yeah, you could, you could tech refresh them, you know, you can bring 'em in Node and then you can just refresh them out, right? But, uh, existing unified, Um, So realizing that could be a issue for some customers, I think.
I think we, we made sure that, um, the data mobility features that we have, customers are loved or years are supported from day one. Everything Is supported today, right? On day one.
So None of the features had to be recoded for this architecture. And that's the beauty of the layered architecture that ONTAP is, Right? Um, We just have to make sure the rate layer had to be converted to the, uh, single pool storage.
I should not oversimplify it. A lot of engineers won't like me when I say that, but I think, but, but we were able to pull that through without changing anything at the app. Uh, at the, uh, data management layer, you will see secure multi-tenancy storage virtual machines that we have supported for last 15, 20 years, which I think is some neo clouds are actually very interested in because they want, they require NCP certification and, and new clouds require secure multitenancy actually built into the storage.
So they're hosting multiple customers on the same infrastructure. It's a great feature. All the security features except aggregate encryption, uh, will have volume level encryption, FlexPro volume level, uh, with your own keys, snapshot, snappier, everything will work.
Flex group, flex wall, flex cash. Um, there was a question about do we always have to copy data from existing unified assets into uh, a FX? Uh, no, not necessarily.
You can burst with flex cash, uh, do your training, do your whatever. Right? And I, I talk to New York clouds all the time, and one of the things they say is that, you know, they always segregate architectures into somewhat of a inner ring and out ring, inner ring being foreign training out ring is where they have all the copies of data, actual data, right.
Um, and all they want is burst and do training and then leave it. Right. Plus cache is a great story for that.
Still technically moving data though, Uh, only portion of data you can, right? Right. Oh yeah.
We're still moving data. Yeah. I mean it's, to me, storage is all about managing cost performance and I'm not Complaining about it.
Yeah. Yeah, Exactly. Um, no change in the way any of the protocols operate.
Uh, volume ownership on single note, right? Um, we talked about it and none the a ps or anything like that changed. I think one of the great feedback we got from our customer base, our, our Brownfield customer base who have been longing for disaggregated storage for some applications, actually even in EDA, for example, where they care for, uh, care for, you know, for, for EDA, um, simulations, they care for less capacity, more performance.
They really loved it. And as I said that boom, like we plan for, uh, uh, early access program of six weeks, but they were done in like two weeks because all the APIs worked as is. Right?
That's the beauty of this. Um, there are a couple of features missing, and I, I, I believe I kind of me alluded to that. One is a fabric pool, which enables us to tier data to object store, right?
Uh, is being worked on, um, dry mixing, you know, I think it's, it's in progress. You know, we are, we are mixing drives and we are unified on tap. And here, uh, we have a strong journey to get to one 20 terabyte SSDs, uh, uh, in a matter of like six months, six to nine months.
Uh, there's, there's a whole lot of things happening, right? With this architecture. This architecture makes a lot of things possible independently, right?
While we improve our standing in terms of the, uh, controller roadmap over time, like highend midrange and everything, um, we can actually address the workloads where capacities, uh, uh, level capacity deployments are actually, could be a, could be a norm, you know, given that AI and they want to get all the data assets into single infrastructure, right? Haven't fabric pool missing as you we're not getting into IDE yet though, but as you start storing embeddings along with your data Yep. And then you change embeddings or update embeddings, you have to re embed everything.
Correct. To take the old versions either you have to keep 'em correct. A snapshot.
So You're have, it's gonna be a whole lot of fabric Pool. Absolutely. Absolutely.
It's been actively worked on, and you should see that in the next, uh, next six to 12 months. Um, I think I've covered all these, uh, I don't know if I have anything specific to cover. Um, Yeah.
Uh, Ray Tech, you know, we, we leverage Ray Tech, uh, for, for, For, Uh, liability of data and actually we are gonna increase the, uh, the size of that. You know, 93 may becomes 1 26, uh, three P becomes six. Right.
You know, so this allows us to actually scale to that level. Right? My point is, um, having aggregate constraint to, to HH Air, we never needed these things, but you know, now we, we are to work on that.
Yeah. So Again, moving to this aggregator on tap today is not required for customers. Right.
But it's a start. If somebody's planning for new AI infrastructure, this is what we would recommend. Right.
Or, you know, other verticals as well. Right. You know, there's not, there's nothing stopping us from deploying it.
Um, unified on type is not going anywhere. I think. We'll, our goal, like what we have done in hybrid cloud, uh, scenarios, our goal is to connect wherever your storage is.
Right. And we take that seriously. Right?
I know changing data management layer and above is not an option. And that's what we have done here, right? That's why we are able to get to this market faster.
Um, I think it's probably a year of two evolution where this will be ready for everything. And this is engineer speaking, so don't hold me to it.