Redefining Data Infrastructure for AI with NetApp
At the Tech Field Day Experience during NetApp INSIGHT 2025, Syam Nair, Chief Product Officer at NetApp, outlines their strategy to redefine data infrastructure for artificial intelligence (AI). The focus is on building intelligent storage systems that unify data across block, file, and object formats, enabling fast and secure access to AI-ready data. By integrating AI capabilities directly into the storage layer, such as metadata enrichment, tokenization, and data governance, NetApp aims to empower storage administrators and extend the usability of data for AI workflows without sacrificing control or security.
In his presentation, Syam Nair emphasized that the AI data landscape is shifting from hype to practical application, with growing unstructured data sources such as machine-generated and generative AI outputs. With the ONTAP platform at its core, NetApp’s vision hinges on making all data AI-ready by embedding intelligence—like security policies, tokenization, and embedding mechanisms—into the data layer itself. This not only ensures data accessibility and governance but also removes the need for complex extraction, transformation, and loading (ETL) processes. NetApp’s AI Data Engine (AIDE) and AFX are designed to streamline this intelligent access while reducing the proliferation of data copies by managing metadata and vectorization in place.
NetApp’s approach aims to elevate the role of storage administrators, transforming them from infrastructure caretakers into enablers of data-centric applications and AI workflows. Instead of pushing AI users to understand the storage backend, NetApp provides APIs and policy-driven data access mechanisms that integrate with tools like Kafka or database systems. Emphasis was placed on security through granular, zero-trust policies and governance over metadata to prevent overhead and sprawl. NetApp aims to support emerging standards such as Apache Iceberg for semantic access and to evolve toward a system where unstructured data can be consumed like structured data—offering semantic reads without altering write formats. Ultimately, NetApp is not attempting to replace databases but rather to unify and enrich data access directly within the intelligent storage infrastructure.
This presentation was recorded at NetApp Insight 2025 in Las Vegas on October 15, 2025. Watch the entire presentation at https://techfieldday.com/event/netappinsight25/ or visit https://netapp.com for more information.
Transcript
Thanks for giving me an opportunity to be here, and thanks to all the listeners. I'll start off by with a few things and I'll take questions. Look, we are really at a time where data is, data has been talked about as the fuel for ai, but, and it has been the fuel for all the workloads, not just ai, but now is a time where all the hype is kind of coming down and people are starting to look at how do you really get value out of the data you have?
And cloud is a reality. Every enterprise, every business is thinking about cloud in one way or the other, as well as unstructured data is growing, machine generated, LLM generated, generated ai. All of these are sources as well as what leverages these unstructured data for us at NetApp, it is the promise of this intelligent data infrastructure is realized through actually having a unified data platform, a platform for blocks file, as well as for objects built on the ontap, uh, platform itself.
Where any data that's in the ONTAP estate is immediately AI ready. What does that mean? I look at through a couple of pieces.
For AI to be successful, data needs to be accessible. It needs to be connected, it needs to be always on. It needs to be always secure.
One of the biggest challenges in the industry is data security, data protection, and data protection. Not through the lens of just backup and being able to restore, but the right data available at the right time to make the right outcome. And is it protected?
Because these days it's not just about data through the lens of PII, it's every data has context. Generative AI models can actually infer context out of data. So the excitement is being able to bring in that intelligence and the security directly to where the data resides, especially in the case of unstructured data, which is a large state that we manage for our customers.
So that's the, uh, that's the path we are on. We talked about some of the products that Jeff talked about, A FX and a IDE. Those are all products in those lines.
I'll stop by saying the vision for me is this is what we are working towards. Just imagine every piece of data is accessible as an iceberg table and, and, and everybody data scientists, they don't need to know what the infrastructure is. And storage administrators and the data providers can actually provide the data access through the tools that the data scientists or the application developers can, uh, need to use.
This is actually much closer to reality than what we think about because we actually have all the technology now and vendors like us, NetApp has actually built in those intelligence into the layer. So that's the vision. That's what we are embarking on.
Uh, quite a few products that we released recently. Happy to take questions or providing more input. Is NetApp trying to almost like redefine what storage is?
'cause I, with the, with the A IDE, you are now doing tokenization and embedding in the storage layer. Are you trying to redefine some of these AI functions as in the storage realm? And what if you succeed in that?
What would be the ramifications if you are in the world of storage, if you're a storage administrator, um, or you're an infrastructure person from an enablement perspective and a learning perspective, how do you think if you succeed in that, how is, how are things gonna change? Yeah. Uh, two things.
One, yes, our vision is to redefine. It's just not about storage. It's about making that storage, making that infrastructure intelligent.
And the second part of it is we really want to provide, empower the administrators, empower the storage professionals to be able to cater to the business needs without having to worry about all the complexities. I'll take an example. Today, if you really want to get value out of data that is in an unstructured, this, many of the storage administrators need to actually open up their anti storage ecosystem for ETLs and pipelines and harmonizations and model.
I mean, I come from a world where we, I have done this multiple times. It's a, it's, it's, you are now letting go of, in a way all the governance and controls you have had on the data is no longer with you because it's moving through the system into multiple layers. Be it a lakehouse, you know, people talk about bronze, silver, gold layer or a delta lakehouse or an iceberg format.
And then, or be it a, an entity modeled database or a graph database. These are all the same data is moved into multiple places where you're losing it. So our vision is by bringing in the intelligence to the storage, we will now empower the storage administrators, the people who really care about this data, who manage the capacity, who manage the performance, to go beyond managing capacity and performance to actually provide these services to the, uh, to the application developers and data scientists.
As an example. The vision would be realized where the data s like data scientists and the application developers may not know about NetApp. They just, they have data.
They just want access to the data. The NetApp storage administrators now have the power to say, look, you don't need to worry about what the infrastructure is. Here's a set of APIs that you use, that you use, which of our tools, you're using Kafka using a different database.
It can actually talk to NetApp directly and be able to access the data. Policies can be set on the NetApp storage site that can actually travel with the data. That's the vision we are looking at.
So in, in short, yes, definitely want to redefine storage as not just capacity and performance of storing the data. It's more about intelligence for the data. And my follow up to that is does this require NetApp to now have conversations with people at enterprises that they had never really had to in the past?
You're not, you know, 'cause storage administrators wouldn't be the people you'd be selling to here. Yeah. Yeah.
I would actually say it's, I, I would, I really want to think about this as making storage administrators and our current users having that supercharge superpower to do it. But to your question directly, we need more people within the organizations to now be influencers. Mm-hmm.
To act like I really want data scientists or, um, uh, chief security officers to be influencers in saying, look, here are the outcomes I want to drive out of the data. And the NetApp administrator, the storage administrator saying is, I can deliver that because I have a platform that is built in ready for it. So I I I, I'm turning it around into, I don't want to go and sell to somebody else, but I want them to be influencers in terms of what the jobs to be done is so that the NetApp storage administrators and our users have been with us, our partners have been with us being able to say, Hey, yeah, I have the platform, I have the tools.
I can actually make that happen. Hmm. I've heard you mention a couple of times security.
Um, how are we defining how, who gets access to what, 'cause obviously in most enterprises, not everybody has access to everything. Um, so is there a way we can say, yes, you've got access to the system, but you can only see this part of the data and and at what layer do you envision, uh, somebody doing that and and who's controlling that? Yeah, So I look at this as the, the, the storage administrator or the storage professional being the provider where they can define, not, I, I don't, it's not opening up access to the system, to everybody else.
It's opening up access to what is the relevant data that workload or that application requires. So I would look at much more granular, uh, policies, granular security policies that are actually ingrained in the system applications, being able to refer to that. Think about this in the context of I have a CRM application that is actually leveraging some data.
Now, if I know the context of the, what the CRMs schema is and if there's a semantic model defining what the output, what what is stored in the NetApp platform, the administrator should be able to say the CRM application has policy access to these kind of data defined by our metadata model. So it is more in line with the zero trust protocols where you can define through an infrastructure as code kind of mechanism, clearly define the policies that gets access to the data without knowing how the data is stored, where the data is stored, or you know, what kind of volumes is what, where the system is. I, I had some confusion when I was listening to a couple of different presentations about the data engine, the AI data engine.
Um, because we got access, we were able to do a, like a little demo, a can demo of how to see things. It was interesting. Right.
It's interesting. Um, but I have a couple of questions because I, I think like the presenter said today in the keynotes, one thing that does happen is, um, in this data pi, this AI pipeline throughout all the stages is there's copies of data everywhere. And this is all without the extra copies.
You don't know about being made, being made, right? So you've got all these extra copies, you've got a lot of copies that are required, but I kind of got from something I heard yesterday and the, the, the lab we did that, it's only one copy of data all the time. So, uh, could you explain a little bit about how you envision using, um, the AI data engine to go through a, a real generic type of AI pipeline?
And is it copying data or are you just restricting access to the same data for everybody? Yeah. And then talk a little bit about what type of, um, uh, management cost there is to that in as far as speed goes.
Okay, great question. I think the confusion could be coming from, because we tend to explain the complexity today of, uh, data moves from unstructured to the lake houses to some kind of structured format, the vision and what we have actually built as part of A IDE, the data engine is your data engine sits right where the storage layer is. So as you are writing data, it's actually form, it is building the metadata layer.
So it is informing what the data is by building the metadata layer on, on another service is actually creating the vectorization data. Data. Another is actually creating the guardrail based on what the policies are.
So in this case, we are not copying anything other than there is one metadata engine that actually has metadata about this particular data that is being returned into. Now, in the context of AI using it, our vision is we already have technologies like Flex Crash and, uh, snap Mirror that actually caches data as required, et cetera. So when AI is using it, this is where we talk about, uh, you can keep the data where it is, you can bring AI to you, or you can keep your compute where it is, and you can take the data to there.
It's not a separate copy. It's using the same technology underlying that is being used in, uh, ONTAP today, like flex C. Okay.
That, that makes a lot more sense. And are, are you seeing, I mean, I think this is such an interesting word. I I am a product marketer, so um, is such an interesting world right now because for a while we've had so much hype about AI and now all of a sudden we are starting to see standards emerging.
We're starting to see, you know, this actually looks, this is what a pipe AI pipeline looks like. So we kind of have this repeatable process that we'll start, we'll start to see, I think, um, the, the market adopt different things. So how, what do you think NetApp stands in that?
Because you guys have like a really awesome opportunity based on ONTAP and all the things that it can do and how you've enhanced it. Yeah. Something that we are working on, and, you know, no, no product announcement as such, but something that we are working on is this concept of every piece of data being, having a semantic model, even if we know this is for specific industries and administrators actually defining based on those industries, building industry models and next step of it, being able to provide an MCP server right on ontap, this is something that we are working on.
That way it actually, the data is not just AI ready. You can access through the protocols that generative AI applications are using. Look, where we wanna differentiate is the following.
And I, like, I don't wanna make a claim that we are the new database. We are not trying to be a database database as its own role. You know, relational databases have its own role.
No SQL database have its own role. We want to be like the database for structured unstructured data, but, and a unified platform. If it is blocked files and objects, we are able to provide that capability.
It's more about making sure that we can add the storage layer, bring meaning to the data that is being stored, which today is a complex task other people have to do. And storage administrators and data infrastructure professionals now actually have the power to act, Hey, look, this data, this is the data, this is the relevance, this is the schema for it. Here, you can leverage it.
That's the vision. That's the first phase of that is our A FX from a disaggregated, which is as a scale standpoint. And a ID that brings in that intelligence there.
Have you, if you, I know it's, I, I know you probably got this in like public review or private review. Have there any been any comments on any type of management task or overhead, uh, tax, I mean, over on the overhead that it might be for getting some of these tasks done? I don't know if we have, that's a very, very fair question.
I think if I understand the question clearly, because now you're adding an a ID engine, is there an additional overhead in terms of performance? Exactly. Uh, I don't have specific numbers.
We do actually have, uh, really large customers are using this platform and given us a lot of feedback. I haven't gotten any feedback in the context of any performance. Look, I would, everything will have some amount of impact.
You put a, you put a new piece of code 11 impact, but nothing that, uh, I would, I, I can, I can tell you as, uh, as something that is impacting the workload. Now, not every system may need this, a ID there are still systems that will just purely track NFS protocol, want the, so those workloads will continue to work. So we, our, our goal is not to replace the entire storage estate with something like that.
We are adding this capability available for AI workloads. The other workloads can continue to work with, uh, on tap the same way it works today without any performance impact. I just take a different side to that from a performance impact.
But when we are building this metadata and these new catalogs, we've gotta be careful about versioning, provenance, and all those things. So over time, the overhead becomes that your metadata layer is larger than your data layer. So have there been thoughts about that?
Yeah, so I'm hoping there is, oh, go ahead. No, no, No. I'm joking.
I'm hoping there is, uh, one of our, or somebody is actually having a conversation today. You could ask that question. You'll, it's been thought through.
Uh, it, it's like governance of the metadata because metadata becomes yet another data to manage, right? The governance of that, that's been thought through, that's part of the system. Somebody can provide you much more details about It.
Great. Thank you. I have one, uh, question, which is, I'm swiveling around this, this idea of, uh, of, you know, a variety of right formats, but semantic read.
And I'm, I'm wondering if, if you're, if you're heading in that direction, most of this semantic access idea seems to be centered on the AI strategy, but, you know, I feel like semantic read is, is it's a little bit of a holy grail. Um, and I'm wondering if you're trying to build towards that in a more general sense. Yeah, so the existing protocol support doesn't change, right?
How our applications are using it. We haven't standardized one, but I think there are standards emerging in terms of, especially for ai, like, as I said, Tableau format from Iceberg is a standard that almost everybody's using now, right? So our, so our goal is to support standardized formats out there, and that way we are not creating a closed ecosystem from that sematic reach standpoint, not to create yet another one, but if Iceberg is the emerging standard, we would just follow that.
That's a great answer. Thanks.