09×05: The Role of Data Infrastructure in Enterprise AI with Ingo Fuchs of NetApp
As customers try to figure out how to present data to Agentic AI applications, many of them are realizing that it’s time for the storage infrastructure team to step up and take a seat at the table. In this episode of Utilizing Tech, recorded live at NetApp Insight in Las Vegas, hosts Stephen Foskett and Guy Currier from The Futurum Group sit down with Ingo Fuchs, Chief Technologist for AI at NetApp, to explore the critical role of data infrastructure in supporting enterprise AI and agentic AI applications. As organizations move AI workloads into production, traditional infrastructures—especially storage teams—must take a more active role in enabling performance, efficiency, and governance. Ingo emphasizes the emerging needs for data quality, control, compliance, and currency, particularly as AI agents begin making decisions and interacting with sensitive enterprise data. The conversation highlights how NetApp’s capabilities, such as AI Data Engine and native infrastructure integrations, enable real-time data pipeline management, enforce guardrails, and ensure consistent and secure data delivery. This shift represents a transformative intersection of storage, infrastructure, and AI operations, paving the way for scalable and reliable enterprise AI solutions.’
Transcript
As customers try to figure out how to present data to Ag agentic AI applications, many of them are realizing that it's time for the storage, uh, team and the storage infrastructure to step up and take a seat at the table. That's the topic of this conversation here at NetApp Insight on this episode of Utilizing Tech. Welcome to Utilizing Tech, the podcast about emerging technology from Tech Field Day.
Part of the Futurum Group. This season focuses on practical applications for ag, agentic ai, and other innovations in artificial intelligence. I'm your host, Steven Foskett, organizer of AI Field Day and other Tech Field Day events.
And joining me this week as my co-host here, live in person at NetApp Insight is Mr. Guy Courier. Welcome Guy.
Thanks, Steven. It's good to be here. It's good to be at NetApp Insight.
It's good to be here on the pod again. Uh, talking about agent ai, I think the best thing about NetApp Insight is that this is a real practitioner oriented kind of crowd. I always say that when I go to an industry event, my favorite thing is talking to the customers that are there and asking them for sort of a reality check.
What are you doing? You know, what do you think of this technology? What do you think of these announcements?
You know, will you kind of get on board with the direction the industry is is facing? Yeah, I think that's, that's true. I think every conference, um, the highlight is to be able to speak to practitioners, also to strategists and architects, uh, who are a different form of practitioner, um, and to see and hear their reactions.
Um, usually frankly, at a vendor conference, pretty enthusiastic about what the vendor has to offer, but then just getting a little deeper and, and seeing, um, how it may apply to whatever it is they're trying to do now across industries in certain specific scenarios, um, where the rubber hits the proverbial road, uh, because that enthusiasm needs to translate into something valuable, valuable into value, into areas where they can innovate in their particular industries. Maybe not have to worry about the tech or maybe geek out into the tech in order to get that value out of it. And getting all that color on it, I think, um, is really helpful for us as we absorb, um, all the different stories, messages, products, and stuff out there.
Yeah. And, and when it comes to agentic ai, you know, many companies are just starting on their AI journey, and so it's hard to know how long, how far along we are, what roadblocks they've found so far, uh, what they've found, what works, what they found that doesn't work. And so this week on the podcast, we've invited on to sort of speak to that, somebody who focuses on this for NetApp.
Uh, so Ingo Fuchs, uh, welcome to the show. Thank You very much. Uh, thanks for having me here today.
Well, my name is Fu as you said. I'm the chief technologist for AI here at NetApp. And so I spend a lot of time talking to customers and partners and analysts and people like yourselves, and just really trying to help our customers cut through all the noise and all the hype and figure out how they can really derive value out of AI and what the right technology bits and pieces are to make that happen.
I wanna start, though, with one thing that I just thought of as you were, uh, as you were talking about this moment ago, Steven, which is a lot of people are just starting out on their AI journey. Some are starting out on their agent AI journey before they've even really gotten very far on their AI journey. The general sense I get is that, um, there's so many things that are different or let's say more extreme about this particular technology revolution.
And one of them I think is that, uh, customers, adopters, users are just all over the place. All of there's more variety right now. There's some who have barely touched it other than of course, employees using chat GPT or what have you.
There are others who are really well advanced. There are many who've been doing AI for years and years. Um, are you getting that same sense from the customers that you, that you work with, that it's just such a wide variety of scenarios?
Yeah, I think I, I, I agree with your observation. I think there are really two things that stand out to me. One is that we see a very, very large trend towards productizing AI actually moving AI workloads into production.
So you're absolutely right that people are in very, very different places in terms of experimenting with ai, having small production workloads, even. Um, you know, some companies very successfully use generative AI to, you know, produce videos and images and blog posts, and maybe making their email processes more effective. But the opportunity with agent AI and building AI factories, that is really where the rubber hits the road, as you said earlier, that is where money can be made, and that is where your competitiveness really comes to fruition.
That's where you outcompete your competition. And so with that move to production come new requirements, and I'm sure we'll go deeper into that, but the other trend that I also wanna talk about here as we are at this conference with storage and infrastructure practitioners, is that the storage team has a massive role to play as companies are looking into deriving value from ai deriving value from their data and moving AI workloads into production. I was just literally, half an hour ago, I was meeting with a customer, and it was a storage team and the infrastructure architecture team, and they started the conversation with, well, we don't care about AI because the I team told us that we don't have a role to play.
And I was like, well, hold on a second. Let me kind of explain to you what the role of a storage and infrastructure team is for production AI workloads, and how much you can help your AI group and your AI practitioners and your line of business to really be way more efficient and make a much, much bigger impact. 'cause there's so much that a storage person and infrastructure person just knows, and so many techniques that they're very familiar with that are just routine for them, they don't even recognize what it can do for production ai.
Yeah, and I think it's interesting that the storage team has that idea about themselves. I I think that, you know, in many cases, uh, storage has been sort of marginalized and, um, you know, they're, they're, they're just over there. They care about like the deepest part of commodity infrastructure, commodity almost, you know, it's, yeah, we don't need to involve those, but to be honest, um, I think that a company that's not engaging their storage team and their storage products in their journey toward AI is really gonna miss out because they're, it's gonna be so inefficient.
Whatever they come up with is, is gonna be expensive and challenging, and, you know, there's gonna be, you know, potential problems where they could head that off by just getting the storage team involved. Yeah. Can we just, can we just break outta this idea?
It just seems ancient at this point. The demands that, that the repositories, the, the resource pools, the applications, the u the demands being put on, on, on, you know, the entire infrastructure are extreme in comparison to, you know, even five years ago. So this idea that there's any corner of that infrastructure that does, that does not have something to contribute to, especially to the latest AI revolution, uh, or the latest revolution of ai.
I mean, that's, I I think we should just put that to rest. There is nobody we talk to in any part of it who can't, um, apply what it is they're doing day to day. And maybe you've been doing day to day for 10 years.
Yeah, they can't, that can't contribute. No. Yeah, no, I, I agree.
I think what we are seeing is that AI is becoming just another enterprise application in a lot of ways. Now, every enterprise application has unique requirements. They need certain things, they need certain functionalities and architectures that are suited for them.
But AI is, especially agenda AI is becoming an enterprise workload. And so from that perspective, it needs to be able to provide value when it comes to that. So creating like an outlier, some shadow ai, shadow IT kind of organization with all their own stuff for an experiment that often works really well, that can be in the cloud, that can be on premises.
But if you're moving into production, it's a little bit like when public cloud first came out, you know, over a decade ago, and people that were early adopters were mostly concerned with, does it work? And then they found, oh, it does work. That's great.
At some point in time, they were like, Ooh, wait, you're not doing any backups of that data. You're not replicating that to another availability zone. You're not replicating to another region.
You're not, um, doing any compliance checks. You're not Relearning all the lessons already. No, Dr.
Right. And so, and all the IT folks were like, duh, you know, this, these are very, very basic things that we have done for decades. We know how to do that.
And then when you apply that to ai, think about vector DB bloat as just as one example, right? So when you're looking at a typical AI workload, you have all of your unstructured data could be petabytes of unstructured data, you need to put structure on top of that, which is vectorization, right? So you're creating these embeddings, and you store that in a database.
A typical vector DB very often will be about 10 times the size of your original data set. So if you have a petabyte of unstructured data, and then you end up with a 10 petabyte vector database, well, but if you do this right with someone like NetApp, you don't need that 10 x factor because we don't need to store the same copy over and over and over and over again. And guess what?
We have done this for decades. We can create clones, we have done snapshots, we have done mirroring, we have done all of these things for decades. So it's really, really simple to really apply these principles to these modern workloads.
And I'm not even talking about compliance and data guardrails and privacy and all of that other stuff that also needs to be considered. Yeah. And, and I, I think that we should zoom in on that a little bit though because, um, to me, the hallmark of an enterprise application, whether it's an AI application or AI agents or just a conventional application, is the data.
It is the enterprise's data. Um, you know, ultimately that is what makes it an enterprise application. And that has so many challenges.
So I, I would say that, you know, generative AI in the form of chatbots or image generation or something like that, it, it's not really inherently an enterprise application. Now, I can see that it might be helpful to have a chat bot to help, um, you know, answer questions based on customer data or something like that. Optimization.
I could see that that might be helpful. But AgTech having AI agents that are assisting your business and doing tasks on, on behalf of the company and the employees, now that's a whole different ballgame, and that's why I get a lot more excited about where we're going with this. But, but, but if, if, if you need data, then that opens up a whole can of worms in terms of figuring out which data that is, uh, you know, like you said, optimizing it, configuring it so that the, uh, AI application can access it and then controlling it.
Because that, I think is really the biggest challenge. Another, another aspect to that. Now, I agree with everything that you said, but another aspect to add to that, if you're thinking about a gentech AI is agents can make decisions.
Mm-hmm. That's one of the key differences for the Gentech ai. Um, and so if an agent is making decisions based on outdated or incorrect data, that can be very, very bad.
It's gonna be the wrong decision, right? And so even if he, if he roll it back to this very simple example of a chat bot, even if you just say like, oh, it's not even an agent, it's just a chat bot. If you have a customer service chat bot that is providing, uh, feedback to a customer, and sometimes these are, they're even pretending to be human agents, right?
So, but okay, so you have a customer service chat bot that is using a customer service policy from two years ago, not gonna lead to that is outdated by now, is not going to be useful if, uh, you are a hospital network and the cust patient has just come in, the blood test was taken, the blood test was updated in the system, but your AI data pipeline hasn't been updated with the latest data yet, you're not going to get good outcomes. So data currency and accuracy is very important. And so with data currency, that is a benefit for a company like NetApp, because the data sits on our systems, we can detect that something has changed and we can propagate that change with the correct guardrails through the ID pipeline, making sure that the customer experience at the end of that pipeline is the right one.
The correct, get the right information with the correct and consistent global guardrails, for example, or other correct factors. Absolutely, yes. What, what are, I wanna follow up a little bit on, on, uh, later on this, you know, um, question of, uh, currency, but what are the different factors can you sort of lay out based on what the work you've been doing?
Uh, guardrails are one factor. Um, I guess currencies is another factor. What are, what are these different factors?
'cause I think it would be really helpful for our audience to see, understand the scope of that. I think we all tend to see our part of the elephant. We don't see the whole thing.
Yeah, absolutely. You know, so from my perspective, it always starts with, like you said, right, data fuels ai, right? So if you have poor data, you're going to get poor Outcomes, those quality, And so, yeah.
And so, um, so start with that data and having the right data. And so that starts with finding the right data, identifying the right data. So that means you need to have, uh, a view into a unified data model where you can find all your data, whether it's on premises in the cloud, and that includes new clouds and sovereign clouds.
So first of all, the question is, can you find the right data that you want to feed into the system? The second question is then can you prevent data that you don't want in the i data pipeline at the beginning of that pipeline? So if you have something that credit card numbers or security numbers, um, you know, medical conditions, you know, something that you don't want the agent to disclose at the end of that pipeline, the best place to stop that information is at the beginning of that pipeline.
So that's where you need data classification and data curation, because Ultimately nothing, no guardrails that have been discovered yet can actually put a guardrail around AI that will keep it a hundred percent of the time from doing something unpredictable. It's a, it's a modern arms race. Yeah, it is.
And one, and then there'll be other. So the only Way to keep from leaking is to keep that data from going in in the first place. Yes, exactly.
So, and, and this way you build like multiple layers. So I'm not saying that this is the only layer of protection that you want. You have multiple layers, same with like ransomware protection, for example, right?
You have multiple layers to protect your core value, your data from being attacked. So you do the same thing with guardrails in the AI world. It's where you need to protect your data.
So you have that data that you feed into the pipeline. You need to think about efficiency, you need to think about reliability. If you're using AI in production, if AI is controlling your manufacturing workflow, you don't wanna shut down the factory for a few days because you have a mistake in your data, right?
So you gotta have that reliability, you gotta have the availability, all those standard things. But then I also want to talk about security, um, because the more value that is in your data, and as you're using ai, your data becomes more valuable. The more important it's to protect your data.
And that is ransomware attacks, that's exfiltration attacks, that's, you know, worrying about things like, uh, post quantum cryptography, which is maybe like a whole nother topic for a different day. But thinking about if somebody comes in today and steals data that is very, very valuable to you, it's probably also very valuable to them, like your customer database. So yeah, maybe they can't decrypt it today, but with quantum computing, maybe they can decrypt it in a few years, and it's probably still valuable to them.
Even if your customer database is two or three years old, your competitors probably still going to find value in that. So security is another really important aspect in this. So a lot of these requirements sound very familiar, don't they?
But even just core, uh, infrastructure considerations, like multi-tenancy suddenly become important. So now think about you have agent ai where you might have thousands of teams of agents that are all over your shared infrastructure with all of your other mission critical applications, all operating in the same environment. With secure multi-tenancy and quality of service enforcement, you can make sure that your mission critical applications perform as expected.
Um, whereas, you know, if you don't have that, you might have some rogue agents that are just gonna grab all of the CPUs and all the network and everything else that they can grab. You wanna prevent that. So you really gotta apply your IT standards, principles and guardrails on, on all of these dimensions for AI workloads, they are your next enterprise workload, and it needs to be treated the same way.
Let me, let me, uh, follow up now on, uh, the concurrency and adjacency part of this data adjacency. Um, I think one of the things that we, we learned, uh, over the course of this season of utilizing tech is the ways in which, uh, you can think of an agenda AI as, um, um, you know, a, let's call it a software replacement. Um, in other words, uh, what you might achieve with a piece of desktop software or a piece of enterprise software or, or, uh, a piece of SaaS, piece of SaaS, uh, a SaaS interface, um, you now can achieve with agentic ai, whether it's scheduled, whether it's responsive.
Um, I'm thinking also the healthcare example that you gave here. Um, data concurrency and adjacency is, uh, really critical, uh, for agentic AI in particular. Because if, if you start to think of it as this is a different way for me to, you know, tabulate or write or be productive in my job or follow a workflow, instead of using a software interface, I'm using an AI agent, then really the matter with which you tend to be working is proximal adjacent local and concurrent data.
The healthcare example really puts it into light, because I don't want to be getting recommendations based on blood tests that are three months old. I want it to be on the blood test that I just took. But if I'm just a knowledge worker, I am making decisions or doing things in the software based on what I'm working on right now, not based on what some AI was trained on over the course of six months of prior data.
Yeah. So, um, how, how do you help to, um, ensure that and solve for that when you're thinking about possibly terabytes of training data, but yet, you know, don't, don't just say rag at me, like we're, we're talking about something that may have, uh, especially with AgTech, multiple input points, multiple types of requirements and, and qualifications And following onto that, I think that one of the interesting things about AI agents is that you, they need to have, um, consistent data from agent to agent to agent, but they don't often necessarily need exactly the same data. And so in many ways, you need to be able to organize and control and present a cons, like a time consistent interface to this agent, and then this other agent and this other one.
So think about like, some, kind of like a, a sales workflow or something like that. Um, you know, if you, if you close the deal or sell one of the widgets in between the first agent and the last agent accessing that data set, then you know, they might be presented with different sets of data, and that may be what you want or it may not be what you want. Yes.
And so you need to be able to have that kind of control. It's like a multi-tenancy almost. Yeah.
And it, it leads to security considerations and privacy considerations and, you know, do you, what can you disclose? What do you want to disclose at particular stages of a process, for example? And so, um, when you're, when you're curating your data sets, uh, as part of your, you know, your AI data pipeline process, and that's something like NetApp data engine for example, for it does for you, is that you can define these guardrails.
You can define, uh, what that curation should look like and what should be presented for, for separate workflows. But I wanted to expand on your point a little bit beyond, uh, what you said in also thinking about today we kind of have this duality of data's either structured or unstructured, or we're using file protocols for unstructured data and block protocols for structured data that we have been following this system for a very long time. I believe that it's fundamentally changing now.
So we now looking at unstructured data saying, well, it's unstructured data, but to make sense of unstructured data, we actually gotta put some kind of structure on top of that. We gotta assign some numerical values and store them in database. So we're putting structure on top of unstructured data, which is this whole process around vectorization and embeddings and all that kind of stuff that I mentioned earlier.
When we do that, what it leads to is that we can have semantic interactions with your infrastructure. You can have a communication with your data, you can talk to your data using agent ai using protocols like MCP. So that is gonna be a whole new world of interfacing.
So you're not using file protocol, you're not using a block protocol, you're using something like MCP, you're using Angen AI protocol for your agents or chat bots to talk to your data. Well, This is core structure. Funny, this is core to agent AI design.
Is this, is this, uh, sort of malleability or adaptability of connection of, of, it's, it's, uh, it's, I don't know, you know, it's prompt engineering, uh, squared. Yeah. Yeah.
I, I guess so I think it goes beyond that, right? It gets probably like a fundamental redefinition of how you interact with your infrastructure and what your infrastructure needs to do. You know, in the past, would you have expected a storage company to understand the context and the, the actually understand the data that's stored in these systems?
Typically not. You would look at some higher level, you know, data company to do that at the end of a process. This is really interesting 'cause it's far, far away.
Now you're building that into the infrastructure itself. You still might have some, um, you know, most companies will still have some broader, you know, data aware application that does specific things for their industry or for the use case. So there's gonna be an other layer of abstraction on top of that.
But that layer can now semantically interface with the infrastructure and get data metadata and context directly out of the infrastructure instead of trying to retrofit it after the fact. That's way more efficient, way more secure. 'cause you can apply guardrails all the way through the entirety of the process.
You can prevent data that you don't want to be seen ever from even getting into the pipeline. All of these things. It's, it requires a fundamental rethinking of how you feel about data, how to process data, and what you want to get out of your infrastructure that sits at the heart of your business.
I was gonna say that this is really interesting 'cause this kind of melting and merging of functionality up and down the stack we're seeing quite generally, but it's really evident in, in, uh, the two of the big announcements at, uh, NetApp Insight, um, which were, uh, like a FX and, um, and the AI data engine, which are taking on at what I think I'm sort of inaccurately calling the device level a great deal of the functionality, um, needed for these, you know, ML and AI systems. Yeah, absolutely. And That's, that's in the device.
It's in the storage, in the storage tier stuff that, you know, was even above the abstraction layer as recently as a year ago. Yeah. And, and as you saw in our sessions, right?
So we, we talk a lot with partners of ours like Informatica or Domino Data Labs, you know, obviously the hyperscalers for their own AI solutions where we have these unique integrations through our first party services. Um, so we are not displacing those tools at all. We are making these tools better by making sure that they get the right data at the right time, at the right quality.
The data is current, it's not outdated. So all of these tools become so much more powerful and so much more effective in the work that they do because they get better data from NetApp from that customer. It's the customer's data.
It's always the customer's data. But we built the infrastructure to make these tools more impactful, more efficient, uh, to drive better outcomes for the customers. And that's, at the end of the day, that's, that's all that matters, is that our customers succeed.
It's really an interesting topic and it's something that we've learned a lot about here at NetApp Insight this week. And as guy points out this whole, uh, season of utilizing tech, uh, we're gonna be continuing this conversation on ai. Uh, we're actually relaunching the original, uh, utilizing tech seasons as utilizing ai.
We're gonna continue utilizing AI as an ongoing weekly podcast and that you can find in your favorite podcast application as utilizing ai, utilizing tech will be back with other topics in it as a seasonal serial podcast in the future as well. So before we go though, uh, let's kind of wrap this one up and talk to a little bit about, uh, where we can continue this conversation and where folks can learn more. So Ingo, uh, where can people find you?
Yeah, the best way to get in touch with me is through LinkedIn. Uh, quite frankly, I'm presenting a lot of conferences, um, and for our customers, go through your, uh, sales team And, uh, where can people learn more about, uh, some of these NetApp capabilities that you mentioned? com and our YouTube channel.
We have some amazing videos out there. Uh, we have demos, uh, for those of you that want to go really, really deep into the technology, we provide hands-on labs that are available, um, where you can actually explore production environment systems that are filled with data. We can experience the entirety of the data pipeline from beginning to end.
And I'll also point out that, uh, here at NetApp Insight, we recorded, uh, a number of tech Field day sessions, which are deep dives into these products. And those are also available on the Tech Field Day YouTube channel, as well as the Tech Field Day website guy. com.
My writings are there. social. And as for me, you'll find me here at AI Field Day, tech Field Day and, and so on.
Along with of course, the weekly Textron Gang podcast. Uh, you know, the Tech Field Day podcast, the rundown, all sorts of different ones that we are producing here, uh, for the Futurum Group. So if you enjoyed this conversation, uh, you can find more episodes, just go to your favorite podcast application and search for utilizing tech.
Uh, this podcast was brought to you by Tech Field Day, which is part of the Future Room Group. com. Or you can find us on X Twitter, blue sky mastodon at utilizing Tech.
Thanks for listening, and we will catch you next week.