Enabling Data using Nutanix Unified Storage
In their presentation at AI Field Day 7, Nutanix, represented by Distinguished Engineer Manoj Naix and SVP Vishal Sinha, laid out their vision for supporting AI workflows using Nutanix Unified Storage. They began by walking through a typical AI pipeline—from data ingestion and cleansing to model fine-tuning, inferencing, and archiving—emphasizing the challenges presented by data fragmentation across edge, cloud, and various formats. The Nutanix platform addresses these challenges by helping platform teams build and manage AI-ready data pipelines, ensuring clean, high-performance data is readily available for training large language models (LLMs) and inferencing operations.
The Nutanix solution includes a wide array of components that form a full-stack enterprise AI platform. Key pieces include Nutanix Cloud Infrastructure (NCI), Nutanix Unified Storage, the Nutanix Database Service (NDB), Nutanix Kubernetes Engine, and management layers such as Nutanix Central and Cloud Manager. Particularly noteworthy is their database management layer, which supports vector databases like PGVector and Milvus, enabling customers to manage LLM applications efficiently. The storage capabilities are tailored for each phase of the AI lifecycle—high-performance for training, rich metadata tiering for archiving, and versioning support via object storage and snapshotting.
The platform is designed to operate seamlessly across hybrid multicloud environments, offering data security, cyber resilience, and robust data mobility through features like global namespaces, immutable snapshots, and tiering based on metadata sensitivity. Nutanix Unified Storage supports multiple protocols (NFS, SMB, S3, iSCSI) and enables app and data colocation to optimize performance. The platform also supports cascading disaster recovery setups—including metro, asynchronous, and near-synchronous replication—to meet global compliance standards. With use cases ranging from edge AI inferencing to deep archival storage, Nutanix’s mature and versatile platform is built to handle the full lifecycle of data-intensive, AI-powered applications.
Recorded live in Santa Clara, California on October 30, 2025 as part of AI Field Day 7. Watch the entire presentation at https://techfieldday.com/appearance/nutanix-presents-at-ai-field-day-7/ or visit https://TechFieldDay.com/event/aifd7/ or https://www.Nutanix.com/enterprise-ai/ for more information.
Transcript
I will talk more about our data journey, like and how we are enabling that using our Nutanix Unified Storage. But before I get deeper into it, I wanted to talk about what a typical AI workflow looks like. If you look at, it all starts with collecting the right data.
Like, and data today is very fragmented. It is in many locations. It could be at the edge or cloud, it could be in different applications, databases, bias objects.
So it becomes important to be able to ingest data from all these different sources and make it available in one location. Now, once you have the data, making that data clean that can be fed to the L LMS becomes important. Which means that you, we need to have the capability to do the transformation of this data ddu of this data, all the cleansing work that needs to be done.
And once you have done that, then this data can be fed into model tuning. Like if you are looking to do the fine tuning of your models, you can, you need to use that data to fine tune it. Again, this can be done in the poor data center in the cloud.
And, and with that finetune model, you use it for inferencing purposes. Like now you may have a rag based model that you're using for inferencing using this FINETUNE model with prompt engineering and with the RAG pipeline, you can make this data usable for reasoning by the LLMs. And for model explainability and other use cases, you may want to checkpoint this data and keep it so that in future, if you want to know how we got to the, the model weights and everything, that data is accessible and you can look at it.
So this is a typical data pipeline and it's challenging to get all of these together on one platform. So if you look at what a platform could look like, if we had to bring all of this thing together, the platform should be able to ingest data from all these different data sources. It should allow some distributed processing engine to be able to transform this data, right?
Then provide ability to train or fine tune models and then pro make it available for inferencing and archive. And all these stages requires different storage characteristics, right? What for training, fine tuning a very high performance storage with low latency becomes important.
The ability to do CHECKPOINTING becomes important for inferencing, supporting vector databases and, and typical any database management becomes important for archiving. Having like a cheap and deep storage to archive the data. And then also providing, yes, Ese Silverton Consulting.
Do you guys actually support a Nutanix Vector database or you are providing the backend storage for other databases? Yeah, so for us, we are supporting Mil or PG Vector. Like, so we, we are providing the management of this database on top of Nutanix.
So we have the Offering right? Management. Yeah.
So, and I will get into a little more detail in the next slide, which is our Nutanix database solution that we have. NDB, that's the one which probably provides the database management pieces. So if you're using PG Vector for example, like we have the support for the database management piece, but the code engine is not something that we built.
Like that's something which customers will bring one of the vector databases that they have. So from a value of the management, help me understand, 'cause most developers consume these things directly. What is the layer, this management layer adding on top of the capabilities of the, So think of you know, the upgrades, the all the database, typical operational piece like you're doing the upgrades, the copy data management piece of it, that's the offering that we have for all databases.
Like, so today, primarily SQL server, Postgres, these are the MongoDB, these are the common deployments that we support there. Now PG Vector is the new one that we added to support this. Before I ask another question, are we gonna get to like a demo mode?
We're gonna see, see this or the, the question that we have then is kind of target audience for this. Are we looking at platform teams building, uh, golden pipelines or are we looking at AI developers consuming this directly? This is the platform team building the pipeline and making it available.
That's the focus for Us. Thank you. Right?
But that's where like, now if we have to support all of this on a platform or turnkey AI enterprise AI platform, the platform needs to provide the data mobility, data curation. You need the scale and performance to support all these different storage requirements. Security and governance becomes important and the full data lifecycle management, like this is what we envision that a platform needs to support.
And, and this is a stack that we, we look at to support all these. So you have a hybrid cloud infrastructure which Provides before you go back, uh, yes. Is data versioning part of the solution here?
I mean obviously there's some capabilities and, and object storage for that sort of thing, but file does not have it and obviously block doesn't have it. Yeah, so we are relying on the snapshotting for that. Okay.
For the snapshot for, uh, on the files and block side and then the versioning on the object side for it. Alright. Yeah.
So this is a turnkey data pla, enterprise AI platform that we are looking at. You have a compute infrastructure which runs either on-prem edge or cloud to Provides a Kubernetes platform to run your distributed compute processing engine for data cleaning and for running your applications. You have an agent AI engine, which helps build the multi-agent workloads they need storage, which is a unified storage offering, and then vector database and database management.
So having the ability to do that as well. And then all of this tied together through a very comprehensive management operations and governance layer. Like this is the stack that we envision and that is what we provide today to our customers, to our Nutanix Cloud infrastructure Provides a very comprehensive virtualization layer with our EHV hypervisor.
It has flow networking like which provides microsegmentation and virtual networking. We support bare metal and hyperscalers. Where the NCI can run this unified storage is the Nutanix Unified storage with data lens and which is the main focus for today talk and we'll d deep dive into that one.
We have n AI offering, which is the agent AI framework offering. That was the focus for the last presentation we did in the tech field. A, we have a Nutanix Kubernetes engine, something that we offer, which came through one of the acquisition that we had done in the past, the D two IQ acquisition that brought that offering to us.
And then the NDB, which is the Nutanix database service, that's the one that we are talking about. That's part of it. And then we have one management console, which does the global management of all of Nutanix, which is Nutanix Central, and then all the operations governance pieces through the Nutanix cloud manager.
So that's the full stack that we provide. Now double clicking on the storage piece and what's our vision there, right? It's to provide a simple, flexible and, and secure data platform to build intelligent applications in a hybrid multi-cloud environment.
That's our focus that how do we really have a very simple platform which can help build intelligent applications, which are AI powered. And this platform is a software defined platform supports multiple protocols, so S-M-B-N-F-S S3 ICE czi. And we also provide the ability to run compute on the same platform.
So you have app and data on the same platform together. It also provides a very comprehensive data security offering. And we'll talk more with data lens, which enables cyber resilience, deep analytics and compliance.
And then this whole thing can be run anywhere either on our own NX hardware or our partners hardware, or you're running in the hyperscalers with AWS Azure platform as well. So that's the, the main data offering that we have. And they're managed through Nutanix Central, as I talked about, the Uber management console.
And then we have data lens, which is the security piece that we support there. So Vishal, um, global Namespace has a lot of different dimensions. I mean, how do you, where does Nutanix Global namespace play in this game?
Yeah, so if you have storage at multiple locations, you can bring all the storage namespace together into one global namespace, for example. You have buckets all over the place. You could have a global namespace to have a view of all the buckets that you have in different geolocations as well, right?
Similarly, if you're doing for files, you have files running across clusters. You can create a global namespace where all the, it stitch stitches together, all the namespace together into one global namespace. So you supply some sort of, uh, orchestration of, uh, data mo mobility movement that's and that sort of thing.
Wow, That's fine. I will get into those details as well. That's Good.
Yeah. And this platform is a very mature platform with very rich data services. If you look at a full data lifecycle management, so if you're looking to tier data to the cloud, it provides the capability not just based on age, but it also gives you rich metadata based tier, like if you want to tier certain kind of files or certain kinds of, you know, non p kind of data, that's the one that we are working on on sensitive data saying, Hey, I want to tier only files which don't have these kinds of sensitive data.
Like, so that is the value add that we like to bring there. Data ingestion, data mobility and data governance, all these pieces, and we have some more information later in the deck will talk about those very comprehensive data protection. Like some of those we leverage from our underlying NCI infrastructure as well.
So supporting metro availability, like one of the few vendors which does that with zero data loss, like we support that neary, async, DR snapshots, immutable snapshots, all of those are supported. And then data security, cyber resilience, permissions management, very like very commonly asked that can I know if I have overprovision the permissions on my shares and on my exports. That's their data encryption in flight at rest, right?
In use that's supported. And quite a few industry certifications around security. So that's all part of the platform that we have and we support a wide range of use cases.
So speaking of replication and synchronization, things of that nature, you, you mentioned, uh, synchronous, asynchronous, near synchronous. Yes. Um, do you support cascading solutions as Well?
We do support cascading solutions too, right? One of the very unique deployments that we have is that you could do metro across two sites and then you can do Async dr on the third site. This is a requirement that we see a lot more a Dora requirement in EMEA and Middle East, that we are seeing quite a few of those and we are one of the few vendors which can support that.
Good. Yeah. And this platform supports a wide range of use cases.
Like at the edge you can have a single node deployment, which supports even your AI inferencing use cases or, uh, inter enterprise app running at the edge. You have high performance use case and we'll get into the architecture of how we support that. Running a data lake on our object storage or supporting another other HPC use cases with general purpose storage, like lots of user profiles, home directories, department shares to machine data to application data we have, and then archive storage also.
So a deep and cheap storage to support the different use cases. So as you can see, it's a very comprehensive platform that we have been building for last, I would say 11 years now. Like, and it has been in the market since 2017, January, and quite a few customers who are on this platform, I.