Data Mobility and Security with Nutanix
At AI Field Day 7, Vishal Sinha of Nutanix presented the company’s comprehensive approach to data mobility and security in the context of AI and enterprise computing. He detailed how Nutanix facilitates seamless data movement across edge, core, and cloud environments by supporting various data operations such as synchronization, replication, and distribution. This data mobility capability allows for use cases such as consolidating data from edge locations to a central data center and then utilizing the cloud for AI model inferencing and tiered storage. The platform maintains a consistent global namespace, simplifying data management across disparate environments and enabling organizations to process, analyze, and store data efficiently regardless of location.
The focus then shifted to security, where Sinha emphasized Nutanix’s architecture as being inherently built with security in mind. The system features cyber-resilience measures that detect both known and novel ransomware threats using built-in file system logic and a behavioral-based detection engine. By continuously monitoring file activities and validating potential threats before alerting users, Nutanix minimizes false positives while maintaining an exposure window currently down to ten minutes, with plans to reduce it further. Integration with third-party security tools like CrowdStrike, Palo Alto Networks, and Splunk enhances the platform’s capabilities within broader enterprise security frameworks. Immutable snapshots and a powerful remediation engine further allow customers to take automated or manual action to protect their data and restore functionality with minimal downtime.
Further elaborating on the benefits to enterprises, especially within AI workflows, Nutanix offers Data Lens—a metadata analytics tool that provides full visibility and governance over data usage. This tool can trace file activity over time, identify anomalous user behaviors, and clarify permissions with bidirectional data access views. These capabilities improve compliance, auditability, and operational transparency. Nutanix’s philosophy centers around four foundational pillars: simplicity, security, ubiquity, and unification of storage formats. By delivering an integrated, software-defined storage platform that supports files, objects, and blocks, Nutanix enables organizations to build and maintain intelligent AI-driven applications while maintaining robust data governance and near real-time threat response.
Recorded live in Santa Clara, California on October 30, 2025 as part of AI Field Day 7. Watch the entire presentation at https://techfieldday.com/appearance/nutanix-presents-at-ai-field-day-7/ or visit https://TechFieldDay.com/event/aifd7/ or https://www.Nutanix.com/enterprise-ai/ for more information.
Transcript
Quickly just, you know, this was one other question you had earlier, like about the data mobility piece. So again, we provide a comprehensive data mobility solution. Like I just took a use case, this is actually a customer use case where customer was doing all the, some of the bio research on-prem and they wanted to do the training and application the cloud for them.
The requirement was, all my data is being generated at the edge. I want to take that data and bring it to my main data center, which is like a data consolidation use case, right? So from multiple sites as end to one kind of consolidation of the data there, they were doing the processing and they were using the cloud to do some further inferencing on it.
And also the tiering and all our solution supports the sync, the replication piece we talked about. It supports the consolidation of data. So you can have N to one, you also have the distribution of data.
So if you want to do, we support that with objects today files is something that we are working on where from one location, if you want to distribute artifacts across sites, across different files or object server, that's where the distribution piece comes in. So this provides end-to-end data mobility between edge code and cloud. Like that's the use case around global data management.
And the best part is all of this is can be in the same global namespace. So it simplifies things even further, like in how you do the data mobility. Next one is security, but if there's any question I can take that otherwise move to the security piece, we have a lot more.
So I just quickly running through a few items there. This is a big piece, like we believe that security, like one of the architectural principles as I mentioned is it's a secure data data pipeline, right? What we provide for security, lots of focus on cyber resilience, being able to not just detect known ransomware, but also new ransomwares, which are not there, right?
And I will walk through how we do it, but at a very high level, any known ransomware, we have built the logic within our file system itself. So in real time we can detect ransomware and block it. And we have customers who have been saved because of the capability for new ransomware.
We have a behavioral based detection engine. And in one of the upcoming slides, I will walk you through how we do that piece. The second part is the data analytics and lifecycle management.
So this is another big ask, like primarily around curing data blindness. Like this is a challenge where customers typically didn't have lots of visibility into their data, what's happening to the data who are accessing. So it fills that gap for customers.
And then the entire audit trails and insights based on the audit trails is what provides the compliance and assurance piece. So let's look at the behavioral based ransomware detection that we have built into the product, right? So our detail lens product gets a continuous stream of all the files and objects activities happening on the, on the store systems, right?
All the reads writes open, close, everything that's happening, change permissions, it's looking at all of them. And, and it runs those events through a a detection engine, I will say ransomware detection engine where it does pattern matching to see if it is matching any of the common patterns around ransomware. And if it matches any of those patterns, instead of just alerting right away, it actually goes and validates whether the file has been encrypted or not.
Right? So this is our way of reducing any false positives by actually validating it. And if that is positive, it triggers the whole remediation engine.
So we provide a policy based engine where you can take actions, your actions could be blocked this IP address from where we are seeing the attack coming or block the user. Or it could be also just make the file object storage read only so that you can no, can't do any further damage to it. So that's the piece which is the exposure window is a very important metric that we look at from the time of detection to the time of blocking, how long it takes.
In the past we were, we had 20 minutes or less, we have reduced it to 10 minutes. So now within a 10 minutes exposure window we can detect new ransomware and block it. Right?
That's the capability. Sorry. So how are you seeing this relate to AI workloads?
Yeah, so primarily we are seeing that all the data that we are putting out there, making sure that we are providing the security of that data. Like, so this is for AI and in general all the other application use cases we are seeing the, the security of that data becomes important. Like if, because part of the AI workflow will be that you're bringing all the data together, storing in one location, making sure that that data is secure there.
Like you're, it's not being like, you're not getting ransomware attack encrypted and then now you have to pay for that part. So essentially securing the data that you have that piece. Michelle, do you, uh, interface with, uh, other security products in the, in the uh, enterprise?
Uh, we had a prior client, uh, client prior yesterday talking about networking, security, SOC knock, all those sorts of things. Do you, do you provide hooks into that if you're detecting some sort of ransomware activity going on in your storage? Yeah, so just actually few weeks ago we published our, like, so we are part of the CrowdStrike marketplace you can get, so we feed all the detection piece into the sea solution for CrowdStrike.
Palo Alto is coming soon. So we are definitely going after some of these bigger solutions and Splunk is the next one that, that's something that we are looking at. So yes, we are feeding into these.
Did you earlier mentioned immutable storage, is that your snapshots are immutable? Yes. Or can be immutable Or it tho those are all immutable.
All This all snapshots are immutable. Yes. Yeah.
Right. So product coherent ENS consulting. So why wait 10 minutes?
How, how long does it take to detect? Yeah, so remember right now this product, we run it in the cloud. So you have your storage systems running across different locations.
We get all the activities, we feed it into the cloud engine that we are running, where it goes through the algorithm, test it out. If it detects an ly, that's when it actually goes back to probe the looks at few of the files to see whether they're really encrypted or not Right. To avoid the false positive.
So that whole thing that it looks through validation and then we just don't wait for just one event. Like we look at few more to make sure that false positive avoidance is a number one thing like for us, because that fatigue of getting false positives makes the customer not trust the system and then some real events will get through. So we spend lots of time making sure that whatever alerts that we generate are not false positives.
Right. Because in 10 minutes, I mean a lot of damage can be done right? With Yeah, That's right.
That's right. And today the best part will be the solution is also now coming on-prem like in by January timeframe will have it in on-prem, then the whole going to the cloud, coming back, testing all those, that cycle will be, so we started with 20 minutes exposure window, we have brought it to 10 and we'll continue to reduce that window. Right.
And these are then like a, if you detect a, an actual positive uh, attack, do you go and quarantine those files and That's right. Recover. So that is the remediation engine piece that you see there.
So that, that's a policy based engine that lets you define the policy. Like the first policy will be how to prevent further damage, right? So for that, it'll give you the ability to block the client, which is doing the encryption block, the IP address.
You can make the files for object storage read only so that you can't do any more encryption pieces. So it provides those capabilities as a runbook for you further. So we don't just stop at detection and blocking.
We also provide the ability to do one click recovery. What that does is that it looks through your snapshots and finds the first unencrypted version of that file and with few clicks it'll bring that back, right? So it's a comprehensive solution for near real time detection, not a backup based solution, which is primarily recovery, but instead being able to detect in real near real time block it and then also provide the recovery.
How is, when you block, uh, an inference query, how is that returned back to the higher level architecture? Is that gonna come back as a regular expression? I mean, as a, um, an error as as An error access at that point?
Yes. Okay. Yeah.
Do y'all trickle that up as as like a 4 0 3 or something? So Right now for the, so the alert actually goes to the platform engineering team saying, Hey, we detect detected something. So yeah, today it, the platform t will know that something is going on.
At that point the client just gets access denied at that point. Is, is that data enriched up to the customer Today? I guess that would fall on Kubernetes or the app or whatever.
That's right. Yeah, whatever the behavior. So, and I I, I'm digging into this because to to your point, 10 minutes doesn't mean squat.
Something happens in 10 minutes. I don't care. Something happens in an hour, I don't care.
What I do care about is now I have to wait until I come back to that operation and restart it. And so what, what I like to see is when something like this happens and we stop the progress, either model building, um, doing some sort of rag pipeline, whatever, once that stops, what is gonna give me the data to get back into it? What is gonna bring me back to normal work?
And today, that's important point because if you look at with the traditional system, that whole workflow is more than two weeks. Like in fact, that's some of the data published there that in a traditional system where you are relying primarily on recovery, now you have to go first. You have to quarantine the entire system, do the contact tracing to figure out which systems are infected.
Are, are we more vulnerable to those attacks? Actually, you still, so it'll not be in minutes because you still have to make sure that whichever systems are infected, are quarantined cleaned before you can turn things back on. What we do is that instead of waiting weeks to get that system back up and running, we could reduce it to hours.
That's the value of the solution. Mm-hmm. And but you're, that's assuming we're talking a fully compromised system or That's right.
One you consider compromised. That's right. Yeah.
Mm-hmm. And, and the, the policy based action is something the client, the customer actually defines what he wants to, that's whether he wants to go back to some snapshot that's not encrypted immediately or decide to try to bring up multiple, uh, you know, files that all represent one system at some consistent point or something like that's which is a different game, Right? Exactly.
You're right. Yeah, because some basic, like what we have seen is customers who start with the system first, we'll put the remediation policy as alert only. Like they don't trust the system yet.
They want to not get clients blocked, as you said. And it was a false positive. And hey, everything blocks and now I have to get a pager and I have to act on it.
But they will generally start in the monitoring mode and then they build trust and they go, lately we see a lot more our clients, uh, customers going more with the actual blocking Policies. You mentioned earlier a lot about disaster recovery as well. Is there a policy based disaster recovery capabilities as well in that case?
That's right. Like when you say policy, it's like the R-P-R-P-O and Well, it's, it's, it's like that. But you know, how am my system brought up in case, you know, one cluster goes down and using another cluster, how are they brought up and sequence?
What's the proper That's the runbook staging. Yes. I'm sorry.
That's part of the runbook. So the Dr. Runbook where when you do the recovery, it runs through the runbook and based on that, it'll bring it up.
That's All, all I wanted to know. Thank You. Yes, yes.
Quickly, this is, and we will do a live demo of this. So that's why I'm just putting the teaser out there. This is our data lens platform.
This is, this looks at all the metadata and user data to build the intelligence. And primarily what it provides is the visibility to what the storage system is doing, what kind of files are stored, what access is happening. And it also provides deep visibility into what the users are doing, right?
One very typical example was that we got, we got a query saying, Hey, looks like we lost some files. Your system is deleting some files we turned the data lens on, and which quickly we found that there was a rogue script which was coming and deleting files and the customer is not aware of it. The whole lineage is what gets tracked in the system, right from the time of file is created to any renames, to any permissions change to, to anything that's happening.
You can put the file name in there, it'll give you the entire history of what has happened to it. But if you want to see a user and what has happened with the user, you can see the whole user activity there. So this gives you a very comprehensive view into your user and into your storage system, right?
And, and the next, okay. So, so that's the part. And the third big piece is also the permissions.
Like for all, it gives a bidirectional view on your files and folders and the permissions that the user has. So you can click on a file and can tell you which all users have access to that file. Or you can click on a user and it'll give you all the files and folders that user has access to.
It provide gives you a bidirectional view on the users and the permissions that they have on different files and folders. So it provides a very comprehensive data analytics, both from user and the storage perspective that you can leverage to make sure that there's full governance built into your, into your data pipeline. Right.
And this is my last side, which brings all of this together on the storage side, which is like the four pillars, like what we have built the platform on, run it simple, right? Which is having a single turnkey enterprise AI platform for all the data and the application that you build, run it secure with some of the capabilities that I just talked about in terms of data lens and what data address data and flight and data and use security run everywhere. That's the foundation for us.
Given we are software defined, the same solution will run at the edge in the code and in the cloud. And then the whole run unified, like you don't need a separate siloed solution for running files, objects of block. All of them are consolidated onto a single platform to give a very turnkey solution to build intelligent applications.
So just before, uh, so you support SSDs disk storage, uh, you know, you don't actually support other storage subsystem behind this, uh, and sort of direct path mode or anything like that. Do you, when you meet, it's a, I wanted to do some Ray Lucchese storage NAS box over here behind your solution. You're not gonna allow that as you're being a front end to that.
That's right. That that's right. Okay.
That's what I wanna know. Yes. Although.