CTERA Enterprise Intelligence: Unify & Activate Your Private Data for Faster & Smarter Outcomes
This session offers insight into the seamless integration of AI within the CTERA Intelligent Data Platform. Embedded AI and analytic Enterprise Data Services are explored, along with the underlying data fabric that facilitates secure, global data connectivity and ensures high-performance access. Additionally, the session demonstrates how agentic AI advances productivity, enabling virtual employees to interact efficiently with private data resources while maintaining robust security and operational effectiveness.
CTERA’s approach to enterprise intelligence is structured around three key pillars. The first pillar focuses on embedding AI and analytics within data management for enhanced security and data quality. This includes real-time inspection of I/O for anomaly detection and integration with security solutions like Purview. The second pillar involves providing a data foundation and fabric for AI training and inferencing, offering a global namespace for data aggregation and access, along with real-time metadata notifications and direct object access via a direct protocol that bypasses the CTERA portal.
The third pillar introduces agentic AI, enabling virtual employees to interact with private data resources efficiently. This pillar provides a semantic layer, allowing connections to various data sources (including non-CTERA systems), enrichment of metadata, and normalization of permissions across different data sources. This facilitates semantic searches and retrieval of relevant documents, empowering users with productivity-enhancing tools within a secure environment and answering questions with a chat-based interface.
Presented by Saimon Michelson, VP Alliances, CTERA. Recorded live on September 11, 2025, at AI Infrastructure Field Day 3 in Santa Clara, California. Watch the entire presentation at https://techfieldday.com/appearance/ctera-presents-at-ai-infrastructure-field-day-3/ or visit https://www.ctera.com or https://techfieldday.com/event/aiifd3/ for more information.
Transcript
For our next segment is the introduction for today's topic, right? We're here to talk about, um, enterprise intelligence as, as we observe that transformation as to, you know, why, uh, and Aaron will cover this. Why a lot of projects, uh, fail, or what is the biggest challenges when it comes to really activating the, the power of our data is lacking.
The, the data management, right? The clean data, the quality data that will help us really reach that promise and that goal. So our key, um, um, driver towards, uh, that, that outcome is to provide faster and smarter decisions, right?
Based on, on the data that you have. Um, and the way we look at this space, uh, of how do we reach enterprise intelligence is, uh, through an, uh, three different pillars, right? And pillar number one has to do everything to do with what elements within our data management solution we can leverage AI for, to provide better security, to provide better data management, and that's what we call embedded AI and analytics.
The second element is providing the data foundation for anyone that's building inferencing capabilities, anyone that's doing AI training, but they need an aggregator, right? They're looking to bridge the distance between where the data gets generated and where it needs to go to, to really empower that team that's building the next model. The third one, which a Aaron's gonna cover is, uh, around providing a agent ai, uh, the ability for us as employees, as subject matter experts to curate, uh, agents that are essentially trained in our private data estate, um, that are best suited to help answer questions and do tasks on our behalf.
We see this as a great tool to provide productivity as a great tool to provide efficiency to an organization, but it all starts with providing that, you know, unified data foundation. So, okay. So you guys have the global namespace that sits here.
Yeah. And I've got a NetApp follower in Boston, a NetApp follower in Atlanta. I've got some S3 buckets over here, and I want to present that unified data through that global namespace as I add those various sources.
Is there a requirement to actually mo is there a requirement at that point to actually move data? Or can, can it just present what's already sitting where it's sitting? So, for the global namespace, for the core product mm-hmm.
You do require migration because if you're looking to reap the benefits of, again, cost security, everything that we deliver, we replace an existing nas. However, in this, in the third pillar where we talk about data intelligence, um, we have built a platform that does not require for you to copy the data over where we can do the extraction of the actual payload of A-P-D-F-I leverage JPEG image of any video, uh, and then move that metadata to our semantic layer, which you can then leverage, um, um, uh, on your AI platforms presented through an MCPA chat interface and so on. So we'll get to that portion, uh, uh, in the third pillar.
Yep. But again, to if you, if you're, if you, if you're strictly looking to, to, again, to, at the benefits of what a global namespace has to offer, our approach is to essentially displace an existing, uh, uh, system. Okay.
Okay. Gotcha. Great.
So starting with our first pillar, um, and, and, and this is anywhere in our platform that we're gonna keep innovating and implementing, um, uh, AI solutions that can further improve on security, uh, or data management. Um, one of our announcements from I think, uh, a couple years ago, three years ago, is around providing, um, an ability to do inspection, real-time inspection of IO as it's happening on the different file shares to determine if it's a legitimate or illegitimate access. Okay.
White listing, of course, applications that are deemed as as trusted and reliable, uh, but trying to do anomaly detection, right? Identifying if, uh, a user or a workstation is compromised, um, and trying to identify are they doing, um, something that should be at least notified to it or enforced. We see a lot of examples working with customers that either experience data exfiltration or ransomware, uh, and helping them minimize the impact, right?
Having to recover perhaps a few tens of files as opposed to recovering. The whole file system means, uh, a lot from an operational standpoint. So a lot of our customers, of course, start with just enabling notification, okay?
They're not implementing the enforcement on day one, but it does provide them with that visibility, that forensic information of what we believe was impacted. Okay? Uh, once you're ready to imply the enforcement, we can also take action to block you from active directory to make sure you're using user session is terminated, uh, until you reinstate access after that incident is being investigated.
Okay. Is this entirely homegrown or do you actually have integrations with some of the existing, uh, you know, data security, ransomware, et cetera, uh, products that are available? The Product I was just describing is homegrown, but we also offer that integration through solutions like Varonis purview and any, anyone that can do, uh, maybe more broader inspection beyond satera as the data source that they can apply other signals for that they're capturing elsewhere.
Uh, so both, so both options are available and are complimentary. Uh, effectively this is a line of defense that we can have within our own file system, um, that we've, uh, implemented. Okay?
Okay. Um, the second element is around providing, uh, global analytics, right? And metadata capture.
So we had put together, um, a centralized platform that we host that, uh, captures all the auditing, uh, all the incidents and all the metadata changes that are happening in real time across all the customer environments. Okay. Um, we're gonna get to that later on in the presentation, but we're able to capture all these different signals and provide forensic information that dates back, uh, up to a year ago.
Okay. With, uh, easy filtering to identify if Jim, uh, between the dates of April 21st, uh, 21st to 22nd, what files were read and accessed, um, has a lot to do from, of course, you know, cyber insurance perspective, understanding what the leak was, uh, what was the level of impact. Uh, so these are all things you're gonna see in the demo later on.
Um, and then third, we had integrated MCP, uh, capabilities directly into the core product, right? Um, um, and MCP is essentially providing, uh, um, you can bring your own large language model into the data management platform. So that means you can automate workflows as a user.
You can ask to summarize documents that you have under management in satara. It can gen, it can author new documents for you. It could, um, create links to content.
It could facilitate a lot of the actions that you would otherwise do manually or try to programmatically do by coding an API, simply using natural language. So you can bring your claw, your cursor, right? Uh, your different platforms, and they're, uh, compatible now with satera.
Okay? So this is one easy thing to do, we can do is be, uh, is, is is provide a natural language way of interacting with your content, okay? Customers take that avenue, um, analyze their data, generate metadata, and feed that into the metadata analytics and create, uh, holistic cycle there.
Yeah. Cycle Y. Yes.
So, so that's where I think the responsibility of every vendor is, is to provide that, uh, MCP interface, uh, where you can do much more complicated tasks. So, uh, in other words, you know, one action could be take the content, the payloads from satera, do some, some level of processing, uh, generate some metadata, and then store that metadata, uh, in a catalog, uh, perhaps create another Confluence article, uh, open up an opportunity in Salesforce. These are all, Or enrich that, uh, file system metadata for subsequent For subsequent access.
Yes. So that's what you're gonna see in Aaron's presentation. Okay.
And this is what we call as the third pillar for, for data intelligence. Great. Okay.
So that is a good, good point. Bar Barton, you had a question? Yeah, I was gonna say, uh, if you read you, uh, today, you'd assume that MCP is everywhere, and that's table stakes.
Yeah. What is the reality? Do you think it's still a different differentiator for you all, or is it becoming table stakes at this Point?
Um, so Aaron will share his, his perspective. I, I believe it's becoming, uh, like the USB, the new USB, the thumb drive. Um, so it is becoming more and more, uh, prominent.
Um, and I, I think, uh, looking just from now, it seems like the industry is really standardizing on that as the, as the main form of access. So, but it's still worth calling out, you're saying, as opposed to, Um, yes. I think it, it, it is still, uh, um, relatively new, new newer capability, you know, in a year from now we might be here and not, not include that, but, um, but nevertheless, it's, um, I think as, as what we can relate to as users or even later on, well, I mean, as machines, uh, can do, it's, uh, our, our tasks are, are across a number of different systems, and I think, uh, it's a responsibility of every system to be able to speak that right language, to provide, to, to allow you to be more effective at the task at hand.
Mm-hmm. Um, so we do a lot of times repetitive things. If it's, uh, summarizing information from Salesforce and curating an email, uh, and opening up a Confluence article, um, or yet one of those data content management systems, and we have to, we, it's our obligation to deliver that to the user.
Okay. Yes. Thank you.
This was pillar number one. Pillar number two is, um, providing a data fabric, right? So, um, you know, we're looking at, at the reasons, uh, customers are, you know, looking at a global namespace.
Uh, and one of them is, is for its aggregation abilities, right? Ability to aggregate data across thousands of different sites into one centralized object storage, but then providing access to it. But how do you provide easy way, uh, and scalable way of accessing all the, all of this distributed data?
One thing is you need to present the entire metadata as a real time notification service. So being able to subscribe for any changes happening on a global level, right? Imagine you're managing a hundred or a thousand sites, understanding every rename, every delete, every upload of new content.
I can subscribe to an API and it will let me know of any change that happens worldwide. Okay? That could be a trigger for a data pipeline workflow that needs to create a, uh, a label for that piece of payload.
Uh, it could, uh, you know, try to produce some derivative product, um, and, and so on. Really, the options are endless. You had a question?
Well, a couple questions, but please, really, on this one, what's the, uh, what's the security around pub sub, Uh, Aaron, do you, uh, wanna Take that security Around? Uh, question. Well, but on pops, I mean, yeah, so we, we, you I'll go here.
Yes. Uh, so, uh, what we're doing is essentially as part of our rest, API, what we have what's called the notification, API, uh, so you, yeah, it works, uh, based on Kafka on the backend. Uh, but we, it's the same, uh, the same, uh, authentication as we have for all our regular APIs, uh, with the OAuth and or, uh, username, password, or shared secret, whatever you prefer.
Well, I mean, it's more that providing a, it's very interesting to provide a, uh, uh, uh, a pub sub feed of file system changes. Mm-hmm. It also opens up a whole bunch of security issues that you're notifying people of stuff that they should never be aware of.
Yeah. So it's, uh, it, of course, it enforces the, the permissions, so you're not able to register on, uh, folders or parts of the file system that you're not allowed to access. Okay.
So it's, uh, it's totally integrated with the whole permissions, uh, uh, mechanism and authorization of, of the global files. Okay. Thank you.
Um, so this is one element and what you would expect to get in that notification, uh, right, as a payload that tells you, here's the file information, here's the metadata, here's the access control for that file too. So if you're looking to respect permissions, uh, of that, of that file, that, you know, that platform can also receive that metadata as part of that same notification. Um, and then the third element, and Jim, you touched this, um, uh, in, in, in your question, is that we can provide access in, in a number of ways.
Of course, you can use the global namespace to leverage N-F-S-S-M-B and even S3 on our front end, but that puts satera in the data plane, right? And perhaps you're looking to run something much, much more scalable, right? You're looking to kind of use, you know, full, full throttle, uh, directing your S3 storage.
And the way we do that is through what we call the direct protocol. Okay? So we, in that same metadata service, we can provide you with, here are the chunks, here's the decryption key, here's the libraries as well for your application to be able to reconstruct that original object.
Okay? And that essentially puts you, uh, completely bypassing satera portal, uh, or the satera namespace. Um, and this is essentially what our customers are using to, uh, whether they're building a web application on top of that data, uh, or they're, um, trying to automate, uh, uh, uh, a data science project, or if it's, uh, you know, leveraging ai, uh, that's what we put together, what we call direct object access.
Um, and we also leverage within our own product, when the edge systems upload the content, they never write the data through a single point of failure, right? They write the data directly to the object storage. And this is how we're able to leverage the performance of all these individual nodes out there, uh, writing simultaneously.
So this would imply that when you said earlier an answer to I think Ray or Andy's question, that you copy the data off of the original storage and bring it into your satera storage platform, that you don't have to do that. You can still use SAT as a global namespace, and then just to have that redirect you to the actual storage location. Yes.
Yes. So it can do both. Does it do both simultaneously or is that either Or to answer, um, um, uh, Andy's question, uh, at the edge, right, when we're migrating, uh, let's say away from a NetApp file system, we need to copy the data over, so it has to move over under our management.
But once it's under Satter's namespace, then you can access the objects directly. You don't have to copy them again. You actually are migrating the data.
Yeah. Yes, Yes. We are displacing the original file system.
Yeah. Okay. I'll clarify a bit for that question.
Uh, so the format where we with the, of the objects is not the native format. So you don't have the whole file system laid out under the destination system because object storage, for example, does not know how to rename a directory or to do other file system things efficiently. Okay.
So it's, everything is stored in, in a ctra format, it's, uh, de-duplicated, right? Uh, but you, but we do provide, it's an open format that we publish and we allow you to access this information direct. So, but, but you'll see the files as, as chunks.
Uh, do you guys have like a metadata service or anything like that? Or is it embedded? So it is a separate metadata service.
So in that architecture, uh, that we've, we've shown the edge is the file presentation, but then the portal, that's the control plane and the metadata catalog, and that's how we optimize operations to Aaron's point, like rename, like move, right? These are things that we can do much more efficiently in that metadata database. And you Pro notification of met Yeah, go ahead.
And that's how we can relay those notifications. Yeah. Um, can as part of that notification, pub sub interface or something like that actually enrich the metadata that you are maintaining?
Or is it something that if they wanted to add metadata, they have to go out to their own metadata service to supply that enrich? So, So to to, to provide that enrichment, we're gonna get to it in Aaron section that we do provide a, a platform to help even further enrich that metadata, or to your point, leveraging a third party that's, uh, designed to handle things like, uh, image recognition, um, you know, certain workloads that are optimized. So yeah, that's all, both, yeah, both.
Okay. So you, you talked about separate libraries. Does that mean that if you actually are using the direct object storage access that you would actually need to link in new libraries to your applications to be able to access it?
You Don't have to. We, we provide those, uh, libraries if, if you choose to, to use them. Otherwise there's a, an a PIA one API call that tells you request a file ID, and you get back the list of objects and the decryption key.
Okay. So you can do it yourself. Okay.
But it's, it's not no, no longer as simple as doing an open on file descriptor. Right. Okay.
You, you need to bypass the concept of doing an open on a file descriptor and then reading that file descriptor. Yes. Okay.
So there's, there's a, there's one action, there's one call that is done before that. Yeah. Because we stored the data in a, in an optimized format, then there has to be that.
Yeah. So, so, uh, again, and I apologize for going back to this, but if you're rer, if you're rewriting the data and storing it in your optimized format, then when you say in direct object storage access that you bypass the tariff, you go direct to the storage, are you going direct to your rewritten version or to the native storage? Like in use case, if my data originally sat on, uh, an S3 or on, uh, right.
Do you have to pull it out? You pull it out of S3, rewrite it, put it back in, and then I talk to S3 directly? I don't, I don't understand how that would No, you would, so what you would do is, so you would invoke an API call on satera, right?
You'll get the list of the de-duplicated encrypted objects mm-hmm. And the decryption key, right? Okay.
And then that, what that library does is essentially reconstructs that file, decrypts it, and provides you with that payload. Now you can use that payload to write to another site, or, you know, you can label it, you don't have to copy it again. Yeah.
Effectively it's a read, it's, it's, it's a read. Yeah, go Ahead. I'll say effectively, instead of like doing the native non-direct version of it, you would send your request to that edge file or the edge file then reaches down in the object storage system, says, here's all the blocks that I need that represent this file that you want.
Bring it back to the edge file or present it out in the other use case, you still talk to the edge file as who you're talking to, but rather than pull the blocks itself, it says, here's the token to get access to that object storage, and here's the blocks you need, go get it. But you have to use their library to yes. Hydrate the data and, you know, construct the data in the proper sequence and stuff like that.
So you're, they are accessing directly the object storage per se, but they're using the library to hydrate the data and where it supposed to Be, like you're using as the metadata plane and then you're going to The master boot record? Yes. You have both options, right?
'cause you have the file that provides the native file interface, right? So it's just regular open and everything is handled for you. Or you can use our optimized, uh, our SDK in order to, uh, directly access the object storage and then you can, uh, enjoy even, uh, you know, better, uh, performance and scalability than what a filer can provide.
Gotcha. And how does the filer present the files at OIC CS N Fsk? Yeah, the filer presents, they're cached on a, on a local file system and CS and nfs.
Okay. Just some sort of process compliant file system. Got it.
Yes. Yeah. Per your direct object, is that using STS to actually get there?
Or is there, do you guys have your own home built implementation? Designed URL. Okay.
So we essentially, it's an authenticated temporary URL to doing that transaction. Yep. Okay.
Okay. So it's pre-authenticated, uh, it's part of our kind of zero trust approach as well. We don't want any, you know, system to have, uh, permanent access of those access keys and secret keys.
Okay. So assuming that you're doing S3 as the front end, is this Elle's own implementation of an S3 compatible protocol? We, we have both.
We have two options. Um, is one is yes, one implementation is that using pre-sign URLs, and we also have a mode that is not direct, not go, not bypassing the, the, the portal, which everything goes through a single node, which does a get or a put. Okay.
Yeah. Yeah. Um, the third pillar, um, is essentially providing that semantic cla, right?
And, and providing an ability to, uh, create that context database of all the information that's under management in satera, but also in other unstructured data sources, right? So using SharePoint, NAS systems, confluence web sources, um, these are all places where you could have content that could be then leveraged for, um, uh, a retrieval augmented generation, uh, database. Okay.
Uh, this is what Aaron is going to cover in his pres presentation, but, um, what are you gonna see, uh, over there at a very high level is a platform that lets you connect to a number of different sources, not only satera. Okay. Do that quality enrichment, uh, Brian to your question mm-hmm.
Of adding additional metadata, uh, into the, to ceras metadata, um, respecting user rights. Okay. So normalizing all the permissions across these different data sources into one permission model that we can then enforce, uh, queries that are being made to this, uh, to this solution.
Uh, and at the end of the day, create those, uh, virtual employees that are connected to your data sets, to the sources that you define as reliable, uh, and deployed to co-pilot to anywhere where you need those agents from a productivity suite standpoint. So imagine I can go into a, a chat based interface and I can ask a question, which is then gonna go and do a semantic search to retrieve the most reliable documents, um, and then provide citations, um, and including all that enriched meta metadata. Okay.
So this is something Aaron is gonna delve into, uh, more detail as we go through this presentation, but this is our approach, right? At a high level, right? Just kind of summarize, implementing AI in the core product, providing a data foundation for customers that are building their own solutions, and then finally taking that, uh, approach to build that semantic layer, uh, for those customers.
Okay. Thank you, Aaron. In this case, you're actually, yeah.
This is one of the only times where you're actually accessing the data and where it's, where it lies. Exactly. Without Copying it in and optimizing it, I billing it and all other stuff.
Yes. You're providing a, a semantic access to the data that we're, wherever It is. Yes.
This case, And we're not tied to satera per se, we can connect to other data switches is look up Right.