CTERA Intelligent Data Management from Edge to Core
Discover how CTERA addresses the complexities of hybrid cloud storage by enhancing operational efficiency and security, advocating a unified platform that extends from edge to cloud to manage increasing data demands. Practical use cases across various industries demonstrate how CTERA leverages AI-powered tools, automated workflows, and an intelligent data platform to improve data insights, security, and governance on a wide scale.
CTERA positions itself as a leader in distributed data management, providing a file system fabric that connects edge locations, core data centers, and cloud environments. Their platform allows for secure data access, protection, and sharing, forming the foundation for future AI applications. CTERA’s DNA comes from the security space and works with highly regulated industries such as healthcare, financial services, and government agencies. These organizations share common challenges of data cybersecurity consciousness, highly distributed data, and the need for information management.
CTERA addresses the increasing complexity of hybrid cloud storage by offering a platform that connects data centers, cloud providers, and edge offices. The company uses object storage as its backbone, providing reliability, durability, capacity, and cost-effectiveness. To overcome the limitations of centralized object storage, CTERA adds a complementary layer that caches data, provides performance, and enables multi-protocol access across various locations. The company modernizes unstructured data across operational efficiency, cyber storage and proactive data protection, productivity through multiprotocol access, and generative AI to improve productivity and security.
CTERA’s architecture consists of edge appliances that present file shares, optimize performance through caching, and provide real-time data protection to a centralized object store. This model enables near zero-minute disaster recovery and facilitates data sharing across geographically dispersed locations. CTERA simplifies data migration from traditional NAS platforms using built-in tools, allowing customers to move large datasets to CTERA-managed object storage, which can reside on-premise or in the cloud. The company excels in industries such as healthcare, public sector, retail, and manufacturing, where security, data distribution, and the need for modernization are critical.
Presented by Saimon Michelson, VP Alliances, CTERA. Recorded live on September 11, 2025, at AI Infrastructure Field Day 3 in Santa Clara, California. Watch the entire presentation at https://techfieldday.com/appearance/ctera-presents-at-ai-infrastructure-field-day-3/ or visit https://www.ctera.com or https://techfieldday.com/event/aiifd3/ for more information.
Transcript
Uh, my name's Simon, Simon Michelson, and Aaron and I are here to tell you a story. Okay? Um, but before we get into, I guess, the meat of today's event, which is surrounding intelligence, um, I'm here to tell you about who we are, what we do.
So for those of you who have seen this already, it's somewhat of a reintroduction is to satara. Why are we're here, why we're in the marketplace, and a chance to maybe see the world from our lens, right? Our perspective.
And we're here to talk about intelligence and intelligent data management. Um, you know, once upon a time, uh, there was a company that set out to really become the leader, uh, when it, when it came to, uh, distributed data management. Mm-hmm.
Uh, we're very passionate about the field of, of, of data and how data, um, goes along with data security. And we've identified a problem that data is generated everywhere. Um, and as a result, the requirement is to provide an ability to manage that data anywhere it gets created.
And we thought, you know, we're the best pioneers to help address that big problem, and we did it. Okay? We are a leader in the space of distributed data management, essentially providing a fabric, a file system that helps connect edge core data centers, cloud environments, with a platform that lets anyone access, protect, share information to really help provide, provide the foundation for what's coming next, which is intelligence.
Okay? So we started with information management and is over the years, as the industry has evolved as intelligence, um, if it's in the, for, in the form of BI data science or AI in, its, in its most recent reincarnation, um, it all, uh, they're all consumers of data, uh, and they require access to those data platforms. So we're a leader in the space.
You know, you've seen us, uh, in the marketplace for the past 17 years or so. I've been in the company for, uh, 15 out of those 17, and I'm responsible for alliances. So I look after all of our global partnerships, uh, and how we go to market.
And we work with some of the biggest providers out there with companies like IBM, Hitachi, uh, hp, as well as the hyperscalers to deliver our solutions to the largest enterprise companies out there in government agencies. Um, you know, when you think about satera, think our DNA is coming from the security space, and it lands itself really well to those more regulated industries. So we work with healthcare organizations, with financial services, with civilian federal defense agencies, um, and to them, they all share in common the cybersecurity, uh, consciousness, uh, with, you know, providing the right protection around their most valuable assets, which is their data.
They're highly, highly distributed, and they require access from any number of different locations. And then finally, they have a lot of information that requires management, which presents challenges. So why did we start this, right?
So one of the certainties we're operating in is that data is growing, and we've heard this for the past 20, 30 years. We know there's gonna be more data tomorrow than there is today. Um, and it's a result of any number of things with devices, with, um, all kinds of elu evolutions in our industry that essentially drive access to more data, more high quality data, um, which we're gonna talk about.
So data is growing, and that is, uh, somewhat of stating the obvious today, but so is architecture complexity, right? I think what we're dealing with today is something so, um, um, yeah, let's say a lot more complicated than what it used to be because we have options, and options provide us with flexibility. But as it people, we have to plan for the worst.
We have to plan for every possible scenario because it's very likely that we're dealing with data centers, with cloud providers, um, and we're always trying to platform from, for every possible scenario we're trying to optimize for, for performance, for accessibility, for cost, right? These are all different strings that the organization is trying to pull and we have to respond to. It's a very difficult task to do.
And that creates an opening for what we call hybrid and hybrid storage or hybrid cloud, because one platform or one solution, whether that's on-prem or in the cloud, may not provide the best, uh, of both worlds. It could perhaps provide a solution today, but we all know requirements change, and we have to keep being agile, responsive for those things as they happen. So the, some of the combinations, uh, you know, we're depicting here on the screen is these relationships between data centers and clouds, between edge offices and cloud providers.
And of course, um, a a handoff between data centers and cloud with edge file systems. So that's where it becomes extremely complicated, and we're trying to put together a platform that helps connect everything that gives us the flexibility that we need. So one of the interesting things as we kind of look at analyst and media coverage is about, um, not shifting away from the, from the thinking of purchasing equipment and purchasing storage and looking at purchasing platforms, right?
So elevating and trying to kind of, uh, put together the right solutions that will not only support us today, but will future proof us for, um, for the, for that, uh, you know, for that future. Simon, what is the, you mind going back, please? What is that?
Uh, the on-prem cloud hybrid, so to up to 2023, does that mean enterprises only have whatever, uh, 12% on-prem and now most of it's hybrid and that's overtaken it? Or what is? So, um, so what we're seeing here is essentially the shift over the years where, if we go back to 2006, a lot of the infrastructure started with on-prem cloud was very, very new at the time, and cloud has started to gain that market share.
And some applications, some workloads were, you know, perfect fit for, uh, transitioning to the cloud completely. But then, you know, around about 2014 or so, we started to see that, well, it maybe is not the best, best fit for every possible workload if it's for security reasons, for performance reasons. We had to have some sort of a combination that, uh, perhaps extends cloud to other sites, uh, that ties on-prem with the cloud and so on.
And that's where we're seeing that segment is growing. And yes, to your point is that the market, if you were to se segment it, um, you got a lot more share of hybrid than ever before, uh, trying to address these, you know, you know, these key requirements, if that makes sense. I think the question Is, is it, is it it spend?
Is it, uh, It's A great question. I mean, what's, what's, what's that percentage we're talking about here? Yeah.
How would you read the numbers? Numerically for cloud is X percent of what? Hybrid?
Percent of what? Yeah. So this is a relative chart, actually, uh, with a hundred, yes, it's, it's percentages, and I apologize.
Uh, we haven't actually included the axis in this version. Um, but you're looking at about here, uh, about 30%, uh, hybrid of what, 30% of what? Simon?
Uh, 30% of essentially it, um, um, IT budget, no, it, infrastructure solutions. Storage, distribution, storage. Okay.
Capacity. Okay. Okay.
Um, and that's what we're looking after is because storage specifically in capacity is the one that maybe presents one of the biggest challenges, right? With we're constantly trying to optimize for cost and performance, and those are somewhat of negative forces, right? Um, but there's a lot of infrastructure options available to us, and we gotta find a way to platform for every possible scenario.
But what's driving, uh, to your point, what's driving that modernization? Um, we've seen, um, kind of four main, uh, forces, okay? That essentially drive an organization to start modernizing their unstructured data, um, um, you know, across the organization, starting with, of course, what we call operational efficiency.
And that has everything to do with doing things cheaper, more effective, leveraging platforms that provide us with the capacity, uh, that we need. Uh, and tying in things like central administration and trying to do more with less resources, right? We constantly talk to it and they're, you know, burnt down with just a lot of different tasks being pulled to many different directions.
And they want solutions that provide the automation, the ease of use, the easy migration of data sets, and really optimizing their cost per gigabyte, per terabyte. Uh, next is resiliency. We see, uh, um, a huge shift towards what's called cyber storage.
Uh, and essentially providing proactive measures for protecting data, trying to minimize the impact of malware, trying to impa, minimize the impact of ransomware, right? Uh, recovery jobs can take a very long time. So if we can put together measures to help, uh, uh, detect and prevent it, uh, we can have less files that we need to recover, less capacity than we need to recover.
That has direct impact towards uptime, disaster recovery procedures, and means that we're, we're up and running a lot more. Uh, third is productivity, and that has everything to do with multi-protocol access from anywhere around the world. Now, whether you're a user or an application, whether you're A-C-I-C-D workflow that is generating a build that needs to be distributed to, uh, a remote site in a different continent, or whether you are Joe trying to share a file with Alice, uh, and you're trying to work together over long distances, right?
We're solving a very difficult problem because we talked last night about, uh, you know, um, pure latency, right? And how fast does it take to send information from one location to another? And we're all trying to work together.
And that's, I'd say, further exacerbated in the pre, in the last few years with a lot of users and, you know, applications accessing data remotely. So think of any producer, consumer architecture, uh, where data gets generated programmatically has to be analyzed on the other end of it. You need a data backbone to help facilitate that bi-directional information sharing.
And so this is what Cera is seeing, or is it what Gartner or Forrester Seeing this is, this is, uh, what Cera ISS seeing is this is why we're in the marketplace. So those three items that I just mentioned are what we call the fundamentals of why an organization today, uh, would choose to select cera as opposed to traditional storage platform buying a NAS system, right? Um, because of these key drivers.
Third is, and the last one, and I guess what we're here today to talk to you about is generative ai, right? Generative AI has, um, a lot of opportunity and promise, uh, to help us further improve productivity, further improve security, but I guess more than anything, we're expecting to turn that one gigabyte into maybe hundreds of millions of dollars because it packs so much information in it that we need to tap, right? And data, or AI or BI or data analytics, uh, solutions, they need access to information.
So, uh, again, they, what they all share in common is the requirement for a data plane, a, a backbone, to help provide that easy access. And the way that we do this is through, um, using object storage. Object storage is the key component that helps us really drive this whole modernization.
We are a software defined solution and a data management company, but we leverage object storage as our backbone, okay? Whether that's on premise, whether that's in the cloud, we use S3 compatible storage to provide the reliability, the durability, the capacity, the cost point that an organization needs to store, um, enormous amounts of data. But what's the problem of that object storage?
It's highly centralized. It's located in perhaps 1, 2, 3, 8 locations, but in many times, you need information to be presented elsewhere, and you also need information to be presented in other protocols, right? Not every application, not every user can maybe write or, you know, read or write information over S3.
They need access over file protocols, like NFS SIFs, maybe web access and so on. And that's the complimentary, uh, portion that we add on, on top of object storage to help cache that data, provide the performance, provide the accessibility over multi-protocol access across any core edge or far edge location. So Is, is the backend storage layer always object storage, or is there diversity of supporting other things?
I would say 95% of our implementations leverage object uhhuh. We do have some smaller implementations of, let's say half a petabyte and below that we allow for using NFS as a backend, okay? Or even block storage in some instances.
But for any workload, typically beyond half a petabyte, we see object storage as being the kind of go-to platform, okay? Um, so we provide that reach, right? And when you look at the types of use cases that we deploy, it could be anything, any, any anywhere from protecting data at a high volume, providing an on-ramp from file to object storage, an easy way to archive enormous amounts of data, uh, providing a fabric where you can securely share data amongst any number of locations, uh, providing the proactive measures in the file system to help detect and enforce things like, uh, malware and ransomware, which we're gonna talk about.
And then finally, as we're venturing in venturing out to the AI space, we're seeing a lot of use for our platform for being that aggregator, being that easy way to capture all this data and bring bridge that gap, that gravity between where the data gets generated and where it gets consumed, okay? And that's why, uh, we're here essentially providing that data foundation for, uh, for all these organizations. And I'm gonna dive into that as we go.
Now, if you take these use cases and apply them to actual workloads, right? If you think about storage as this pyramid, right? Where you're trying to optimize at the very top for performance, and as you're going down the pyramid, you're looking for more capacity and improvement of cost, right?
And you can probably think of just five use cases each and every one of you and everyone here joining us today of where this might be a great fit. So here's some examples of workloads that we've implemented that this architecture is actually terrific for, and these, these are all also, um, workloads that you can see why AI bi data science solutions, uh, can help even provide more value in the future. But first task is manage, right, protect and, and provide the cost and performance, right?
Not introducing any change to how machines and users access, but at the same time, creating that single source of truth in object storage that we can then later capitalize on. So here's some examples where we've implemented PAX Medical Imaging Hospitals, uh, across 150, um, hospitals, for example, here in the United States, MRIs, CAT scans, right? Think about all that intelligence or all that information that captures things like diseases, right?
Can we, can we provide better patient care? But table stakes first step is help me create that single pool of storage that I can then present to my data science teams, my research teams. Well, same, same goes for, was there a question?
Yeah. Well, so I mean, yeah, the, the before the animation of this slide, you were representing, you know, effectively class full object storage capabilities, but then you're showing things like here, here, like general purpose NAS and software distribution that are gonna be read heavy. How are you handling the issue of I'm gonna create more fees as I go down that stock?
Is it gonna be in, is it intelligent of this kind of capability lands here, or is it a manual management of what goes into what class of storage? So, um, and, and maybe I'll try to respond. You tell me if I respond to the, to the, did I understand the question correctly?
Um, but in that model before, um, the access points, right? Um, for reading and writing data are through satera platforms or SATERA presentations, CED representations of the data. Mm-hmm.
Okay? This is how we guarantee to provide the performance, uh, for those users. Um, the secondary copy, or the authoritative copy essentially resides in S3.
Okay. That is the copy that is then leveraged for, uh, any derivative product, any application, any data science, um mm-hmm. It's out of band, essentially, and not impacting, um, the actual workloads that are generating the data.
Okay. But so, I mean, you had hot, cold and archive. Yeah.
So I, you know, aligning that to S3 standard S3 storage, infrequent access and laser. Yep. So if I'm doing software distribution and I've got, you know, my iso if my software out there, you know, is it gonna automatically tear itself up into standards so that you aren't getting nailed on read fees?
Yeah. So it is exactly how we do it. We do this from file to object, and then also within the classes of object storage.
We don't support today glacier, however, but we do support infrequent access and we use information lifecycle management to essentially move data that's less frequently accessed into a cheaper tier. Okay. Again, optimizing for cost and, and performance.
Okay. Um, we actually work with customers to model their access patterns. Um, sometimes if you accessing cold storage more often, then you know your price will completely change.
Uh, but we, we work with our customers. We have a way of discovering the data and modeling for that. Um, uh, but yeah, this, yeah, very Good question.
And then when you present outwards, does it present outwards as S3 as well, so Exactly. 'cause You're doing a unified namespace here, it's effectively what it is, right? It's effectively, effectively what?
It's Okay. So if I'm a writing application that is class full Yes. And is aware of the app of the, of the class that I'm writing to mm-hmm.
Do you abstract that class from it? So if I think I'm writing infrequent access, you put it in whatever tier you actually think it needs to be in, but what you respond back to the application is always whatever that class is. Yeah.
So we provide on the front end, both file and object. So we would completely abstract that to you as the, as the user creating or reading the content. And we can also provide you with direct access to the object storage if you wanna bypass cera and not have satera in the data plane.
Mm-hmm. You can do that as well. Now, while we don't write native format of objects, we provide you with the metadata and the ability to essentially read those optimized objects directly.
And this is what enables those large scale data retrievals, um, metadata labeling, uh, and so on. We are an aggregator, but then you can bypass the terra and work directly on the, on the underlying infrastructure. Okay.
Okay. Thank You. Great.
Um, just to continue on that, just to, I guess to complete our introduction portion. Um, so where are we complimentary and where we're competitive? Right?
So we work with, um, as a software defined solution. Uh, we have a lot of partnerships both at the edge as well as at the core. So we leverage a lot of server storage hyperconverged platforms.
Uh, we run as a virtual appliance, um, and it's very easy to deploy in an automatic fashion. Um, and then on the other side of it, on the backend, we support practically every S3 compatible target. Okay?
Um, any S3 solution that presents get put deletes, um, is fine for us. We have those certifications, of course, with all the on-premise players as well as the clouds providers. Where do we compete?
Most of our total addressable market is essentially going after traditional nas. Um, and that includes, uh, vendors like the ones you can see on the screen where you're dealing with a lot of complexity, a lot of gravity of that infrastructure. It's not easy to protect, it's not optimized for cost.
And you're looking to get off from the, the solutions for the reasons that I've mentioned previously. Now, we're not alone in this market too, right? There are other cloud gateway solutions that provide this, um, unified namespace capability that we may go up against, but our day-to-day business is really going after business as usual, right?
Everyone that has a NAS appliance, how does it look, right? So we spoke in a lot of abstract terms, right? So if we distill what are the main key architecture elements that make up a global namespace, Jim, to your point, um, it's, um, an edge, uh, appliance or an endpoint that gets deployed to a user's laptop or a VM that gets deployed on a commercial hypervisor, VMware Hyper V and K vm, and it's responsible for presenting file shares, right?
That's task number one, file presentation. Second is optimizing for performance caching, frequently access data to provide the best performance at the edge, and then providing that real time protection of all that information back to a centralized object store. Okay?
Through this model, we're able to provide nearly, nearly, you know, zero minute disaster recovery because we're able to provide any number of representations, instances of that same data over file protocols, right? So this is how you have an easy handoff from a site that experiences an outage to a, a core data center that has a cached ready and available presentation of that same data. This is how we facilitate sharing, right?
You can deploy an edge appliance in San Francisco, you can deploy another edge appliance in New York, in Hong Kong, and they all subscribe to a global namespace, and they're starting to present that same data across any location. Now, one of the key elements that we do that it's it's extremely difficult is distributed locking, right? How do you lock data over long distances, right?
So these are one of the elements that we addressed as part of this global namespace guaranteeing that strong consistency that if you open up a file in one location, you can only open up for read only at the other site. Now, of course, it's not a bulletproof solution that the, the network does go down and there are requirements for offline access. Our fallback is always to use things like version and control, right?
To be able to store those different versions and reconcile not automatically, but at least store the versions in the system. The last thing I'll talk about change, change can be difficult in every storage product, every storage product, we make it very easy. We have built-in tooling inside the product to discover and migrate data sets out from all these traditional NAS platforms.
So think of targeting an SMB share and NFS share, copying all that data, uh, over automatically. And we've delivered on projects that moved 15 petabytes in six months, right? And we were able to do them with just one, one point.
Like Let's say you have a NetApp filer today and you wanna move to satara, you're actually migrating the data off the old filer to Sentara object storage. Yes. Yes.
To, so not Satra object. The object storage is we leverage from either hyperscaler or an on-prem three compatible target, but we would, yes, we would replace, displace that existing file system that filer and copy the data over for those reasons we mentioned before. If it's for optimizing cost, if it's providing the ability to replicate that data to other sites, if you bring it under our management, this is how you get all the capabilities.
And the reason we can, we're also, um, you know, best suited to, to, to, to innovate is because once we own the file system, we can then provide those things like proactive measures against things like ransomware, right? So we have a lot of, so It's, I mean, I think I, I want a little bit more follow up on Ray's question. It is are, if you move something off a NetApp filer, what is your backing store at that point?
Okay, That's a great question. It could be that same NetApp filer. Okay.
That's, it could be that same NetApp filer. It could be a server that's presenting directly at that store, Just moving it off the NetApp NFS filer and putting it on NetApp S3. I mean, how does this work?
I mean, are you actually moving the data is the question? We are moving the data, yes. We're moving the data.
The data is moved and, but the data is presented or backed on our edge by directly attached storage. So this could be a data store on VMware that, that is presented to via NFS or, uh, or, you know, directly attached storage. And then to us, it looks like a virtual disc.
Okay. Yeah. Okay.
But you, you actually do require the storage overhead necessary to make a copy of it? Yes. Okay.
The data gets copied over, uh, and then we then replicate it to an S3 backend, which could be a NetApp as well, it could be a NetApp storage grid, S3, it could be, uh, IBM Cloud object storage. It could be an Amazon S3 bucket, it could be Microsoft. Uh, and the last element is to what are those, some of those industries that you see us, you know, excel in?
Right. Uh, I mentioned our emphasis or background is coming from, uh, security. Uh, most of our contracts tend to be very large in size.
Uh, we're dealing with half a petabyte and above. Um, and the industries that typically find us extremely, uh, relevant, uh, are the ones that have, uh, they're very security minded. Uh, they ha they're highly distributed and they have a lot of information that they need to manage.
They're starving for modernization. Um, so that's where we excel in a lot in healthcare organizations, pub, public sector, defense, civilian agencies. Uh, we work with, uh, uh, retail with highly distributed stores and warehouses.
Uh, and it could be anything from video surveillance to, uh, again, medical imaging anywhere where you have, uh, a nas, uh, file front end.