Reimagining Data Management in a Hybrid-Cloud World with Qumulo
The presentation by Qumulo at Cloud Field Day 23, led by Douglas Gourlay, focuses on the challenges and opportunities of modern data management in hybrid cloud environments. The presentation emphasizes the need for a unified, scalable, and intelligent approach across both on-premises and cloud infrastructures. The speakers prioritize customer stories and use cases to illustrate how Qumulo’s unique architecture provides enhanced performance, visibility, and simplicity for organizations.
A key theme of the presentation is the importance of innovation, specifically in addressing the evolving needs of customers. Qumulo focuses on unstructured data, highlighting its work with diverse clients, including those in the movie production, scientific research, and government sectors. The presentation highlights how Qumulo’s approach enables both data durability and high performance, particularly in scenarios involving edge-to-cloud data synchronization, disaster recovery, and AI-driven data processing.
The presentation showcases how Qumulo enables freedom of choice by supporting any hardware and any cloud environment. Their solutions are designed to manage large-scale data, extending file systems across various locations with strict consistency and high performance. By leveraging cloud elasticity for backup and tiering, Qumulo offers cost-effective options for disaster recovery and provides the agility to adapt to changing business needs.
Presented by Douglas Gourlay, President and CEO, Qumulo. Recorded live in Millbrae, California, on June 4, 2025, as part of Cloud Field Day 23. Watch the entire presentation at https://techfieldday.com/appearance/qumulo-presents-at-cloud-field-day-23/ or https://techfieldday.com/event/cfd23/ for more information.
Transcript
Outta curiosity for you guys this session before I go burning through slides at, you know, a high rate of speed, what would make it actually interesting? What would make it the most fun or exciting session you've had today? I, I'd like to hear about your customers.
You must hear about our customers. Okay. Uh, what about 'em?
What type they are? Um, the big problems you help them solve. Okay.
You could definitely focus on that. Yeah. Talking stories.
What's That? Talking stories like, you know, like talk stories. Stories, yeah.
Like, you know, like stories about clients, stories about the company, like stories about the product. I could do that. Yeah.
Stories about clients doing cool stuff. Yeah. Alright.
Gimme one other topic or one other interesting thing. Surprises. Surprises.
Like if I took a grump quad and threw it at Brandon over, that was a bad Throw your life on. Wait. Oh yeah.
So we could poise and surprises. All right. We'll try to do some interesting, surprising technology.
I will definitely put it in the context of stories about customers doing interesting things and some of the challenges they've had. Um, this is a fluffy slide of people we can ignore that. Um, I always loved this quote though, by the way, innovation distinguishes between a leader and a follower.
A fun question I like asking customers of some of the more legacy suppliers in our industry is name a feature delivered by vendor X in the last five years that's made your job better. It's made it easier. It's made it more fun and more interesting because one of the challenges, what's that?
If it work, what's that? If it's, if it's working, yes. If it's work and if it's working and if it's, I mean, I came from an industry that was, had many vendors challenged with software quality.
Uh, what I have generally found in this network attached storage space is generally pretty good. Software quality across the board. You know, there's functions that sometimes don't work, but like data durability and keeping things available is generally pretty good as an industry.
So I'm not gonna sit there and use that to like bash competitors. But what I have seen is in many cases, a lack of innovation. A just establishing that status quo is somehow okay, we shouldn't be too thrilled about, wow.
A new software release that had a few more point features. But what truly disrupts what I've also found is interesting is listening to the stories of our customers and asking them what are their priorities? Not just what are their pain points, but what are their business priorities?
What are their priorities from an engineering perspective? Sometimes from an operations perspective and from the executive side where technology intersects the business, what are they doing to grow the business with technology and vice versa? And what we find in those scenarios is oftentimes by paying attention to the operator, you get a very different perspective on what is important to them than the person necessarily designing the system or solution.
And oftentimes an extremely different perspective than the person that's writing the check or opening the purchase order, right? One believes sometimes a lot of the marketing hype. The other is where the rubber beats the road and has to interact with your system every day.
How you make that person's job easier, how you make it a lot more fun. How you let them sleep at night is sometimes a very different set of innovations than the thing that creates a really glorious headline. So it's finding the ability to do all of the above and balancing our investments to continue to innovate, but supporting all of those missions of our clients.
So yes, we focus largely on the left on structured data, which is lots of cool stuff. Structured data tends to be a lot more ones and zeros, a lot more block oriented, a lot more online transaction processing databases. Uh, we are not focused on that side of the world.
We stick with the 80% plus that is unstructured and growing rapidly. And we have customers who do all of these things. Customers who make movies.
That's kind of cool. Uh, most of the movies you've probably seen in the last year have in some form or fashion been built or stored or developed or edited or composited or rendered on a Qumulo system or in I'd say probably 10 of the 10 largest video production houses of the world. If you're taking cool, fun drugs, okay, maybe cool drugs, maybe interesting drug.
Mm-hmm. Maybe not the fun ones. Um, there's a very high chance those were developed and stored on one of our systems.
Uh, 10 of the 10 largest NIH and NSF grant recipients run our systems to store things like one customer said, we're storing the cure for cancer on a Qumulo system. Now it's our job to keep it there. Adorable enough, accessible enough, perform it enough for them to unpack that, make it repeatable, test it against multiple genotypes.
We have another customer using the protein folding at home system that runs on us. Another customer who said, um, they build animated movie, uh, creatures and characters in movies that are really popular with kids today. And they said, our job is to preserve those so that your grandchildren can watch them someday in their perfect unaltered unedited state.
Be they named Kevin or Bob or whatever. You have another customer who says the wildfires that sweep across Southern California are modeled and that data is sent to the firefighters. So they determine where to build the fire breaks, where to evacuate, and that data is stored on us.
Um, how old, uh, this product is? Pardon? How How old this product your of yours SLO has been in business for 13 years.
13 years. And when you, uh, decide to invest in the, um, certification for healthcare, The certification for healthcare, there's multiple. Yeah.
Like in the country, usually it's like this that if, um, if the corporation, healthcare corporation have some, um, storage solution or any solution, they need to kind of certificate inside of their infrastructure. Uh, we store multiple millions to billions of PAX images for many healthcare systems across the us. So this was your, uh, one of the first, uh, The design point was initially how do we store small files efficiently and how do we store large files efficiently and erase your code them?
One of the legacy systems would triplicate small files, which sounded like a really good idea until we had a customer with 30 billion files in one cluster storing more storing mortgage data. Okay. Right Now another customer said, we have lots of mortgages.
Our normal mortgage underwriter gets three on a day. Uh, they're running more efficiently, they're getting six. They started using AI processing technologies with 30,000 plus concurrent connections connected to one of our clusters.
And now they're getting 14 underwritten a day per person. We had one customer who had PAX images. They wanted to go to A-V-N-A-A vendor neutral archive.
And Brandon did the math on this one. I'll introduce Brandon here in a second. But the math they were doing said that it was gonna cost them $800,000 in API charges because each file was written out and then written back in as an object.
Read another file, write it back as an object, an object store. It's modern, it's cloud native. We must go with the object store because it's cloud native and cool.
If we bin pack the files into the most efficient size for the object store to read or to write to and read from, we reduced it from $800,000 of API calls to $180, not 180,800,000 to 180 by doing it more efficiently. Read caching, let us improve the performance by having a read cache that is 92 to 98% hit rate. If I can conserve 92 to 98% of the data out of a read cache without calling back to a, a spitting disc or even a high performance NVME drive, can I then stretch the distance between where the cache is and where the data store is?
Could the data store sit in the public cloud in a coherent caching instance, stay at your house. I have a 500,000 image Lightroom library that when I go home, I access locally, I go to my parents' house and I'm accessing it locally, get it stored on a culo cluster sitting in our office. So it's high performance, yet it's coherent.
It's strictly consistent and it's highly durable, but it's also elastically local and where I need it to be instantly when I need it to be there. CAD cam look at public utilities and public works or one of our, one of my favorite ones is we have a customer which operates a real time crime center. I probably publicly can't say what city, but you'll figure it out really quick.
There was a car crash into a bunch of people. There was 800 plus cameras in the city that was recorded on us. Well then they said we have to add more cameras.
So one of our employees flew a additional couple pieces of equipment down to help install it by 18 hours to a snowstorm that was in the south randomly and got them installed before a major football event was happening in that town. Then the president decides they wanna show up in that city and that changes the security risk profile. But they weren't allowed to use algorithmic processing in that city.
The federal law enforcement didn't care. So how do you give access to federal law enforcement to use algorithmic processing to run counter-terrorism missions, doing image recognition and facial recognition against tens to hundreds of thousands of people during one of the most high profile and high risk events. Technologies that keep data strictly consistent and very correct, help tremendously in that.
And lastly, as we have to use the obligatory, I cannot do a presentation without using the word ai. That might be the only time you see it to be AI infrastructure is really cool. I spent a lot of time working on fun clusters that do lots of GPUs before I came to Mula.
What's interesting to me though is when I think of how most enterprises I talk to are implementing ai, it's a software package. It's a tool. I was talking to this company glean the other day.
They're neat. They import all of your unstructured data, your structured data, different systems and systems of record across the enterprise. They let people like me ask questions like, who's my most productive salesperson?
Who did the most sales calls this week? Which customer support agents having the best impact with clients and having the best results? And it analyzes the data and gives me the answers.
It's amazing stuff. Like that's really cool. I don't have to deploy 8,000 GPUs in my data center for me to use that application.
'cause the foundational model was already trained, it was tuned, it was consumed by gle. They map our data, they keep my data and my data, but let me answer questions about my business very efficiently and very effectively. If a radiologist back to your medical example, is it using an AI assistant to answer a question?
Does this picture have a met metastasizing cancer cell in it or not? Do you want it to use the most current correct image? You think it should have the latest file?
If that file changed before the agent analyzed it, should the agent know the file is updated shortly after the agent analyzed it? Wouldn't it be cool if the agent was notified that it subscribed to those changes? Or if I put payroll XLS in the wrong folder and something like glean analyzed it and then I did the right thing 30 seconds later and said, oh shoot, I really screwed that one up and I move it to my private folder.
Wouldn't it be neat if we told the indexers of the world, the file you index has moved? Can you clear your cash out now? I should probably record that that happened, but probably not distribute payroll that XLS.
When somebody said, Hey, what's payroll look like this quarter? You probably don't want that information available through the AI system, even though somebody screwed up and put the file in the wrong place. So building a strictly consistent, very correct, very durable file system.
Taking all the authentication and authorization that runs in the file system and extending it to the systems around us and making that data available to who it's supposed to be available to in its most correct form and most consistent form is incredibly important. Uh, let's see. So yep.
Legent AI is interesting. Agents depend on the truth. They also are really risky because the companies building them are really rich.
If there's a car wreck, it doesn't make the front page news. If an autonomous driving company drags somebody 20 feet, they get shut down from the city and the litigations have multiple commas in them because they're worth more in their insurance policies cover more. Mm-hmm.
It's why because of that, we're seeing unprecedented growth in the amount of data consumed. The amount of data generated. If you're running agentic AI or generative ai, not only can it create a lot of content, you probably wanna log every decision that goes into that model.
So you know why the car swerve left instead of right in a self-driving system. Seven outta the eight largest of those run on us. The data's scattered all over the place.
We have one customer who operates, uh, seven research campuses, two data centers. They have a single uh, microscope that generates 750 terabytes of information a week doing pro uh, protein modeling. They got an AI processing facility.
It's in Texas. They're not. They had to get the AI processing facility where the power was.
Now they have to get the data 350 plus petabytes, some of it to that AI processing facility, but it's also scattered across Google, Amazon, and Azure. So you have three clouds, two data centers, seven research campuses, 25 external research affiliates in an AI processing facility in Dallas. How do you get your arms around all of that data and get just the right data that you need to at the performance levels necessary to that AI facility so it can do the job to help accelerate the things they're doing, which is HIV vaccine research, cancer research, A LS, Alzheimer's and Parkinson's research.
All pretty worthwhile causes. So it was a recent paper, uh, by a former government employee talking about the need in their missions for data superiority. Who has the best data will likely win the next war.
Every U-A-S-U-A-V mission flying generates 10 to 50 terabytes of data perion. If you're flying a thousand to 1500 UAS UAVs and running three to four sorties a day, how do I handle the hundreds of petabytes a day that is generated by that? If you have an orbital or hypersonic threat, you have five seconds to map it out and determine where it's coming from, where it's going to, and determine what countermeasure to use.
How do you map the exabytes of data to plan the countermeasure and the appropriate response to minimize the impact to both military and civilian targets? Those kind of things also run on us. So unifying the data that has to become the approach.
So what do we do? Well, we said we, if we wanna support any data in any location, well that was really fancy and fluffy. That really means we wanna support any hardware and any cloud.
Give our customers the freedom of choice. That seemed like a really neat thing to do. And then in the current sort of macroeconomic trade war tariff environment, it seemed like a really, really smart thing to do because the vendor you might choose to support operations in Asia might be very different than the vendor in Europe who might be very different than the hardware vendor in the US depending upon local regulatory and tariff constructs.
Being able to normalize and give our customers the freedom of choice to use the hardware that's best for them helps tremendously. But we didn't build a system dependent upon things like, uh, sort of legacy storage, cost memory, the 3D Crosspoint memory, why it's only made in one fab in the world and Dally and PRC US government would have a problem if we were building hardware that depended upon components that stored unencrypted uncompressed data on memory made in China. That was be viewed as a significant security risk.
So we wanted to avoid things like that. Again, giving our customers choice, we say we wanna control and extend data that that means two things. It means can I treat large pools of data across different clouds and different clusters like they're one big pool.
Could I tier data from one cloud to another or one pool to another, one cluster to another, but not change the expression of the file system? To me, my F drive still looks like my F drive, even though the data's moved, locations on the backend, users machines don't have to change their workloads and workflows stay the same even though I'm putting the data where it's more efficient for my business or maybe locating it more efficient for the user to improve their experience. The inverse is also true.
Taking data that's centralized and distributing it out to the edges, allowing for a large central pool of data to be expressed in a geographically distributed manner. Like someone making a movie. The Deadpool Wolverine was an awesome movie.
Really fun to watch, awesome financially for the parent company. Disney, if that had leaked out ahead of time, that'd be a multi-hundred million dollar risk to an organization like that. Maintaining centralized control over that's very valuable.
Yet enabling a globally distributed network of VFX experts, artists, storytellers, compositors, Foley operators, making cool sounds to all work together and collaborate on that file and not have the merge conflict resolution of forking a 2030 petabyte chunk of data and putting in all different locations. It's really important. A lot of times in those VFX worlds, you have to schedule jobs a week or two in advance.
I have to know what job's gonna be done in your location. So I start moving that data there a week early. If the producer or director schedules a change, I'm either scrambling to do it or we can't work on that project.
And until we get the data moved to who's gonna work on it next. If you can extend that data geographically, you overcome those boundaries. If we can allow people to access a strictly consistent version of a file from 20 or 30 different locations concurrently without right locking each other, it changes the world from their ability to become a global company and not have to fly hard drives around the world in order to enable employees to do their jobs.
Let's see. That's why we built our cloud data fabric. It's effectively what it does.
Accelerating AI and other critical applications, performance applications. We built our neural cache. And the goal here was again, how do I achieve a really high hit rate for read cache and how to do bin packing and stuffing on right caching into the object layer effectively and efficiently.
But it's also how do you learn? Can I treat every cache hit and every cache miss like a supervised learning model and learn from it and improve the performance. And we add heuristics against 180 day time series and know what, when the fiscal period ends for a business.
'cause the data used at the end of quarter is usually very different than the data used at the beginning of quarter. And then how do we do all of this in a cloud native model to deliver cloud scale economics, elasticity of IO elasticity of data storage and elasticity of locality and geography, all without pre-provisioning, all built into true cloud native models. We're not just saying, let's grab a big chunk of EBS and throw that behind a VM and trust me it's in the cloud now.
Or even worse, taking my legacy hardware systems and bolting them into the cloud and going, I can sell you part of a virtual file system, but it'll be so prohibitively expensive, you'll never want to use it. How can we deliver a modern file system on the cloud at price points that are elastic and work the way the cloud's supposed to by using the the lowest level cloud primitives? And then when I say any server, any clouding location in six minutes, if you want, we can boot and add a new cluster in five continents.
I don't have to tell you if it's a 10 petabyte cluster or 50 terabyte cluster or 500 petabyte cluster. There's no pre-provisioning required. You simply add the data you want and you only pay for what you use.
So we had to test the performance of it. So we got it to eight terabits per second in a single cloud and over 5 million iops. I ran outta budget before I ran outta capacity, said Please guys stop.
I I actually wanna be profitable again this quarter. So we paused it at that point we figured eight terabits was enough for most of the missions we're supporting today. Now what we're gonna show first is an implementation of our cloud data fabric where we can take data across multiple public clouds and data centers and stretch and extend it to different edge devices.
An edge device that could be running on a Cisco Nexus 8,300 router, an edge device that can run on an intel knock at my house or an HPE one U server sitting in a branch office. Different performance rates, different levels of durability. Uh, the KNUCK is not exactly gonna be extremely durable, right through cache this one.
Yeah, I would, I would not trust that with uh, right, you know, a high right cycle io but really good for reading off, really good for my house. But I extend my Lightroom library to it. The municipal government use case.
It's kinda like a Swiss Army knife and we have, you know, dozens of municipal governments working on us. And what does that, what does that look like? The Splunk data could be ingested in, has wonderful observability capabilities.
Us exporting logs back to things like Splunk using open metrics and open telemetry. Who's accessing what files, who's failing to access what files? Who's trying to access what files and not doing it successfully.
Putting a nice virtuous cycle there. Rubric backup or very common target for Rubrik or Commvault or Veeam. But rubriks doing a great job Genetech for video surveillance data.
And again, extending that out to federal law enforcement or whatever's necessary or extending that to the public cloud for image recognition and pattern recognition and object recognition, ARC GIS data or all the GIS data and all the public works data, all the CAD cam data. You'll notice those land into different folders and then extend it out. We can extend across multiple autonomous systems to the architecture and engineering firms.
The workstation interacting with the CAD cam GIS folders 'cause the external architecture engineering firms is usually an external firm to most mid-size cities and smaller. It's not always in-house. Sometimes they support multiple clients.
Their ability during a emergency, during a potential natural disaster where a hurricane's hitting, being able to interact and understand the impact of flooding against the entire environment and what public facilities are impacted. Being able to work from home just as effectively as they're working from the office during those types of natural disasters is incredibly important. And then at the same time, how can we use the data center and use the cloud?
Can I use the cloud as a tier three or tier four backup up to my data center where I only have to pay for what I'm storing in it and only have to pay for the compute in it when I'm using it. As opposed to how many times have we built in our lives And I can, I've lost count now. Two tier three data centers within about a hundred kilometers of each other designed to be active.
Active or active standby. Where I have to over-provision the capacity by 50% in each one so they can handle all the workloads when they move over. We all say it's gonna work whenever every time we test it, it doesn't work.
Something invariably fails. We fight out during the drill that that we've something we haven't planned for or oops, we grew capacity too much over here. We don't have the capacity to handle it over there.
The most wonderful attribute of the cloud is elasticity. The thing you can never get on premises that the cloud delivers invariably is elasticity. I can vote a really reliable data center.
I can vote a really adorable one. I could build something that's pretty darn secure. I could build something that is geographically where I need it to be.
I could build it at high scale, but what I can't do is snap my fingers and add 5,000 virtual machines tomorrow into my data center unless I had the capacity for it already deployed. I can do that in the cloud in a few hours. So being able to leverage the elasticity of the cloud to use the cloud as a backup or tier three data center to stream all of the data to it.
Imagine the CapEx savings and the only thing you're paying for is the four or five days when you need to be active is that hurricane blows through and you've reassured your facilities. Okay, restore service locally, then you have a choice they want in the cloud, they want it local today, most customers don't have that choice. It costs too much to store the files in the cloud or I, the system I'm using here doesn't work with the one there.
It's the first time customers will have that type of choice. We have customers doing this today. We a photo photo.
The Challenge of something like that is the data gravity. I mean if you're talking petabytes of data, Hundreds of Yes. Yeah, it move very quickly Today.
What's that? It won't move very quickly. You know, the perfect time to boat a life raft before the boat takes off.
Yes. But you know, hundreds of petabytes we're talking days. Oh yeah.
5 petabytes in two days into one of our cloud offerings. Absolutely. And so you're absolutely right.
It could take days. So the right time to do it's today not when you know the hurricane's coming. But if I start today, the rate of change is the core question you're asking.
How much data changes a day? And if I only have to stream the changes even at a block level, not a full file level up, how much data changes in that enterprise today? Very rarely do we have an enterprise that's cha that's changing so much data so quickly in a day that they can't stream it to their cloud provider that day's diffs.
So you're, you're managing the change tracking as well as the uh, Full data log and extending the file system coherently between the cloud and the on-prem environment in a single global file system. Where's the log? The log stored in a distributed fashion across all of the systems in the cluster in quorum.
And if I answer that wrong, please fix me. 'cause Brandon's a lot smarter. I'm mm-hmm.
Okay. So let's take this use case though and let me ask Brandon to dive into it in more detail, walk you through the specifics. What I'm gonna ask is I need each of you to do me a favor, try to stump him, please.
Mm-hmm. Ask him the really hard question. 'cause he is a hell of a lot smarter about this than I am.
All right. I only had to learn this file stuff in the last 11 months. So long.
I've been with Defer before that. I was a networking guy for 25, 27 years. So I'm still figuring things out.
I'm glad I know what we talk log, I was thinking sis log when you first said it, sir. So I'm glad I got dodged the bullet on that one. But, uh, I'll be here to occasionally chime in and ask Brandon hard questions as well, making a little more fun, but please pick on 'em.
Okay. And then Mike's gonna stand up and come up here and do a fun demo of this. Working in a production environment, moving data back and forth, but also making changes in both the cloud and on-prem environments and stretching a file system from this little box clusters and to public clouds and showing it all working.
Interoperably. Brandon, take it away my friend. Oh, thank you.